What is AI usage-based billing?
AI usage-based billing charges customers for what an AI product actually consumes rather than a flat subscription fee. Each request is recorded as a usage event, priced against a rate, and either added to an invoice or deducted from a prepaid balance. Common billable units include tokens, API calls, inference requests, compute time and agent actions.
It's the pricing model behind most modern AI products, and it anchors any serious AI monetization strategy. The rest of this guide breaks down what you can meter, how to shape pricing around those units, and how to keep usage and revenue reconciled.
What You Can Meter in an AI Product
The first decision in usage-based pricing is which unit you charge for. Most AI products have several candidates, and the best one tracks two things at once: the value a customer gets and the cost you carry to serve them.
Tokens. The standard unit for large language model features. Most products count both input tokens (the prompt) and output tokens (the response), often at different rates, because output is more expensive to generate.
API calls. A count of requests to an endpoint, useful when the work per call is fairly uniform or when customers think in terms of requests rather than tokens.
Compute and GPU-hours. For training jobs, fine-tuning or self-hosted inference, you can bill the GPU time or compute-seconds consumed, which maps closely to your underlying infrastructure cost.
Inference requests. A charge per prediction, generation or classification, common for vision, speech and embedding models where a single call produces one discrete result.
Storage. Vector databases, document indexes and retrieval layers consume storage that grows over time, so gigabytes stored or vectors indexed can be a billable line of its own.
Agent actions and outcomes. As products move toward autonomous agents, you can meter discrete actions an agent takes, or price on outcomes such as a resolved ticket or a completed task, tying the bill to delivered results rather than raw compute.
Common AI Billing Metrics
The table below maps the most common metrics to the unit you bill and a representative use case. Most products combine two or three of these rather than relying on any single one.
| Metric | Unit billed | Example use case |
|---|---|---|
| Tokens | Per 1K or 1M input/output tokens | LLM chat, summarisation and content generation APIs |
| API calls | Per request | Classification, moderation or lookup endpoints with uniform work |
| GPU time | Per GPU-hour or compute-second | Model training, fine-tuning and self-hosted inference |
| Inference requests | Per prediction or generation | Image generation, speech-to-text and embeddings |
| Storage | Per GB stored or vectors indexed | Vector databases and retrieval-augmented generation |
| Agent actions | Per action or completed outcome | Autonomous agents resolving tasks or tickets |
Designing Usage-Based Pricing for AI
Once you know what to meter, the next step is turning units into a price that is fair, predictable and profitable. A few design choices do most of the work.
Choose the billing unit deliberately
Pick the unit that customers understand and that tracks your cost. Tokens suit raw language models; outcomes or actions suit agents. If the unit is too abstract, customers struggle to forecast spend and trust erodes.
Use tiers and volume rates
Tiered and volume pricing lower the unit rate as consumption grows, rewarding larger customers and giving them a predictable marginal cost. Committed-use discounts trade a lower rate for a minimum monthly commitment, which smooths your revenue.
Offer prepaid credits
Selling a balance of credits up front brings in cash early, caps your exposure to runaway usage, and gives customers a clear budget. The balance draws down as usage happens, which is why credits pair so well with real-time balance management. Our guide to AI credits and prepaid billing goes deeper on this model.
Plan for overage and allowances
Many AI features ship with a free or included monthly allowance and charge overage beyond it. This lowers the barrier to adoption while still capturing revenue from heavy users. A recurring platform fee plus metered overage is one of the most common hybrid structures for embedded AI.
Avoiding Revenue Leakage
Usage-based pricing only pays off if every billable event actually reaches the invoice. At AI volumes, small gaps between usage and billing compound quickly, so the controls below matter as much as the pricing itself.
Meter in real time
Collect and price events as they happen rather than batching them at month-end. Real-time metering keeps running totals, balances and dashboards accurate, and it is the only way to enforce a prepaid balance before a customer over-consumes. The mechanics of this pipeline are covered in our AI and token-based billing guide.
Enforce quotas and limits
Rate limits, spend caps and quotas protect both sides: they stop a misbehaving integration from running up an unpayable bill, and they let you cut off or throttle usage the moment a credit balance or plan limit is reached.
Reconcile usage against billing
Deduplicate retried events, validate that nothing is missing, and make sure the number in the usage dashboard, the number on the invoice and the number in the ledger all agree. Clean reconciliation is what turns metered usage into revenue you can recognise and audit. For language-model specifics, see our guide to LLM billing, and the broader AI monetization platform that brings metering, rating and balances together.
Frequently Asked Questions
What is AI usage-based billing?
AI usage-based billing charges customers for what an AI product actually consumes rather than a flat subscription fee. Each request is recorded as a usage event, priced against a rate, and either added to an invoice or deducted from a prepaid balance. Common billable units include tokens, API calls, inference requests, compute time and agent actions.
How do you meter AI usage?
You meter AI usage by emitting a usage event for every request that records the customer, the model or service, a timestamp and the billable quantity, such as input and output tokens, API calls, GPU-seconds, inference requests, storage or agent actions. Those events are collected, deduplicated and normalised into a single format, then priced by a rating engine. At AI volumes, metering needs to run in real time so running totals, balances and quotas stay accurate between invoices.
How is usage-based billing different from subscriptions?
A subscription charges a fixed recurring amount regardless of how much the customer uses. Usage-based billing charges for actual consumption, so the bill rises and falls with the volume of tokens, calls or compute consumed. Usage-based models require event-level metering and real-time rating, while many AI products blend the two with a recurring platform fee plus metered usage and overage on top.
Ready to Bill Your AI Product by Usage?
Talk to our billing experts about metering tokens, calls, compute and agent actions, and pricing AI consumption in real time with EarnBill.