Token billing is easiest to reason about when measurement and pricing are separate. First establish exactly how many input, output, cached, or reasoning tokens each customer consumed. Then let a pricing system decide what those quantities are worth.
Do not collapse every token into one meter
Modern model APIs price different token classes differently. A durable integration records the classes you may need to distinguish later rather than baking today's price table into application code.
Typical meters include input tokens, output tokens, cached input, model-specific premium units, and sometimes completed tasks or agent runs alongside the token counts.
Attribute usage at the expensive boundary
Emit the event where the model call is made or where the provider response is received. Include the customer/subscription context available at that point. For agents, keep a run or task identifier as a dimension so multiple model and tool calls can later be explained as one workflow.
Retries must not become revenue
If your application retries the same operation, decide whether the underlying provider actually performed the work twice. The transport retry to the meter should always be idempotent; the billable business event is a separate question.
Pricing belongs downstream
Once reliable token quantities exist, you can implement per-token rates, included allowances, model multipliers, prepaid credits, or task-based pricing. Those are rating and commercial-policy decisions.
UsageBox currently provides the meter, not that rating engine. It stores usage and monthly quantities; it does not currently maintain a prepaid-credit wallet or calculate monetary invoice lines.
Why the separation matters
Model prices change much faster than application instrumentation should. If your code emits stable usage facts, a pricing change becomes a downstream configuration problem instead of a release across every service that calls an LLM.
For implementation details, see AI API billing: tokens, credits, requests and outcomes and the usage metering hub.