Documentation
¶
Overview ¶
Package pricing values observed token usage in micro-USD (millionths of a dollar) using an offline-first ladder of price tables.
The ladder, most authoritative first:
- config overrides — the user's own rates, always win. A PARTIAL override (say output only) replaces just the rates it names: the rest come from the rung below, because "override this rate" is not "delete the others";
- the runtime-refreshed LiteLLM table cached in the data dir;
- the runtime-refreshed, provider-scoped Models.dev table;
- the embedded LiteLLM and Models.dev snapshots for offline use;
- nothing. A model no rung can price stays unpriced. Zero requires an explicitly verified free service.
VENDOR COST vs ESTIMATE: Some adapters emit cost already stamped by the provider (copilot's nano-aiu, crush's session cost, pi/openclaw via the vendor, goose's provider-reported figure). This cost, stamped by the adapter, is the source of truth and is never overwritten by the price ladder — the ladder is a public-rate-card estimate of the same charge, and letting the estimate win is a strict loss of fidelity. Adapters that do not emit cost rely on pricing.Pricer to consult the ladder at collect time. An event left unpriced (Pricer returns ok=false or is nil) is stored with nil CostMicroUSD — the honest state, since a stored 0 would claim the request was free.
LONG-CONTEXT TIERS are whole-request rate cards selected by prompt size, not marginal breakpoints. Every provider measures the boundary on input tokens (cached or not) and switches cards for the FULL request — OpenAI over 272K input, Google over 200K, etc. Charged at Rates.Long if the request's ContextTokens exceed the model's threshold, never marginal. The Charge type unifies per-request and aggregated pricing, and Aggregate marks the latter so a thousand short turns can never add up to one long-context request.
EMBEDDED SNAPSHOT: the go:embed'ed filtered LiteLLM table is regenerated by internal/pricing/gensnapshot.go and is guaranteed to provide a baseline price for any model that appears in published documentation. It keeps a firewalled install pricing while offline. Partial config overrides can reach tables below the topmost rung (overrides fill missing fields from the rung they displace), and the snapshot is the floor: a model it cannot price remains unpriced.
Integer micro-USD is the unit everywhere: it sums exactly across millions of rows, where float dollars drift. The package imports only model.
Index ¶
Constants ¶
const ModelsDevURL = "https://models.dev/api.json"
const RefreshURL = "https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json"
RefreshURL is the upstream LiteLLM price table. Hardcoded on purpose: the config exposes only an off switch and per-model overrides, which already cover the air-gapped and the disagrees-with-the-table cases without turning the price feed into an arbitrary user-controlled fetch target.
const SourceOverride = "override"
SourceOverride is the price_source stamped when a config override priced the event. The other two rungs stamp "litellm-<fetch date>" and "embedded-<snapshot date>", built by the table loaders.
Variables ¶
This section is empty.
Functions ¶
func BuildModelsDevSnapshot ¶
BuildModelsDevSnapshot filters Models.dev API data into the same envelope consumed by the runtime cache. Used by the offline snapshot generator.
func SnapshotField ¶
SnapshotField reports whether a LiteLLM price key is one this decoder reads, i.e. one the embedded snapshot must KEEP. It is exported for the snapshot generator (gensnapshot.go), which is the only caller: the generator's filter stripping a field the decoder reads is exactly the bug that left every air-gapped install flat-pricing long-context turns, and one predicate for both halves makes the two unable to disagree.
Types ¶
type Charge ¶
type Charge struct {
Model string
Provider string
ServiceTier string
Input int64
Output int64
Reasoning int64
CacheRead int64
// CacheWrite5m / CacheWrite1h split the cache-creation tokens by the
// requested entry lifetime. A source that reports no split puts everything
// in CacheWrite5m, which is the documented fallback (the 1h write is the
// dearer of the two, so this never over-bills).
CacheWrite5m int64
CacheWrite1h int64
// AdditiveReasoning bills Reasoning at the output rate on top of Output.
// False means reasoning is already contained in Output (see
// model.ReasoningModeFor) and billing it again would double charge.
AdditiveReasoning bool
// Aggregate marks a charge that stands for SEVERAL requests summed together
// — report's display-time valuation of unpriced groups, which arrives
// pre-aggregated per model — rather than describing one. Its token counts
// are a total, not a prompt size, so no long-context tier is selected for
// it: a thousand short turns must never add up to one long-context request.
// The zero value is therefore "one request", which is what every ledger path
// prices.
Aggregate bool
}
Charge is the token shape being priced: the counts, plus the three facts that change which rate applies (the model/provider identity, the service tier, and whether reasoning tokens are billed on top of output or already inside it).
func ChargeFor ¶
func ChargeFor(e model.UsageEvent) Charge
ChargeFor derives the Charge for a stored usage event: token counts straight off the event, the reasoning-billing rule from the event's tool, and the cache-write split from the transient CacheTTL enrichment. A split that does not add up to the recorded cache-creation count is discarded in favour of "all 5m" — a partial split would silently drop billable tokens.
func (Charge) ContextTokens ¶
ContextTokens is the prompt size the long-context tier is measured against: every bucket that is part of the request's input, CACHED OR NOT. Output is excluded because every provider measures the boundary on the prompt (OpenAI: ">272K input tokens"; Google: "query input context longer than 200K tokens").
Cache reads are the load-bearing term. The adapters normalize a provider's prompt into disjoint buckets — codex stores `input = raw_input - cached` because the Responses API reports cached_tokens as a SUBSET of input_tokens — so Input + CacheRead is exactly the number OpenAI's own sentence is about, and counting Input alone would be reading a normalization artifact as a prompt size. Measured on this machine's ledger: 17 rows cross 272K on uncached input, 551 cross on input + cache read. Cache writes join them for the same reason — Anthropic's context-window doc counts input_tokens, cache_read_input_tokens and cache_creation_input_tokens alike toward the window, and on this ledger they add no crossings on top.
type Engine ¶
type Engine struct {
// contains filtered or unexported fields
}
Engine resolves charges through the ladder. It is safe for concurrent use: Refresh swaps the middle rung under a write lock while Price reads it.
func New ¶
New loads embedded tables and local caches without touching the network. Models.dev uses only a fresh valid cache; otherwise its embedded table serves.
func (*Engine) Price ¶
Price resolves a charge against the ladder and returns the cost in micro-USD plus the price_source that produced it. ok=false means unpriced: NO rung could value this charge shape. A rung that knows the model but publishes no rate for the tokens being charged (an output-only row against an input-only charge) is a gap in that table, not a free request, and gets the same recovery an unpriceable row gets — the next rung down is tried. Only when the ladder runs out is the charge reported unpriced, never $0.00.
A charge billed off the model's long-context card carries longContextStamp on the end of whatever the rung stamped.
func (*Engine) PriceEvent ¶
PriceEvent prices a usage event. It is the collector-facing entry point and satisfies collect.Pricer.
func (*Engine) PriceStoredEvent ¶ added in v0.1.1
PriceStoredEvent prices an event without assuming the transient CacheTTL split survived storage. Cache writes are priced only when both possible TTL extremes produce the same cost and provenance from one table snapshot. This preserves verified free and equal-rate prices; ambiguous costs stay unknown. Collection uses this optional method for historical price sync. PriceEvent retains its existing behavior for fresh observations with source enrichment.
func (*Engine) Refresh ¶
Refresh independently updates the LiteLLM and Models.dev feeds when due. Disabled refresh or an empty data dir prevents all downloads. Models.dev attempts at most once per 24 hours per engine, reuses a fresh cache across restarts, and falls back to its embedded snapshot on download/parse failure. LiteLLM retains its last loaded table on failure. Errors do not stop pricing.
type LongContext ¶
type LongContext struct {
// Threshold is the prompt size, in tokens, above which this card applies.
// 0 means the model publishes no long-context tier.
Threshold int64
Input float64
Output float64
CacheRead float64
CacheWrite5m float64
CacheWrite1h float64
}
LongContext is a model's above-threshold rate card. Providers switch cards for the WHOLE request rather than billing a band of tokens separately — OpenAI: "Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request"; Google: "If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates" — so this is a rate card selected per request, never a marginal breakpoint. Reading LiteLLM's "_above_200k_tokens" field naming as marginal is an artifact of the name and matches no documented provider behaviour; it also under-bills by a factor of ten on real traffic, because a crossing request is typically broad (a large cache read plus a modest input) rather than one bucket deep past the line.
Threshold is per model and comes from the LiteLLM field name (see longContextCard): the live table publishes five distinct boundaries — 128K, 200K, 256K, 272K and 512K — so a hardcoded 200K would be wrong for four of them. A zero rate here falls back to the same-named base rate, except an absent 1h write rate uses the resolved long-context 5m write rate: LiteLLM publishes an above-threshold rate only for the buckets it prices, and inventing a multiple of the base rate for the rest would be a made-up number.
type Options ¶
type Options struct {
// DataDir holds the refreshed table's cache. Empty disables both reading
// the cache and refreshing it, leaving overrides plus the embedded floor.
DataDir string
// Refresh enables the runtime refresh (config pricing.refresh, default
// true). It gates the network only: an existing cache is still read.
Refresh bool
// Overrides are per-model rates that beat every table.
Overrides map[string]Rates
}
Options configures the pricing engine.
type Rates ¶
type Rates struct {
// Free distinguishes a verified free service from an unpublished zero rate.
Free bool
Input float64
Output float64
CacheRead float64 // 0 -> the input rate: a cache read is a discounted input token
CacheWrite5m float64 // 0 -> the input rate: writing a 5m entry costs at least the input
CacheWrite1h float64 // 0 -> the resolved 5m write rate
InputBatch float64 // 0 -> the standard input rate (no batch tier published)
OutputBatch float64 // 0 -> the standard output rate
// Long is the second rate card, used for a request whose prompt crosses
// this model's published threshold. Zero Threshold means the model has no
// long-context tier, which is the common case.
Long LongContext
}
Rates are the per-token USD rates for one model, in LiteLLM's units (dollars per single token). A zero field means "not published", not "free"; Cost resolves each zero to the documented fallback rate.
func (Rates) Cost ¶
Cost values a charge at these rates, in micro-USD, rounding half up. The result is never negative and saturates at math.MaxInt64 micro-USD.
A request over the model's long-context threshold is billed ENTIRELY off the second rate card — every bucket, not the excess above the line.
type Table ¶
type Table struct {
Source string
Models map[string]Rates
// ProviderScoped tables require the recorded provider as well as the model.
ProviderScoped bool
}
Table is one rung of the ladder: a model->rates map plus the price_source string to stamp when it prices something.
func (*Table) Lookup ¶
Lookup resolves rates for a (provider, model) pair, trying the provider's namespaced LiteLLM key before the bare model id so a proxied model is priced at the proxy's rates when the table publishes them. A table row with no usable price counts as a miss so the next rung gets a chance.