pricing

package
v0.1.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 3, 2026 License: MIT Imports: 17 Imported by: 0

Documentation

Overview

Package pricing values observed token usage in micro-USD (millionths of a dollar) using an offline-first ladder of price tables.

The ladder, most authoritative first:

  1. config overrides — the user's own rates, always win. A PARTIAL override (say output only) replaces just the rates it names: the rest come from the rung below, because "override this rate" is not "delete the others";
  2. the runtime-refreshed LiteLLM table cached in the data dir;
  3. the runtime-refreshed, provider-scoped Models.dev table;
  4. the embedded LiteLLM and Models.dev snapshots for offline use;
  5. nothing. A model no rung can price stays unpriced. Zero requires an explicitly verified free service.

VENDOR COST vs ESTIMATE: Some adapters emit cost already stamped by the provider (copilot's nano-aiu, crush's session cost, pi/openclaw via the vendor, goose's provider-reported figure). This cost, stamped by the adapter, is the source of truth and is never overwritten by the price ladder — the ladder is a public-rate-card estimate of the same charge, and letting the estimate win is a strict loss of fidelity. Adapters that do not emit cost rely on pricing.Pricer to consult the ladder at collect time. An event left unpriced (Pricer returns ok=false or is nil) is stored with nil CostMicroUSD — the honest state, since a stored 0 would claim the request was free.

LONG-CONTEXT TIERS are whole-request rate cards selected by prompt size, not marginal breakpoints. Every provider measures the boundary on input tokens (cached or not) and switches cards for the FULL request — OpenAI over 272K input, Google over 200K, etc. Charged at Rates.Long if the request's ContextTokens exceed the model's threshold, never marginal. The Charge type unifies per-request and aggregated pricing, and Aggregate marks the latter so a thousand short turns can never add up to one long-context request.

EMBEDDED SNAPSHOT: the go:embed'ed filtered LiteLLM table is regenerated by internal/pricing/gensnapshot.go and is guaranteed to provide a baseline price for any model that appears in published documentation. It keeps a firewalled install pricing while offline. Partial config overrides can reach tables below the topmost rung (overrides fill missing fields from the rung they displace), and the snapshot is the floor: a model it cannot price remains unpriced.

Integer micro-USD is the unit everywhere: it sums exactly across millions of rows, where float dollars drift. The package imports only model.

Index

Constants

View Source
const ModelsDevURL = "https://models.dev/api.json"
View Source
const RefreshURL = "https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json"

RefreshURL is the upstream LiteLLM price table. Hardcoded on purpose: the config exposes only an off switch and per-model overrides, which already cover the air-gapped and the disagrees-with-the-table cases without turning the price feed into an arbitrary user-controlled fetch target.

View Source
const SourceOverride = "override"

SourceOverride is the price_source stamped when a config override priced the event. The other two rungs stamp "litellm-<fetch date>" and "embedded-<snapshot date>", built by the table loaders.

Variables

This section is empty.

Functions

func BuildModelsDevSnapshot

func BuildModelsDevSnapshot(data []byte) ([]byte, error)

BuildModelsDevSnapshot filters Models.dev API data into the same envelope consumed by the runtime cache. Used by the offline snapshot generator.

func SnapshotField

func SnapshotField(key string) bool

SnapshotField reports whether a LiteLLM price key is one this decoder reads, i.e. one the embedded snapshot must KEEP. It is exported for the snapshot generator (gensnapshot.go), which is the only caller: the generator's filter stripping a field the decoder reads is exactly the bug that left every air-gapped install flat-pricing long-context turns, and one predicate for both halves makes the two unable to disagree.

Types

type Charge

type Charge struct {
	Model       string
	Provider    string
	ServiceTier string

	Input     int64
	Output    int64
	Reasoning int64
	CacheRead int64
	// CacheWrite5m / CacheWrite1h split the cache-creation tokens by the
	// requested entry lifetime. A source that reports no split puts everything
	// in CacheWrite5m, which is the documented fallback (the 1h write is the
	// dearer of the two, so this never over-bills).
	CacheWrite5m int64
	CacheWrite1h int64

	// AdditiveReasoning bills Reasoning at the output rate on top of Output.
	// False means reasoning is already contained in Output (see
	// model.ReasoningModeFor) and billing it again would double charge.
	AdditiveReasoning bool

	// Aggregate marks a charge that stands for SEVERAL requests summed together
	// — report's display-time valuation of unpriced groups, which arrives
	// pre-aggregated per model — rather than describing one. Its token counts
	// are a total, not a prompt size, so no long-context tier is selected for
	// it: a thousand short turns must never add up to one long-context request.
	// The zero value is therefore "one request", which is what every ledger path
	// prices.
	Aggregate bool
}

Charge is the token shape being priced: the counts, plus the three facts that change which rate applies (the model/provider identity, the service tier, and whether reasoning tokens are billed on top of output or already inside it).

func ChargeFor

func ChargeFor(e model.UsageEvent) Charge

ChargeFor derives the Charge for a stored usage event: token counts straight off the event, the reasoning-billing rule from the event's tool, and the cache-write split from the transient CacheTTL enrichment. A split that does not add up to the recorded cache-creation count is discarded in favour of "all 5m" — a partial split would silently drop billable tokens.

func (Charge) ContextTokens

func (c Charge) ContextTokens() int64

ContextTokens is the prompt size the long-context tier is measured against: every bucket that is part of the request's input, CACHED OR NOT. Output is excluded because every provider measures the boundary on the prompt (OpenAI: ">272K input tokens"; Google: "query input context longer than 200K tokens").

Cache reads are the load-bearing term. The adapters normalize a provider's prompt into disjoint buckets — codex stores `input = raw_input - cached` because the Responses API reports cached_tokens as a SUBSET of input_tokens — so Input + CacheRead is exactly the number OpenAI's own sentence is about, and counting Input alone would be reading a normalization artifact as a prompt size. Measured on this machine's ledger: 17 rows cross 272K on uncached input, 551 cross on input + cache read. Cache writes join them for the same reason — Anthropic's context-window doc counts input_tokens, cache_read_input_tokens and cache_creation_input_tokens alike toward the window, and on this ledger they add no crossings on top.

func (Charge) Tokens

func (c Charge) Tokens() int64

Tokens is the total token count being charged, used to tell "no usage, so genuinely zero" apart from "priced at zero", which is never reported.

type Engine

type Engine struct {
	// contains filtered or unexported fields
}

Engine resolves charges through the ladder. It is safe for concurrent use: Refresh swaps the middle rung under a write lock while Price reads it.

func New

func New(opt Options) *Engine

New loads embedded tables and local caches without touching the network. Models.dev uses only a fresh valid cache; otherwise its embedded table serves.

func (*Engine) Price

func (e *Engine) Price(c Charge) (int64, string, bool)

Price resolves a charge against the ladder and returns the cost in micro-USD plus the price_source that produced it. ok=false means unpriced: NO rung could value this charge shape. A rung that knows the model but publishes no rate for the tokens being charged (an output-only row against an input-only charge) is a gap in that table, not a free request, and gets the same recovery an unpriceable row gets — the next rung down is tried. Only when the ladder runs out is the charge reported unpriced, never $0.00.

A charge billed off the model's long-context card carries longContextStamp on the end of whatever the rung stamped.

func (*Engine) PriceEvent

func (e *Engine) PriceEvent(ev model.UsageEvent) (int64, string, bool)

PriceEvent prices a usage event. It is the collector-facing entry point and satisfies collect.Pricer.

func (*Engine) Refresh

func (e *Engine) Refresh(ctx context.Context) error

Refresh independently updates the LiteLLM and Models.dev feeds when due. Disabled refresh or an empty data dir prevents all downloads. Models.dev attempts at most once per 24 hours per engine, reuses a fresh cache across restarts, and falls back to its embedded snapshot on download/parse failure. LiteLLM retains its last loaded table on failure. Errors do not stop pricing.

func (*Engine) Revision

func (e *Engine) Revision() string

Revision identifies the loaded rates. Collectors use it to retry older unpriced usage only when the pricing information changes.

type LongContext

type LongContext struct {
	// Threshold is the prompt size, in tokens, above which this card applies.
	// 0 means the model publishes no long-context tier.
	Threshold int64

	Input        float64
	Output       float64
	CacheRead    float64
	CacheWrite5m float64
	CacheWrite1h float64
}

LongContext is a model's above-threshold rate card. Providers switch cards for the WHOLE request rather than billing a band of tokens separately — OpenAI: "Prompts with >272K input tokens are priced at 2x input and 1.5x output for the full request"; Google: "If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates" — so this is a rate card selected per request, never a marginal breakpoint. Reading LiteLLM's "_above_200k_tokens" field naming as marginal is an artifact of the name and matches no documented provider behaviour; it also under-bills by a factor of ten on real traffic, because a crossing request is typically broad (a large cache read plus a modest input) rather than one bucket deep past the line.

Threshold is per model and comes from the LiteLLM field name (see longContextCard): the live table publishes five distinct boundaries — 128K, 200K, 256K, 272K and 512K — so a hardcoded 200K would be wrong for four of them. A zero rate here falls back to the same-named base rate, except an absent 1h write rate uses the resolved long-context 5m write rate: LiteLLM publishes an above-threshold rate only for the buckets it prices, and inventing a multiple of the base rate for the rest would be a made-up number.

type Options

type Options struct {
	// DataDir holds the refreshed table's cache. Empty disables both reading
	// the cache and refreshing it, leaving overrides plus the embedded floor.
	DataDir string
	// Refresh enables the runtime refresh (config pricing.refresh, default
	// true). It gates the network only: an existing cache is still read.
	Refresh bool
	// Overrides are per-model rates that beat every table.
	Overrides map[string]Rates
}

Options configures the pricing engine.

type Rates

type Rates struct {
	// Free distinguishes a verified free service from an unpublished zero rate.
	Free         bool
	Input        float64
	Output       float64
	CacheRead    float64 // 0 -> the input rate: a cache read is a discounted input token
	CacheWrite5m float64 // 0 -> the input rate: writing a 5m entry costs at least the input
	CacheWrite1h float64 // 0 -> the resolved 5m write rate
	InputBatch   float64 // 0 -> the standard input rate (no batch tier published)
	OutputBatch  float64 // 0 -> the standard output rate

	// Long is the second rate card, used for a request whose prompt crosses
	// this model's published threshold. Zero Threshold means the model has no
	// long-context tier, which is the common case.
	Long LongContext
}

Rates are the per-token USD rates for one model, in LiteLLM's units (dollars per single token). A zero field means "not published", not "free"; Cost resolves each zero to the documented fallback rate.

func (Rates) Cost

func (r Rates) Cost(c Charge) int64

Cost values a charge at these rates, in micro-USD, rounding half up. The result is never negative and saturates at math.MaxInt64 micro-USD.

A request over the model's long-context threshold is billed ENTIRELY off the second rate card — every bucket, not the excess above the line.

func (Rates) Priceable

func (r Rates) Priceable() bool

Priceable reports whether the rates carry a price or a verified free status. An ordinary all-zero entry remains unknown.

type Table

type Table struct {
	Source string
	Models map[string]Rates
	// ProviderScoped tables require the recorded provider as well as the model.
	ProviderScoped bool
}

Table is one rung of the ladder: a model->rates map plus the price_source string to stamp when it prices something.

func (*Table) Lookup

func (t *Table) Lookup(provider, name string) (Rates, bool)

Lookup resolves rates for a (provider, model) pair, trying the provider's namespaced LiteLLM key before the bare model id so a proxied model is priced at the proxy's rates when the table publishes them. A table row with no usable price counts as a miss so the next rung gets a chance.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL