Documentation
¶
Overview ¶
Package calc holds the pure token and cost arithmetic that runs between an OpenRouter response and a persisted assistant message: it normalises the provider usage block, prices a call from the model catalog, and derives the compaction budget and its watermarks from the model limits and the compaction config.
Index ¶
- Constants
- func EffectiveInputCapacity(input UsableInput) float64
- func IsOverflow(input OverflowInput) bool
- func MaxOutputTokens(model Model) float64
- func SetModuleEnvForTesting(env map[string]string) func()
- func ValidatePolicy(cfg Config) error
- type CacheCost
- type CompactionConfig
- type CompactionWatermarks
- type Config
- type GetUsageInput
- type InputTokenDetails
- type LanguageModelUsage
- type LanguageModelV3InputTokens
- type LanguageModelV3OutputTokens
- type LanguageModelV3Usage
- type Model
- type ModelCapabilities
- type ModelCost
- type ModelLimit
- type OpenRouterCompletionTokensDetails
- type OpenRouterPromptTokensDetails
- type OpenRouterUsage
- type OutputTokenDetails
- type Over200KCost
- type OverflowInput
- type ProviderMetadata
- type TokenCache
- type Tokens
- type UsableInput
- type UsageCache
- type UsageResult
- type UsageTokens
Constants ¶
const COMPACTION_BUFFER float64 = 20_000
COMPACTION_BUFFER bounds the output reservation taken off the window.
const COMPACTION_LOW_TO_HIGH_RATIO float64 = 2.0 / 3.0
COMPACTION_LOW_TO_HIGH_RATIO is the low-watermark half of the 40/60 hysteresis: low is two-thirds of high.
const DefaultCapacityTokens float64 = 500_000
DefaultCapacityTokens caps the working set when no capacity_tokens is configured. A model with a smaller window is bounded by the window.
const OUTPUT_TOKEN_MAX_DEFAULT float64 = 32_000
OUTPUT_TOKEN_MAX_DEFAULT is the output-token ceiling when SENIOR_DEV_OUTPUT_TOKEN_MAX is unset.
const PolicyWindow = "window"
PolicyWindow is the compaction policy: the window is the budget.
const TRIGGER_PCT float64 = 0.6
TRIGGER_PCT is the fraction of the capacity at which auto-compaction fires: the high watermark.
Variables ¶
This section is empty.
Functions ¶
func EffectiveInputCapacity ¶
func EffectiveInputCapacity(input UsableInput) float64
EffectiveInputCapacity is the model-visible input capacity after reserving output space and applying the capacity cap. It deliberately does not apply the trigger percentage; Watermarks derives both marks from this one underlying capacity.
func IsOverflow ¶
func IsOverflow(input OverflowInput) bool
IsOverflow reports whether the assistant's token count has reached the high watermark.
func MaxOutputTokens ¶
MaxOutputTokens is `min(model.limit.output, OUTPUT_TOKEN_MAX)`, falling back to OUTPUT_TOKEN_MAX when the minimum is 0 or NaN. A negative limit.output is passed through as is.
func SetModuleEnvForTesting ¶
SetModuleEnvForTesting re-runs the package-init evaluation of OUTPUT_TOKEN_MAX against the supplied environment. Returns a restore func.
func ValidatePolicy ¶
ValidatePolicy refuses a policy name this binary does not implement, and a budget field outside its range. Both are refused at config load so a misspelled block fails before any model call.
Types ¶
type CompactionConfig ¶
type CompactionConfig struct {
// Policy names how the compaction budget is derived. The only policy is
// "window": the budget is the model's own context window, capped by
// CapacityTokens. It may be spelled out or left empty; any other name is
// refused by ValidatePolicy.
Policy string `json:"policy,omitempty"`
Auto *bool `json:"auto"`
Prune *bool `json:"prune"`
// PreserveRecentTokens overrides the verbatim tail budget a compaction
// keeps ahead of the summary. The tail is sized in tokens after
// truncation, never in turns.
PreserveRecentTokens *float64 `json:"preserve_recent_tokens"`
// PreserveRecentFraction sizes the verbatim tail as a fraction of the
// high watermark instead of a fixed token count, so it scales with the
// window. PreserveRecentTokens wins when both are set.
PreserveRecentFraction *float64 `json:"preserve_recent_fraction"`
// CapacityTokens caps the working set below the model's window: a cost
// decision, or a model known to degrade before its advertised context.
// Absent means DefaultCapacityTokens.
CapacityTokens *float64 `json:"capacity_tokens"`
Reserved *float64 `json:"reserved"`
}
CompactionConfig is the `compaction` block of project config. Every field is optional, so every field is a pointer: `auto` is tested strictly (an absent value is NOT false) and `reserved` nullishly (an explicit 0 wins).
type CompactionWatermarks ¶
CompactionWatermarks describes the preferred post-compaction target and the occupancy at which another compaction becomes necessary.
func Watermarks ¶
func Watermarks(input UsableInput) CompactionWatermarks
Watermarks returns the 40/60 hysteresis around the capacity.
type Config ¶
type Config struct {
Compaction *CompactionConfig `json:"compaction"`
}
Config is the project-config projection this package needs: only the `compaction` block is read.
type GetUsageInput ¶
type GetUsageInput struct {
Model Model
Usage LanguageModelUsage
Metadata ProviderMetadata
}
GetUsageInput is GetUsage's parameter object.
type InputTokenDetails ¶
type InputTokenDetails struct {
NoCacheTokens *float64 `json:"noCacheTokens,omitempty"`
CacheReadTokens *float64 `json:"cacheReadTokens,omitempty"`
CacheWriteTokens *float64 `json:"cacheWriteTokens,omitempty"`
}
InputTokenDetails is the flattened input breakdown.
type LanguageModelUsage ¶
type LanguageModelUsage struct {
InputTokens *float64 `json:"inputTokens,omitempty"`
InputTokenDetails *InputTokenDetails `json:"inputTokenDetails,omitempty"`
OutputTokens *float64 `json:"outputTokens,omitempty"`
OutputTokenDetails *OutputTokenDetails `json:"outputTokenDetails,omitempty"`
TotalTokens *float64 `json:"totalTokens,omitempty"`
Raw json.RawMessage `json:"raw,omitempty"`
ReasoningTokens *float64 `json:"reasoningTokens,omitempty"`
CachedInputTokens *float64 `json:"cachedInputTokens,omitempty"`
}
LanguageModelUsage is the flattened usage GetUsage consumes. The details blocks are optional; the two trailing fields are flat aliases GetUsage falls back to when the details are absent.
func AsLanguageModelUsage ¶
func AsLanguageModelUsage(usage LanguageModelV3Usage) LanguageModelUsage
AsLanguageModelUsage flattens the token groups. totalTokens is recomputed as input + output; the provider's own total survives only inside Raw.
type LanguageModelV3InputTokens ¶
type LanguageModelV3InputTokens struct {
Total *float64 `json:"total,omitempty"`
NoCache *float64 `json:"noCache,omitempty"`
CacheRead *float64 `json:"cacheRead,omitempty"`
CacheWrite *float64 `json:"cacheWrite,omitempty"`
}
LanguageModelV3InputTokens is the provider-side `inputTokens` block.
type LanguageModelV3OutputTokens ¶
type LanguageModelV3OutputTokens struct {
Total *float64 `json:"total,omitempty"`
Text *float64 `json:"text,omitempty"`
Reasoning *float64 `json:"reasoning,omitempty"`
}
LanguageModelV3OutputTokens is the provider-side `outputTokens` block.
type LanguageModelV3Usage ¶
type LanguageModelV3Usage struct {
InputTokens LanguageModelV3InputTokens `json:"inputTokens"`
OutputTokens LanguageModelV3OutputTokens `json:"outputTokens"`
Raw json.RawMessage `json:"raw,omitempty"`
}
LanguageModelV3Usage is ComputeTokenUsage's return shape.
func ComputeTokenUsage ¶
func ComputeTokenUsage(usage *OpenRouterUsage) LanguageModelV3Usage
ComputeTokenUsage splits the provider's usage block into input and output token groups. An absent cache-write count stays absent (nil) so that the provider-metadata fallbacks in GetUsage can still supply it.
type Model ¶
type Model struct {
Cost *ModelCost `json:"cost"`
Limit ModelLimit `json:"limit"`
Capabilities ModelCapabilities `json:"-"`
}
Model is the catalog projection this package needs: the limit block (for the budget and the output reservation) and the cost block (for usage).
type ModelCapabilities ¶
type ModelCapabilities struct {
Attachment bool `json:"attachment"`
Reasoning bool `json:"reasoning"`
Temperature bool `json:"temperature"`
ToolCall bool `json:"toolcall"`
Input map[string]bool `json:"input"`
Output map[string]bool `json:"output"`
}
ModelCapabilities is the models.dev capability slice retained alongside cost and limits so provider request assembly does not invent support.
type ModelCost ¶
type ModelCost struct {
Input float64 `json:"input"`
Output float64 `json:"output"`
Cache *CacheCost `json:"cache"`
ExperimentalOver200K *Over200KCost `json:"experimentalOver200K"`
}
ModelCost is a model's price block. `cache` is optional in practice, so it is a pointer.
type ModelLimit ¶
type ModelLimit struct {
Context float64 `json:"context"`
Input *float64 `json:"input"`
Output float64 `json:"output"`
}
ModelLimit is a model's context, input and output limits. Input is optional: a catalog entry that names none is budgeted from Context.
type OpenRouterCompletionTokensDetails ¶
type OpenRouterCompletionTokensDetails struct {
ReasoningTokens *float64 `json:"reasoning_tokens"`
}
OpenRouterCompletionTokensDetails is `usage.completion_tokens_details`.
type OpenRouterPromptTokensDetails ¶
type OpenRouterPromptTokensDetails struct {
CachedTokens *float64 `json:"cached_tokens"`
CacheWriteTokens *float64 `json:"cache_write_tokens"`
}
OpenRouterPromptTokensDetails is `usage.prompt_tokens_details`.
type OpenRouterUsage ¶
type OpenRouterUsage struct {
PromptTokens *float64 `json:"prompt_tokens"`
CompletionTokens *float64 `json:"completion_tokens"`
PromptTokensDetails *OpenRouterPromptTokensDetails `json:"prompt_tokens_details"`
CompletionTokensDetails *OpenRouterCompletionTokensDetails `json:"completion_tokens_details"`
// Raw is the untouched wire object, including fields such as `cost` and
// `is_byok` that senior-dev never reads.
Raw json.RawMessage `json:"-"`
}
OpenRouterUsage is the numeric projection of OpenRouter's `usage` object that the token arithmetic reads, plus the untouched original carried through as Raw.
func (*OpenRouterUsage) UnmarshalJSON ¶
func (u *OpenRouterUsage) UnmarshalJSON(data []byte) error
UnmarshalJSON decodes the numeric projection and keeps the original bytes.
type OutputTokenDetails ¶
type OutputTokenDetails struct {
TextTokens *float64 `json:"textTokens,omitempty"`
ReasoningTokens *float64 `json:"reasoningTokens,omitempty"`
}
OutputTokenDetails is the flattened output breakdown.
type Over200KCost ¶
type Over200KCost struct {
Cache *CacheCost `json:"cache"`
Input float64 `json:"input"`
Output float64 `json:"output"`
}
Over200KCost is the price block a provider applies above 200K context. The key order is cache, input, output.
type OverflowInput ¶
OverflowInput is what the trigger decides on.
type ProviderMetadata ¶
ProviderMetadata is the per-provider metadata map a response may carry.
type TokenCache ¶
TokenCache is the persisted cache-token pair, declared read then write. Contrast UsageCache, which is the same data in the order usage builds it.
type Tokens ¶
type Tokens struct {
Total *float64 `json:"total,omitempty"`
Input float64 `json:"input"`
Output float64 `json:"output"`
Reasoning float64 `json:"reasoning"`
Cache TokenCache `json:"cache"`
}
Tokens is the persisted assistant token block. `total` is optional, so it is a pointer and the key is dropped when it is absent.
type UsableInput ¶
UsableInput is what the budget is derived from: the compaction config and the model's limits.
type UsageCache ¶
UsageCache is the cache block of a usage result.
type UsageResult ¶
type UsageResult struct {
Cost float64 `json:"cost"`
Tokens UsageTokens `json:"tokens"`
}
UsageResult is GetUsage's result: the call's cost in USD and its tokens.
func GetUsage ¶
func GetUsage(input GetUsageInput) UsageResult
GetUsage derives the billed token counts and the cost of one model call. Cached input tokens are subtracted from the input count, since providers report inputTokens inclusive of cache reads and writes.
type UsageTokens ¶
type UsageTokens struct {
Total *float64 `json:"total,omitempty"`
Input float64 `json:"input"`
Output float64 `json:"output"`
Reasoning float64 `json:"reasoning"`
Cache UsageCache `json:"cache"`
}
UsageTokens is the token block of a usage result. Total is optional.