calc

package
v0.7.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 5, 2026 License: Apache-2.0 Imports: 6 Imported by: 0

Documentation

Overview

Package calc holds the pure token and cost arithmetic that runs between an OpenRouter response and a persisted assistant message: it normalises the provider usage block, prices a call from the model catalog, and derives the compaction budget and its watermarks from the model limits and the compaction config.

Index

Constants

View Source
const COMPACTION_BUFFER float64 = 20_000

COMPACTION_BUFFER bounds the output reservation taken off the window.

View Source
const COMPACTION_LOW_TO_HIGH_RATIO float64 = 2.0 / 3.0

COMPACTION_LOW_TO_HIGH_RATIO is the low-watermark half of the 40/60 hysteresis: low is two-thirds of high.

View Source
const DefaultCapacityTokens float64 = 500_000

DefaultCapacityTokens caps the working set when no capacity_tokens is configured. A model with a smaller window is bounded by the window.

View Source
const OUTPUT_TOKEN_MAX_DEFAULT float64 = 32_000

OUTPUT_TOKEN_MAX_DEFAULT is the output-token ceiling when SENIOR_DEV_OUTPUT_TOKEN_MAX is unset.

View Source
const PolicyWindow = "window"

PolicyWindow is the compaction policy: the window is the budget.

View Source
const TRIGGER_PCT float64 = 0.6

TRIGGER_PCT is the fraction of the capacity at which auto-compaction fires: the high watermark.

Variables

This section is empty.

Functions

func EffectiveInputCapacity

func EffectiveInputCapacity(input UsableInput) float64

EffectiveInputCapacity is the model-visible input capacity after reserving output space and applying the capacity cap. It deliberately does not apply the trigger percentage; Watermarks derives both marks from this one underlying capacity.

func IsOverflow

func IsOverflow(input OverflowInput) bool

IsOverflow reports whether the assistant's token count has reached the high watermark.

func MaxOutputTokens

func MaxOutputTokens(model Model) float64

MaxOutputTokens is `min(model.limit.output, OUTPUT_TOKEN_MAX)`, falling back to OUTPUT_TOKEN_MAX when the minimum is 0 or NaN. A negative limit.output is passed through as is.

func SetModuleEnvForTesting

func SetModuleEnvForTesting(env map[string]string) func()

SetModuleEnvForTesting re-runs the package-init evaluation of OUTPUT_TOKEN_MAX against the supplied environment. Returns a restore func.

func ValidatePolicy

func ValidatePolicy(cfg Config) error

ValidatePolicy refuses a policy name this binary does not implement, and a budget field outside its range. Both are refused at config load so a misspelled block fails before any model call.

Types

type CacheCost

type CacheCost struct {
	Read  float64 `json:"read"`
	Write float64 `json:"write"`
}

CacheCost is the per-token price of prompt-cache reads and writes.

type CompactionConfig

type CompactionConfig struct {
	// Policy names how the compaction budget is derived. The only policy is
	// "window": the budget is the model's own context window, capped by
	// CapacityTokens. It may be spelled out or left empty; any other name is
	// refused by ValidatePolicy.
	Policy string `json:"policy,omitempty"`
	Auto   *bool  `json:"auto"`
	Prune  *bool  `json:"prune"`
	// PreserveRecentTokens overrides the verbatim tail budget a compaction
	// keeps ahead of the summary. The tail is sized in tokens after
	// truncation, never in turns.
	PreserveRecentTokens *float64 `json:"preserve_recent_tokens"`
	// PreserveRecentFraction sizes the verbatim tail as a fraction of the
	// high watermark instead of a fixed token count, so it scales with the
	// window. PreserveRecentTokens wins when both are set.
	PreserveRecentFraction *float64 `json:"preserve_recent_fraction"`
	// CapacityTokens caps the working set below the model's window: a cost
	// decision, or a model known to degrade before its advertised context.
	// Absent means DefaultCapacityTokens.
	CapacityTokens *float64 `json:"capacity_tokens"`
	Reserved       *float64 `json:"reserved"`
}

CompactionConfig is the `compaction` block of project config. Every field is optional, so every field is a pointer: `auto` is tested strictly (an absent value is NOT false) and `reserved` nullishly (an explicit 0 wins).

type CompactionWatermarks

type CompactionWatermarks struct {
	Capacity float64
	Low      float64
	High     float64
}

CompactionWatermarks describes the preferred post-compaction target and the occupancy at which another compaction becomes necessary.

func Watermarks

func Watermarks(input UsableInput) CompactionWatermarks

Watermarks returns the 40/60 hysteresis around the capacity.

type Config

type Config struct {
	Compaction *CompactionConfig `json:"compaction"`
}

Config is the project-config projection this package needs: only the `compaction` block is read.

type GetUsageInput

type GetUsageInput struct {
	Model    Model
	Usage    LanguageModelUsage
	Metadata ProviderMetadata
}

GetUsageInput is GetUsage's parameter object.

type InputTokenDetails

type InputTokenDetails struct {
	NoCacheTokens    *float64 `json:"noCacheTokens,omitempty"`
	CacheReadTokens  *float64 `json:"cacheReadTokens,omitempty"`
	CacheWriteTokens *float64 `json:"cacheWriteTokens,omitempty"`
}

InputTokenDetails is the flattened input breakdown.

type LanguageModelUsage

type LanguageModelUsage struct {
	InputTokens        *float64            `json:"inputTokens,omitempty"`
	InputTokenDetails  *InputTokenDetails  `json:"inputTokenDetails,omitempty"`
	OutputTokens       *float64            `json:"outputTokens,omitempty"`
	OutputTokenDetails *OutputTokenDetails `json:"outputTokenDetails,omitempty"`
	TotalTokens        *float64            `json:"totalTokens,omitempty"`
	Raw                json.RawMessage     `json:"raw,omitempty"`
	ReasoningTokens    *float64            `json:"reasoningTokens,omitempty"`
	CachedInputTokens  *float64            `json:"cachedInputTokens,omitempty"`
}

LanguageModelUsage is the flattened usage GetUsage consumes. The details blocks are optional; the two trailing fields are flat aliases GetUsage falls back to when the details are absent.

func AsLanguageModelUsage

func AsLanguageModelUsage(usage LanguageModelV3Usage) LanguageModelUsage

AsLanguageModelUsage flattens the token groups. totalTokens is recomputed as input + output; the provider's own total survives only inside Raw.

type LanguageModelV3InputTokens

type LanguageModelV3InputTokens struct {
	Total      *float64 `json:"total,omitempty"`
	NoCache    *float64 `json:"noCache,omitempty"`
	CacheRead  *float64 `json:"cacheRead,omitempty"`
	CacheWrite *float64 `json:"cacheWrite,omitempty"`
}

LanguageModelV3InputTokens is the provider-side `inputTokens` block.

type LanguageModelV3OutputTokens

type LanguageModelV3OutputTokens struct {
	Total     *float64 `json:"total,omitempty"`
	Text      *float64 `json:"text,omitempty"`
	Reasoning *float64 `json:"reasoning,omitempty"`
}

LanguageModelV3OutputTokens is the provider-side `outputTokens` block.

type LanguageModelV3Usage

type LanguageModelV3Usage struct {
	InputTokens  LanguageModelV3InputTokens  `json:"inputTokens"`
	OutputTokens LanguageModelV3OutputTokens `json:"outputTokens"`
	Raw          json.RawMessage             `json:"raw,omitempty"`
}

LanguageModelV3Usage is ComputeTokenUsage's return shape.

func ComputeTokenUsage

func ComputeTokenUsage(usage *OpenRouterUsage) LanguageModelV3Usage

ComputeTokenUsage splits the provider's usage block into input and output token groups. An absent cache-write count stays absent (nil) so that the provider-metadata fallbacks in GetUsage can still supply it.

type Model

type Model struct {
	Cost         *ModelCost        `json:"cost"`
	Limit        ModelLimit        `json:"limit"`
	Capabilities ModelCapabilities `json:"-"`
}

Model is the catalog projection this package needs: the limit block (for the budget and the output reservation) and the cost block (for usage).

type ModelCapabilities

type ModelCapabilities struct {
	Attachment  bool            `json:"attachment"`
	Reasoning   bool            `json:"reasoning"`
	Temperature bool            `json:"temperature"`
	ToolCall    bool            `json:"toolcall"`
	Input       map[string]bool `json:"input"`
	Output      map[string]bool `json:"output"`
}

ModelCapabilities is the models.dev capability slice retained alongside cost and limits so provider request assembly does not invent support.

type ModelCost

type ModelCost struct {
	Input                float64       `json:"input"`
	Output               float64       `json:"output"`
	Cache                *CacheCost    `json:"cache"`
	ExperimentalOver200K *Over200KCost `json:"experimentalOver200K"`
}

ModelCost is a model's price block. `cache` is optional in practice, so it is a pointer.

type ModelLimit

type ModelLimit struct {
	Context float64  `json:"context"`
	Input   *float64 `json:"input"`
	Output  float64  `json:"output"`
}

ModelLimit is a model's context, input and output limits. Input is optional: a catalog entry that names none is budgeted from Context.

type OpenRouterCompletionTokensDetails

type OpenRouterCompletionTokensDetails struct {
	ReasoningTokens *float64 `json:"reasoning_tokens"`
}

OpenRouterCompletionTokensDetails is `usage.completion_tokens_details`.

type OpenRouterPromptTokensDetails

type OpenRouterPromptTokensDetails struct {
	CachedTokens     *float64 `json:"cached_tokens"`
	CacheWriteTokens *float64 `json:"cache_write_tokens"`
}

OpenRouterPromptTokensDetails is `usage.prompt_tokens_details`.

type OpenRouterUsage

type OpenRouterUsage struct {
	PromptTokens            *float64                           `json:"prompt_tokens"`
	CompletionTokens        *float64                           `json:"completion_tokens"`
	PromptTokensDetails     *OpenRouterPromptTokensDetails     `json:"prompt_tokens_details"`
	CompletionTokensDetails *OpenRouterCompletionTokensDetails `json:"completion_tokens_details"`

	// Raw is the untouched wire object, including fields such as `cost` and
	// `is_byok` that senior-dev never reads.
	Raw json.RawMessage `json:"-"`
}

OpenRouterUsage is the numeric projection of OpenRouter's `usage` object that the token arithmetic reads, plus the untouched original carried through as Raw.

func (*OpenRouterUsage) UnmarshalJSON

func (u *OpenRouterUsage) UnmarshalJSON(data []byte) error

UnmarshalJSON decodes the numeric projection and keeps the original bytes.

type OutputTokenDetails

type OutputTokenDetails struct {
	TextTokens      *float64 `json:"textTokens,omitempty"`
	ReasoningTokens *float64 `json:"reasoningTokens,omitempty"`
}

OutputTokenDetails is the flattened output breakdown.

type Over200KCost

type Over200KCost struct {
	Cache  *CacheCost `json:"cache"`
	Input  float64    `json:"input"`
	Output float64    `json:"output"`
}

Over200KCost is the price block a provider applies above 200K context. The key order is cache, input, output.

type OverflowInput

type OverflowInput struct {
	Cfg    Config
	Tokens Tokens
	Model  Model
}

OverflowInput is what the trigger decides on.

type ProviderMetadata

type ProviderMetadata map[string]map[string]any

ProviderMetadata is the per-provider metadata map a response may carry.

type TokenCache

type TokenCache struct {
	Read  float64 `json:"read"`
	Write float64 `json:"write"`
}

TokenCache is the persisted cache-token pair, declared read then write. Contrast UsageCache, which is the same data in the order usage builds it.

type Tokens

type Tokens struct {
	Total     *float64   `json:"total,omitempty"`
	Input     float64    `json:"input"`
	Output    float64    `json:"output"`
	Reasoning float64    `json:"reasoning"`
	Cache     TokenCache `json:"cache"`
}

Tokens is the persisted assistant token block. `total` is optional, so it is a pointer and the key is dropped when it is absent.

type UsableInput

type UsableInput struct {
	Cfg   Config
	Model Model
}

UsableInput is what the budget is derived from: the compaction config and the model's limits.

type UsageCache

type UsageCache struct {
	Write float64 `json:"write"`
	Read  float64 `json:"read"`
}

UsageCache is the cache block of a usage result.

type UsageResult

type UsageResult struct {
	Cost   float64     `json:"cost"`
	Tokens UsageTokens `json:"tokens"`
}

UsageResult is GetUsage's result: the call's cost in USD and its tokens.

func GetUsage

func GetUsage(input GetUsageInput) UsageResult

GetUsage derives the billed token counts and the cost of one model call. Cached input tokens are subtracted from the input count, since providers report inputTokens inclusive of cache reads and writes.

type UsageTokens

type UsageTokens struct {
	Total     *float64   `json:"total,omitempty"`
	Input     float64    `json:"input"`
	Output    float64    `json:"output"`
	Reasoning float64    `json:"reasoning"`
	Cache     UsageCache `json:"cache"`
}

UsageTokens is the token block of a usage result. Total is optional.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL