router

package
v0.2.1-rc.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 17, 2026 License: Apache-2.0 Imports: 24 Imported by: 0

Documentation

Overview

Package router turns one model into a panel of them.

It exists because of a measured result rather than an intuition. Two labs (bench/routerlab, bench/probelab) asked whether a harness can spend less and get more by not sending every call to the same model, and the answer they converged on is narrow and specific: predicting which model will fail does not pay, but *trying the cheap one and escalating on a verified failure* does. On a 36-task suite that policy reached 0.974 success at $0.000148 a task against the incumbent single model's 0.875 at $0.000076 — and it beat the best single model on the panel outright, at a sixth of that model's price.

Everything here follows from that one sentence. The cascade needs a verifier, so verification is the centre of the design and an unverifiable call is treated as evidence about nothing. The order of the rungs needs an ability estimate, so there is a ledger — a one-parameter Rasch rating per model and call class, because the labs tested the two-parameter alternative and it did not pay for itself. And the ability estimate may not be seeded from price: measured against ability, log output price correlated at r = 0.46 with n = 7, and the second-cheapest model on the panel had the second-highest ability while the second-dearest sat sixth of seven.

The whole package is inert unless CODEAF_MODELS names a panel. With it unset the harness builds the same single adapter it always did, and nothing below this line runs.

Index

Constants

View Source
const MinGraded = 8

MinGraded is how much graded evidence a rating needs before the ordering may prefer it to the cold-start prior.

Eight, which is profile.MinSamples — the number this codebase already uses for exactly this decision, "may a measurement overwrite a default". Deliberately the same number rather than a new one, and it is comfortably above what went wrong: arm B's collapse was **five** graded observations, all of them budget stops on a single oversized task, outvoting a prior and rerouting every leaf on the panel. The report's own conclusion is that this gate alone would have prevented the whole thing.

The gate is on *graded* observations specifically, and that is the point. 103 of 108 leaf outcomes in that arm were unverified successes, which move nothing and are correct to move nothing — so a count that looked like plenty of experience was five failures wearing a hundred and eight's clothes. Below the gate the ordering falls back to the cold-start prior, which is the honest statement that nothing is known yet.

Variables

This section is empty.

Functions

func Ability

func Ability(rating float64) float64

Ability is the rating read as a success probability against an average call of its class, which is the form the cascade actually orders on.

func Shaped

func Shaped(class provider.CallClass, shape string) provider.CallClass

Shaped is the ledger key for a class that is divided into sub-populations. The class stays readable — `exec.leaf/oversized` — because a ledger nobody can read is a ledger nobody checks.

It is EXPORTED because the router is no longer the only thing that writes into the ledger: the chat engine grades a settled task node under provider.ClassTaskNode with the work's own kind as its shape (internal/session's taskgrade.go). Two spellings of "how a shape becomes a key" would be two keys for one population, which is the one fault a ledger cannot recover from — so there is one function and both callers use it.

func WithAvoidModel

func WithAvoidModel(ctx context.Context, model string) context.Context

WithAvoidModel asks a cascade to prefer a different opener when the panel permits it. The named model stays available as a fallback, so a preference never turns a healthy single-model configuration into a failure.

Types

type Catalog

type Catalog struct {
	Fetched time.Time               `json:"fetched"`
	Models  map[string]CatalogEntry `json:"models"`
}

Catalog is what the provider says about the panel: prices, context windows, and the parameters each model claims to accept.

The claims are recorded and not believed. The probe lab asked all three of its models for logprobs, all three advertised support in `supported_parameters`, and one of them returned null on every request. So this file is used for economics — price is a real number and the cascade's ordering needs it — and never as a capability check. What a model can actually do is measured.

func LoadCatalog

func LoadCatalog(dir, baseURL, apiKey string, client *http.Client) *Catalog

LoadCatalog returns the panel's economics, preferring a fetch and falling back to whatever was cached.

The order matters and so does the tolerance. A stale price makes the cascade order slightly wrong; a hard failure here makes the harness refuse to start because a metadata endpoint was slow. So every failure is soft: fetch, fall back to the cache however old it is, and fall back again to knowing nothing, which the router handles by ordering the panel as the operator wrote it.

func (*Catalog) Entry

func (c *Catalog) Entry(slug string) (CatalogEntry, bool)

Entry looks a model up, tolerating OpenRouter's floating-alias prefix: the catalog lists `deepseek/deepseek-v4-flash-latest` while the configured slug is `~deepseek/deepseek-v4-flash-latest`, and a lookup that missed on the tilde would silently price the model at zero.

type CatalogEntry

type CatalogEntry struct {
	Slug          string   `json:"slug"`
	PromptPrice   float64  `json:"prompt_price"` // $/M input tokens
	OutputPrice   float64  `json:"output_price"` // $/M output tokens
	ContextLength int      `json:"context_length,omitempty"`
	Parameters    []string `json:"supported_parameters,omitempty"`
}

CatalogEntry is one model's advertised economics.

type Client

type Client interface {
	CompleteWithMessages(ctx context.Context, messages []ai.Message, options ...ai.Option) (*ai.Response, error)
	Model() string
}

Client is the slice of the model adapter everything above the transport consumes. Both the single-model adapter and the Router satisfy it, which is what lets a panel be a configuration choice instead of a second code path through the planner and the executor.

type Entry

type Entry struct {
	Model   string             `json:"model"`
	Class   provider.CallClass `json:"class"`
	Rating  float64            `json:"rating"`
	Count   int                `json:"count"`
	Updated time.Time          `json:"updated"`
}

Entry is one model's measured ability at one kind of call.

Keyed by class because ability is not one number. The router lab's panel ordered differently on structured planning than on reasoning, and a single pooled rating would have learned the average of two things it could have known separately. Keyed by resolved snapshot rather than by configured slug because a floating alias is a moving target: `~vendor/model-latest` is a different set of weights this month than last, and pooling them would let a regression hide behind the record of the model it replaced.

type Event

type Event struct {
	At    time.Time `json:"at"`
	Call  string    `json:"call"` // ties every row about one unit of work together
	Run   string    `json:"run"`  // the run cache key, which is a run identity
	Class string    `json:"class"`
	// Shape is the sub-population within the class the rating was keyed on —
	// which leaves this leaf was ordered against. Empty for an undivided class.
	Shape string `json:"shape,omitempty"`

	// Candidates is the ordered rung list the choice was made from, and Rung is
	// where in it Model sits. Escalation is what had already been tried and
	// failed before this attempt.
	Candidates []string `json:"candidates,omitempty"`
	Rung       int      `json:"rung"`
	Escalation []string `json:"escalation,omitempty"`

	Model string `json:"model"`
	// Resolved is the response's own model field. For a floating alias it is a
	// dated snapshot and it is what the ledger keys on, because a rating pooled
	// across two sets of weights measures neither.
	Resolved string `json:"resolved,omitempty"`

	// Explore marks an attempt the router took to buy evidence rather than
	// because it expected the best answer, and Propensity is the probability it
	// would have. The pair is what makes the log usable for offline policy work:
	// evidence gathered by a rule that chose it for its own reasons is biased in
	// the direction of the rule, and only a recorded propensity lets a later
	// estimator weight it back out. It is one number and it is written down at
	// the moment it applied, which is the only moment it can be known.
	Explore    bool    `json:"explore,omitempty"`
	Propensity float64 `json:"propensity,omitempty"`

	Verdict provider.Reading `json:"verdict"`
	// Final marks the row that carries the settled verdict. An attempt is
	// written when it returns, before the call site has had a chance to check
	// the answer; when the check lands it is appended as a second, short row
	// against the same call id. The log is append-only, so a correction is an
	// append — a reader takes the last row for a call id as the truth.
	Final bool `json:"final,omitempty"`

	// Outcome is the settle's own plain word for how a whole piece of work
	// ended — `landed`, `not accepted`, `stopped` — and it is empty on every
	// row about a single model call. A verdict says what may be LEARNED from an
	// outcome; this says what the outcome WAS, and the two stopped being the
	// same question when the chat engine started grading settled task nodes:
	// a node somebody stopped and a node nobody could check are both
	// unlearnable and are not the same news (internal/session's taskgrade.go).
	Outcome string `json:"outcome,omitempty"`
	// Retries is how many times the work was handed back before it settled —
	// the repair rounds a task node spent. Zero on a routed call, which retries
	// by escalating rather than by repeating, and where Escalation already says
	// what was tried.
	Retries int `json:"retries,omitempty"`

	PromptTokens     int     `json:"prompt_tokens,omitempty"`
	CompletionTokens int     `json:"completion_tokens,omitempty"`
	CachedTokens     int     `json:"cached_tokens,omitempty"`
	Cost             float64 `json:"cost,omitempty"`
	LatencyMS        int64   `json:"latency_ms,omitempty"`
}

Event is one routed attempt, written as it happened.

This file is not telemetry for a dashboard; it is the substrate for the offline policy work the two labs could only simulate. Both of them had to reconstruct what a router *would* have done from an offline matrix, and the thing neither could recover was the counterfactual: which models were in the running when a choice was made. So the candidates are recorded alongside the choice, which is the one field that makes a logged decision analysable after the fact rather than merely auditable.

type Events

type Events struct {
	// contains filtered or unexported fields
}

Events is the append-only record of every routing decision.

One row per turn is the design — a leaf may take two hundred of them — so the rows are buffered and land in batches. What is never batched is a *row*: see Append for why half a row on disk would be worse than no row at all.

func OpenEvents

func OpenEvents(dir string) (*Events, error)

OpenEvents opens the log. A log that cannot be opened is not a reason to refuse a run: the router still routes, it just stops keeping a diary, so the error is returned for reporting and a nil Events is usable.

func (*Events) Append

func (e *Events) Append(event Event)

Append writes one row. Rows are written whole and O_APPEND is atomic for a write of this size on every platform this runs on, so two concurrent codeaf processes interleave rows without ever interleaving bytes.

That is the one thing the buffering must not break, and it is why the buffer is landed before a row that would not fit rather than after: a row split across two writes is a row another process is free to land inside, and half a JSON object is worse in the log than no object. Every write the buffer makes therefore carries whole rows only — including a row larger than the buffer, which goes straight to the file in one call.

Every read of the file handle happens under the mutex, including the one that asks whether there is a file at all. A leaf's settled verdict is appended from whichever goroutine reported it, which may be after the run has begun shutting down, so Close and Append genuinely do race — the check outside the lock was a data race on the handle that only the race detector would ever have shown, since a closed *os.File returns an error rather than panicking.

func (*Events) Close

func (e *Events) Close() error

Close releases the log.

type Ledger

type Ledger struct {
	// contains filtered or unexported fields
}

Ledger is what the harness has learned about its panel, across runs and across processes.

Two locks, and the split is the point. mutex guards the maps and is held for map operations only — reads take it shared, because Rating, Resolve, Entries and Aliases only look. flushing serialises the file work: a lock file with retries, a re-read, a re-parse, a re-encode and a rename, which is on the order of a second and used to happen with the ranking mutex held. Routing a single call reads the ledger once per rung, so every one of those reads was queued behind whatever observation happened to be writing to disk.

func LoadLedger

func LoadLedger(dir string) (*Ledger, error)

LoadLedger reads what has been learned, returning an empty ledger when there is nothing yet. A corrupt file is treated as an empty one, exactly as the profile does: it is an accumulation of observations, not a source of truth, and refusing to run over it would trade a small loss for a total one.

func (*Ledger) Alias

func (l *Ledger) Alias(slug, resolved string)

Alias records which dated snapshot a floating slug actually served. The resolved model is the response's own `model` field, so this is the provider telling us what it ran rather than us guessing.

func (*Ledger) Aliases

func (l *Ledger) Aliases() map[string]string

Aliases returns the floating-slug to snapshot map, for reporting.

func (*Ledger) Close

func (l *Ledger) Close() error

Close saves and then stops waiting for anything. An observation still trying to reach the file after this abandons its lock wait rather than holding the process open for it.

func (*Ledger) Entries

func (l *Ledger) Entries() []Entry

Entries returns everything known, ordered for reading: strongest first within a class, classes alphabetically.

func (*Ledger) Observe

func (l *Ledger) Observe(model string, class provider.CallClass, prior float64, verdict provider.Reading)

Observe folds one outcome in, and only the outcomes that are evidence: a provider failure says nothing about a model and an unverified success says nothing about an answer.

func (*Ledger) Rating

func (l *Ledger) Rating(model string, class provider.CallClass, prior float64) (float64, int)

Rating reports what is known about a model at one kind of call. The prior is used only when nothing is known, which is what makes a new model start somewhere sensible instead of at the bottom.

func (*Ledger) Read

func (l *Ledger) Read(class provider.CallClass, queries []Query) []Reading

Read rates a whole panel in one acquisition.

Ranking used to take the lock twice per rung — resolve the alias, then read the rating — which is twenty acquisitions for a five-model panel on every ordering, and an ordering happens several times per routed call. Each of them could land behind a flush. One snapshot per call is both cheaper and more honest: every rung in an ordering is now read from the same instant.

func (*Ledger) Resolve

func (l *Ledger) Resolve(slug string) string

Resolve maps a configured slug onto the snapshot last seen behind it. Ratings are read through it so that selection and recording agree about which model they are talking about.

func (*Ledger) Save

func (l *Ledger) Save() error

Save flushes anything still queued, now rather than at the end of the window. It is called on the way out; the ordinary path flushes as it goes, so this is only ever picking up a debounced burst or a file lock that was busy at the time.

type Panel

type Panel struct {
	Models []Spec `json:"models"`

	// ClientConfig resolves each model's own account. It is runtime wiring and
	// never part of the JSON an operator writes.
	ClientConfig func(model string) provider.Config `json:"-"`

	// MaxOutputPrice refuses a model dearer than this, in $/M output tokens. It
	// is a sanity cap rather than a budget: a typo in a slug that resolves to a
	// frontier model would otherwise be discovered on the invoice. It can only
	// act on a price it knows, so it does not fire for a model the catalog has
	// never described and the operator did not price — which is the offline case,
	// where refusing to run would be the worse failure.
	MaxOutputPrice float64 `json:"max_output_price,omitempty"`
}

Panel is the set of models one run may use, in the order the operator wrote them. That order is the tie-break of last resort: when nothing has been measured and no price is known, the panel is tried as written.

func LoadPanel

func LoadPanel(value string) (Panel, error)

LoadPanel reads CODEAF_MODELS. An empty value means no panel, which is not an error — it is the kill switch, and it is the default.

Two forms are accepted. A comma-separated list of slugs is the one that gets typed, and it is deliberately the whole of what a first-time user needs to know. A path to a JSON file is for a panel worth keeping, where roles and a price cap are worth writing down. YAML was specified and is not implemented: this module has no YAML dependency and adding one to read a six-line file would be the most expensive line in the go.mod.

type Query

type Query struct {
	Slug  string
	Prior float64
}

Query is one member to read: the configured slug, and the cold-start prior to answer with when nothing is known about it.

type Reading

type Reading struct {
	Rating float64
	Count  int
}

Reading is one panel member's standing at one kind of call: the rating and the graded evidence behind it, read through the alias map.

type Router

type Router struct {
	// contains filtered or unexported fields
}

Router sends each call to the cheapest model that can do it, and escalates when it turns out one could not.

It is the same shape as the adapter it wraps, which is the whole design: the planner and the executor call CompleteWithMessages and never learn that there is more than one model behind it.

func New

func New(panel Panel, base provider.Config, dir string) (*Router, error)

New builds a router over a panel. base carries everything the adapter needs except the model, which each rung supplies for itself.

func NewPinned

func NewPinned(panel Panel, base provider.Config, dir, opener string) (*Router, error)

NewPinned builds a router whose first attempt honours an explicit model choice. The choice joins the panel when it is not already present, so a chat picker may name any model without giving up ledger observation or the panel's terminal escalation rung.

func (*Router) Close

func (r *Router) Close() error

Close flushes what the run learned and releases the diary.

A Router owns a file handle and a queue of observations the run has already paid for, so a router that is replaced — a model switched mid-session — has to be closed rather than dropped. Close is idempotent: a second call saves nothing new and closes nothing twice.

func (*Router) CompleteWithMessages

func (r *Router) CompleteWithMessages(ctx context.Context, messages []ai.Message, options ...ai.Option) (*ai.Response, error)

CompleteWithMessages routes one call.

The two shapes of work are routed differently and the difference is not an optimisation. A planning call is one request whose answer can be checked the moment it arrives, so it cascades: cheap model, verify, escalate on a failure the verifier caught. A leaf is a whole conversation whose worth is only known at the end, so it is routed once and pinned — escalation for a leaf is the scheduler re-running it, not the loop changing model between turns.

func (*Router) Ledger

func (r *Router) Ledger() *Ledger

Ledger exposes what has been learned, for reporting.

func (*Router) Model

func (r *Router) Model() string

Model reports the configured opener. It is what the harness prints and what the profile is keyed on; which model actually served a given call is in the events log, where a per-call answer belongs.

func (*Router) Rungs

func (r *Router) Rungs() int

Rungs is how many models the panel holds. The scheduler reads it to decide whether re-running a failed leaf could possibly help.

func (*Router) StreamComplete

func (r *Router) StreamComplete(ctx context.Context, prompt string, options ...ai.Option) (<-chan ai.StreamChunk, <-chan error)

StreamComplete streams from the first rung and never cascades. A stream is committed the moment its first byte is delivered, so there is nothing to escalate to: the answer has already started arriving. Nothing in the harness streams today; this exists so that the router is a drop-in for the adapter.

type Spec

type Spec struct {
	Slug  string  `json:"slug"`
	Role  string  `json:"role,omitempty"`  // base, mid or top
	Price float64 `json:"price,omitempty"` // $/M output tokens, overriding the catalog
}

Spec is one model on the panel.

Role is an operator's claim about where a model sits, not a measurement, and it is used for exactly one thing: the cold-start rating of a model nothing has been recorded about yet. It is worth about one observation and is updated away immediately. Price is an override for the catalog, for the case where an operator knows something the catalog does not — a negotiated rate, a private deployment — and not a place to encode an opinion about quality.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL