modelrouter

package
v1.2.8 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 2, 2026 License: MIT Imports: 16 Imported by: 0

Documentation

Overview

Package modelrouter implements the pure decision engine of model auto mode: it turns (config, prompt, short history) into a routing Decision using a System One decision provider. It holds no per-session state.

Index

Constants

View Source
const (
	SkipUnknown       = "unknown model"
	SkipDisabled      = "provider disabled or unconfigured"
	SkipNoAttachments = "no attachment support"
	SkipContextWindow = "context window too small"
	SkipDuplicate     = "duplicate"
)

Skip reasons reported by FilterCandidates.

View Source
const (
	ReasonMatched        = "matched"
	ReasonNoMatch        = "no_match"
	ReasonLowProbability = "low_probability"
	ReasonLowConfidence  = "low_confidence"
	ReasonRouterError    = "router_error"
	ReasonNoRoutes       = "no_routes"
	ReasonDisabled       = "disabled"
)

Decision reasons.

View Source
const (
	ErrClassUnreachable   = "unreachable"
	ErrClassUnauthorized  = "unauthorized"
	ErrClassModelNotFound = "model_not_found"
	ErrClassTimeout       = "timeout"
	ErrClassTooLarge      = "too_large"
	ErrClassBadRequest    = "bad_request"
	ErrClassMalformed     = "malformed"
	ErrClassServer        = "server"
)

Error classes carried in Decision.ErrClass.

View Source
const (
	ReasonNoPersonas = "no_personas"
	ReasonNoRouter   = "no_router"
)

Persona decision reasons not shared with model routing.

View Source
const (
	RelevanceReasonNoRouter   = "no_router"
	RelevanceReasonHosted     = "hosted_provider"
	RelevanceReasonNoCands    = "no_candidates"
	RelevanceReasonPartialPfx = "partial:"
)

Relevance filter reasons (rag.FilterResult.Reason) besides the error classes.

View Source
const HealthTTL = 60 * time.Second

HealthTTL is the lifetime of a cached router health report.

View Source
const PersonaThreshold = 0.60

PersonaThreshold is the minimum probability of the chosen persona.

Variables

This section is empty.

Functions

func BuildState

func BuildState(in Input, historyPrompts int, budgetTokens int) string

BuildState renders the System One state for in as a JSON object string {"request", "previous_requests"?, "attachments"?}. budgetTokens is the token budget available for the state (the caller has already subtracted the question and criteria overhead); a safety margin is reserved on top of that so the estimated size of the result is at most budgetTokens minus SafetyMargin(budgetTokens). The prompt keeps its head and tail joined by an ellipsis marker; only the last historyPrompts previous prompts are included, each truncated harder; attachments are listed by file name only.

func ClassifyError

func ClassifyError(err error) string

ClassifyError maps a decision-model failure to one of the ErrClass* values.

func DecisionConsumers

func DecisionConsumers(cfg *config.Config) []string

DecisionConsumers lists the features that currently use the decision model.

func DefaultLookup

func DefaultLookup(cfg *config.Config) func(models.ModelID) (CandidateInfo, bool)

DefaultLookup returns a lookup backed by the models registry and the provider enabled/disabled state in cfg. A nil cfg treats every known model as enabled.

func EstimateTokens

func EstimateTokens(s string) int

EstimateTokens is the cheap token estimator used by the router.

func FilterCandidates

func FilterCandidates(cands []models.ModelID, lookup func(models.ModelID) (CandidateInfo, bool), hasAttachments bool, historyTokens int) (usable []models.ModelID, skipped map[models.ModelID]string)

FilterCandidates removes duplicates, unknown models, models whose provider is disabled or unconfigured, models without attachment support when the prompt carries attachments, and models whose context window cannot hold historyTokens. The order of cands is preserved. skipped maps each removed model to the reason. A ContextWindow <= 0 means "unknown" and is never skipped.

func MaskedAPIKey

func MaskedAPIKey(r config.DecisionRouterConfig) string

MaskedAPIKey returns a display-safe form of the router API key: empty when none is stored, "$VAR" references verbatim (masked), the masked tail otherwise. The unmasked key is never returned.

func NewRelevanceFilter

func NewRelevanceFilter(opts RelevanceOptions) rag.RelevanceFilter

NewRelevanceFilter returns a rag.RelevanceFilter driven by the shared decision model.

Rule for candidates that are not judged (resolves "max candidates"): only the first MaxCandidates non-pinned candidates, taken in the deterministic order code -> kb -> events -> memory, are asked. Candidates beyond the cap, and candidates that do not fit any request (token budget, MaxBodyBytes), are KEPT: the filter never drops what the model has not seen. Kept candidates stay subject to the existing *MaxChars/TotalMaxChars budgets. Pinned candidates are never asked and always kept.

The filter is fail-open: any error, timeout, malformed answer, empty router model or engine build failure keeps everything with Applied=false and the error class in Reason. When only some batches fail, the failed batch is kept, the others still apply, Applied=true and Reason is "partial:<class>".

func PersonaCap

func PersonaCap(n int) (offered, dropped int)

PersonaCap reports how many of n personas are offered to the decision model and how many are dropped by the 25 cap.

func ProviderFor

ProviderFor builds the decision provider described by the decision model, bounded by dec.EffectiveTimeout().

func ProviderWithTimeout

func ProviderWithTimeout(r config.DecisionRouterConfig, timeout time.Duration) (systemone.DecisionProvider, error)

ProviderWithTimeout is ProviderFor with an explicit per-call bound (warm-up needs a longer one: the first call loads the model).

func RenderHealth

func RenderHealth(rep systemone.HealthReport) string

RenderHealth formats a health verdict as one line plus the problems.

func RenderModels

func RenderModels(ctx context.Context, showAll bool) (string, error)

RenderModels lists the decision models the provider reports.

func RenderStatus

func RenderStatus(ctx context.Context, withHealth bool) (string, error)

RenderStatus returns a text report of the saved decision model: provider, effective URL, masked key, model, timeout, consumers and, when withHealth is set, a live health probe. The API key is always masked.

func RouterHealth

RouterHealth returns the (cached) health report for the decision model.

func SafetyMargin

func SafetyMargin(budgetTokens int) int

SafetyMargin returns the number of tokens kept free below a budget for the model's own answer and estimator error.

func SharedHealth

func SharedHealth() *systemone.HealthCache

SharedHealth returns the process-wide router health cache.

func StartWarmupOnReload

func StartWarmupOnReload()

StartWarmupOnReload subscribes to the config bus (once per process) and, on every change of a section in warmupSections, drops the cached health reports and warms the decision model up in the background when it is configured and at least one consumer is enabled.

Types

type CandidateInfo

type CandidateInfo struct {
	ID                  models.ModelID
	Known               bool
	Enabled             bool
	SupportsAttachments bool
	ContextWindow       int
}

CandidateInfo describes a candidate model for filtering.

type Decision

type Decision struct {
	RouteID       string // "" when no match
	Matched       bool
	Probability   float64
	Confidence    float64
	Probabilities map[string]float64
	// Candidates is [primary, fb1, fb2] of the route, or [coder] on no match or failure.
	Candidates     []models.ModelID
	Reason         string
	Err            error
	ErrClass       string
	RouterProvider string
	RouterModel    string
	LatencyMs      int64
	CostUSD        *float64
	InputTokens    int
}

Decision is the outcome of routing one prompt.

type Engine

type Engine struct {
	// contains filtered or unexported fields
}

Engine routes prompts. It is safe for concurrent use.

The engine is keyed only on the shared decision model (provider, model, timeout); the per-consumer policy (threshold, routes, history) is passed to each call.

func ForConfig

func ForConfig(dec config.DecisionModelConfig) (*Engine, error)

ForConfig returns the process-wide engine for dec, rebuilding it when the decision model content changes (hot reload).

func NewEngine

func NewEngine(dec config.DecisionModelConfig, opts ...EngineOption) (*Engine, error)

NewEngine builds an engine from the shared decision model (provider, model and timeout).

func (*Engine) Ask

Ask is the generic entry point for any consumer: it sends one System One request with the given JSON state (a JSON object string) and questions to the configured decision model. Answers are NOT validated (the lenient decode is used) so the caller judges each one. The caller is responsible for keeping state within ContextBudget (BuildState does that for the routing state). More than systemone.MaxQuestions questions, or a state/request over systemone.MaxBodyBytes, is rejected before anything is sent, with an error that ClassifyError maps to "bad_request" / "too_large". A non-nil error is a transport/protocol failure that affects every question.

func (*Engine) ContextBudget

func (e *Engine) ContextBudget(ctx context.Context) int

ContextBudget returns the context window, in tokens, of the configured decision model (cached for a few minutes, with a conservative default when the provider cannot tell). Callers subtract their question overhead and keep SafetyMargin free.

func (*Engine) Provider

func (e *Engine) Provider() systemone.DecisionProvider

Provider returns the decision provider used by the engine.

func (*Engine) Route

func (e *Engine) Route(ctx context.Context, policy config.ModelAutoModeConfig, in Input) Decision

Route classifies in.Prompt against the routes and thresholds of policy. It never fails: on any router failure it returns the coder model with Reason "router_error".

func (*Engine) RoutePersona

func (e *Engine) RoutePersona(ctx context.Context, in PersonaInput) PersonaDecision

RoutePersona classifies in.Prompt into one of in.Personas. It never fails: any router problem yields Reason ReasonRouterError (or ReasonNoRouter when no router model is configured) with Err/ErrClass set.

func (*Engine) RouteWithPersona

func (e *Engine) RouteWithPersona(ctx context.Context, policy config.ModelAutoModeConfig, in Input, p PersonaInput) (Decision, PersonaDecision)

RouteWithPersona takes the model-route decision and the persona decision from ONE /v1/systemone request carrying two questions ("task" and "persona"). The state comes from in; p.Prompt and p.History are ignored in that case. The request's CostUSD and InputTokens are attributed to the model Decision only and left nil/0 on the PersonaDecision so the cost is not counted twice; LatencyMs is reported on both. When either question has nothing to ask (no enabled routes, no personas, no router model) the decisions are taken separately, which sends at most one question.

type EngineOption

type EngineOption func(*Engine)

EngineOption customises NewEngine.

func WithProvider

func WithProvider(p systemone.DecisionProvider) EngineOption

WithProvider injects a decision provider (tests).

type Input

type Input struct {
	Prompt          string
	History         []string // previous user prompts, newest last
	AttachmentNames []string
	HasAttachments  bool
	CoderModel      models.ModelID
}

Input is what the engine needs to route one prompt.

type PersonaDecision

type PersonaDecision struct {
	Persona     string // "" when not matched
	Matched     bool
	Probability float64
	Confidence  float64
	// Probabilities is keyed by persona NAME plus the reserved "none". A persona
	// literally named "none" is reported under its criterion key (e.g. "none_2")
	// so it never overwrites the real "none" entry.
	Probabilities  map[string]float64
	Reason         string
	Err            error
	ErrClass       string
	RouterProvider string
	RouterModel    string
	LatencyMs      int64
	CostUSD        *float64 // 0/nil when the request was shared with the model decision (see RouteWithPersona)
	InputTokens    int
	Offered        int      // personas actually sent
	Dropped        []string // personas left out by the 25 cap
}

PersonaDecision is the outcome of classifying one prompt into a persona.

type PersonaInput

type PersonaInput struct {
	Prompt   string
	History  []string // previous user prompts, newest last (may be empty)
	Personas []PersonaOption
	// HistoryPrompts caps how many History entries reach the state (the
	// auto mode policy value; 0 sends none).
	HistoryPrompts int
}

PersonaInput is what the engine needs to pick a persona for one prompt.

type PersonaOption

type PersonaOption struct {
	Name        string // persona name as known to the persona manager
	Description string // when this persona applies (<= 500 chars)
}

PersonaOption is one persona the router may pick.

type RelevanceOptions

type RelevanceOptions struct {
	// Decision returns the shared decision model. Default: config.Get().DecisionModel.
	Decision func() config.DecisionModelConfig
	// Threshold is the minimum p(useful) to keep a candidate. Default: remembrances config.
	Threshold func() float64
	// MaxCandidates caps how many candidates are judged per call. Default: remembrances config.
	MaxCandidates func() int
	// MaxCandidateChars caps the text of one candidate. Default: remembrances config.
	MaxCandidateChars func() int
	// LocalOnly skips hosted (typesafe/custom) providers. Default: remembrances config.
	LocalOnly func() bool
	// HistoryPrompts returns previous user prompts (oldest first) added to the state.
	// Default: none.
	HistoryPrompts func() []string
	// HistoryCount is how many previous prompts to include. Default: config
	// ModelAutoMode.HistoryPrompts.
	HistoryCount func() int
}

RelevanceOptions configures NewRelevanceFilter. Every func is optional and is evaluated on each Filter call, so hot reload needs no rebuild; a nil func reads the current global config (config.Get()).

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL