Documentation
¶
Overview ¶
Package modelrouter implements the pure decision engine of model auto mode: it turns (config, prompt, short history) into a routing Decision using a System One decision provider. It holds no per-session state.
Index ¶
- Constants
- func BuildState(in Input, historyPrompts int, budgetTokens int) string
- func ClassifyError(err error) string
- func DecisionConsumers(cfg *config.Config) []string
- func DefaultLookup(cfg *config.Config) func(models.ModelID) (CandidateInfo, bool)
- func EstimateTokens(s string) int
- func FilterCandidates(cands []models.ModelID, lookup func(models.ModelID) (CandidateInfo, bool), ...) (usable []models.ModelID, skipped map[models.ModelID]string)
- func MaskedAPIKey(r config.DecisionRouterConfig) string
- func NewRelevanceFilter(opts RelevanceOptions) rag.RelevanceFilter
- func PersonaCap(n int) (offered, dropped int)
- func ProviderFor(dec config.DecisionModelConfig) (systemone.DecisionProvider, error)
- func ProviderWithTimeout(r config.DecisionRouterConfig, timeout time.Duration) (systemone.DecisionProvider, error)
- func RenderHealth(rep systemone.HealthReport) string
- func RenderModels(ctx context.Context, showAll bool) (string, error)
- func RenderStatus(ctx context.Context, withHealth bool) (string, error)
- func RouterHealth(ctx context.Context, dec config.DecisionModelConfig) (systemone.HealthReport, error)
- func SafetyMargin(budgetTokens int) int
- func SharedHealth() *systemone.HealthCache
- func StartWarmupOnReload()
- type CandidateInfo
- type Decision
- type Engine
- func (e *Engine) Ask(ctx context.Context, state string, qs map[string]systemone.Question) (*systemone.Response, time.Duration, error)
- func (e *Engine) ContextBudget(ctx context.Context) int
- func (e *Engine) Provider() systemone.DecisionProvider
- func (e *Engine) Route(ctx context.Context, policy config.ModelAutoModeConfig, in Input) Decision
- func (e *Engine) RoutePersona(ctx context.Context, in PersonaInput) PersonaDecision
- func (e *Engine) RouteWithPersona(ctx context.Context, policy config.ModelAutoModeConfig, in Input, ...) (Decision, PersonaDecision)
- type EngineOption
- type Input
- type PersonaDecision
- type PersonaInput
- type PersonaOption
- type RelevanceOptions
Constants ¶
const ( SkipUnknown = "unknown model" SkipDisabled = "provider disabled or unconfigured" SkipNoAttachments = "no attachment support" SkipContextWindow = "context window too small" SkipDuplicate = "duplicate" )
Skip reasons reported by FilterCandidates.
const ( ReasonMatched = "matched" ReasonNoMatch = "no_match" ReasonLowProbability = "low_probability" ReasonLowConfidence = "low_confidence" ReasonRouterError = "router_error" ReasonNoRoutes = "no_routes" ReasonDisabled = "disabled" )
Decision reasons.
const ( ErrClassUnreachable = "unreachable" ErrClassModelNotFound = "model_not_found" ErrClassTimeout = "timeout" ErrClassTooLarge = "too_large" ErrClassBadRequest = "bad_request" ErrClassMalformed = "malformed" ErrClassServer = "server" )
Error classes carried in Decision.ErrClass.
const ( ReasonNoPersonas = "no_personas" ReasonNoRouter = "no_router" )
Persona decision reasons not shared with model routing.
const ( RelevanceReasonNoRouter = "no_router" RelevanceReasonHosted = "hosted_provider" RelevanceReasonNoCands = "no_candidates" RelevanceReasonPartialPfx = "partial:" )
Relevance filter reasons (rag.FilterResult.Reason) besides the error classes.
const HealthTTL = 60 * time.Second
HealthTTL is the lifetime of a cached router health report.
const PersonaThreshold = 0.60
PersonaThreshold is the minimum probability of the chosen persona.
Variables ¶
This section is empty.
Functions ¶
func BuildState ¶
BuildState renders the System One state for in as a JSON object string {"request", "previous_requests"?, "attachments"?}. budgetTokens is the token budget available for the state (the caller has already subtracted the question and criteria overhead); a safety margin is reserved on top of that so the estimated size of the result is at most budgetTokens minus SafetyMargin(budgetTokens). The prompt keeps its head and tail joined by an ellipsis marker; only the last historyPrompts previous prompts are included, each truncated harder; attachments are listed by file name only.
func ClassifyError ¶
ClassifyError maps a decision-model failure to one of the ErrClass* values.
func DecisionConsumers ¶
DecisionConsumers lists the features that currently use the decision model.
func DefaultLookup ¶
DefaultLookup returns a lookup backed by the models registry and the provider enabled/disabled state in cfg. A nil cfg treats every known model as enabled.
func EstimateTokens ¶
EstimateTokens is the cheap token estimator used by the router.
func FilterCandidates ¶
func FilterCandidates(cands []models.ModelID, lookup func(models.ModelID) (CandidateInfo, bool), hasAttachments bool, historyTokens int) (usable []models.ModelID, skipped map[models.ModelID]string)
FilterCandidates removes duplicates, unknown models, models whose provider is disabled or unconfigured, models without attachment support when the prompt carries attachments, and models whose context window cannot hold historyTokens. The order of cands is preserved. skipped maps each removed model to the reason. A ContextWindow <= 0 means "unknown" and is never skipped.
func MaskedAPIKey ¶
func MaskedAPIKey(r config.DecisionRouterConfig) string
MaskedAPIKey returns a display-safe form of the router API key: empty when none is stored, "$VAR" references verbatim (masked), the masked tail otherwise. The unmasked key is never returned.
func NewRelevanceFilter ¶
func NewRelevanceFilter(opts RelevanceOptions) rag.RelevanceFilter
NewRelevanceFilter returns a rag.RelevanceFilter driven by the shared decision model.
Rule for candidates that are not judged (resolves "max candidates"): only the first MaxCandidates non-pinned candidates, taken in the deterministic order code -> kb -> events -> memory, are asked. Candidates beyond the cap, and candidates that do not fit any request (token budget, MaxBodyBytes), are KEPT: the filter never drops what the model has not seen. Kept candidates stay subject to the existing *MaxChars/TotalMaxChars budgets. Pinned candidates are never asked and always kept.
The filter is fail-open: any error, timeout, malformed answer, empty router model or engine build failure keeps everything with Applied=false and the error class in Reason. When only some batches fail, the failed batch is kept, the others still apply, Applied=true and Reason is "partial:<class>".
func PersonaCap ¶
PersonaCap reports how many of n personas are offered to the decision model and how many are dropped by the 25 cap.
func ProviderFor ¶
func ProviderFor(dec config.DecisionModelConfig) (systemone.DecisionProvider, error)
ProviderFor builds the decision provider described by the decision model, bounded by dec.EffectiveTimeout().
func ProviderWithTimeout ¶
func ProviderWithTimeout(r config.DecisionRouterConfig, timeout time.Duration) (systemone.DecisionProvider, error)
ProviderWithTimeout is ProviderFor with an explicit per-call bound (warm-up needs a longer one: the first call loads the model).
func RenderHealth ¶
func RenderHealth(rep systemone.HealthReport) string
RenderHealth formats a health verdict as one line plus the problems.
func RenderModels ¶
RenderModels lists the decision models the provider reports.
func RenderStatus ¶
RenderStatus returns a text report of the saved decision model: provider, effective URL, masked key, model, timeout, consumers and, when withHealth is set, a live health probe. The API key is always masked.
func RouterHealth ¶
func RouterHealth(ctx context.Context, dec config.DecisionModelConfig) (systemone.HealthReport, error)
RouterHealth returns the (cached) health report for the decision model.
func SafetyMargin ¶
SafetyMargin returns the number of tokens kept free below a budget for the model's own answer and estimator error.
func SharedHealth ¶
func SharedHealth() *systemone.HealthCache
SharedHealth returns the process-wide router health cache.
func StartWarmupOnReload ¶
func StartWarmupOnReload()
StartWarmupOnReload subscribes to the config bus (once per process) and, on every change of a section in warmupSections, drops the cached health reports and warms the decision model up in the background when it is configured and at least one consumer is enabled.
Types ¶
type CandidateInfo ¶
type CandidateInfo struct {
ID models.ModelID
Known bool
Enabled bool
SupportsAttachments bool
ContextWindow int
}
CandidateInfo describes a candidate model for filtering.
type Decision ¶
type Decision struct {
RouteID string // "" when no match
Matched bool
Probability float64
Confidence float64
Probabilities map[string]float64
// Candidates is [primary, fb1, fb2] of the route, or [coder] on no match or failure.
Candidates []models.ModelID
Reason string
Err error
ErrClass string
RouterProvider string
RouterModel string
LatencyMs int64
CostUSD *float64
InputTokens int
}
Decision is the outcome of routing one prompt.
type Engine ¶
type Engine struct {
// contains filtered or unexported fields
}
Engine routes prompts. It is safe for concurrent use.
The engine is keyed only on the shared decision model (provider, model, timeout); the per-consumer policy (threshold, routes, history) is passed to each call.
func ForConfig ¶
func ForConfig(dec config.DecisionModelConfig) (*Engine, error)
ForConfig returns the process-wide engine for dec, rebuilding it when the decision model content changes (hot reload).
func NewEngine ¶
func NewEngine(dec config.DecisionModelConfig, opts ...EngineOption) (*Engine, error)
NewEngine builds an engine from the shared decision model (provider, model and timeout).
func (*Engine) Ask ¶
func (e *Engine) Ask(ctx context.Context, state string, qs map[string]systemone.Question) (*systemone.Response, time.Duration, error)
Ask is the generic entry point for any consumer: it sends one System One request with the given JSON state (a JSON object string) and questions to the configured decision model. Answers are NOT validated (the lenient decode is used) so the caller judges each one. The caller is responsible for keeping state within ContextBudget (BuildState does that for the routing state). More than systemone.MaxQuestions questions, or a state/request over systemone.MaxBodyBytes, is rejected before anything is sent, with an error that ClassifyError maps to "bad_request" / "too_large". A non-nil error is a transport/protocol failure that affects every question.
func (*Engine) ContextBudget ¶
ContextBudget returns the context window, in tokens, of the configured decision model (cached for a few minutes, with a conservative default when the provider cannot tell). Callers subtract their question overhead and keep SafetyMargin free.
func (*Engine) Provider ¶
func (e *Engine) Provider() systemone.DecisionProvider
Provider returns the decision provider used by the engine.
func (*Engine) Route ¶
Route classifies in.Prompt against the routes and thresholds of policy. It never fails: on any router failure it returns the coder model with Reason "router_error".
func (*Engine) RoutePersona ¶
func (e *Engine) RoutePersona(ctx context.Context, in PersonaInput) PersonaDecision
RoutePersona classifies in.Prompt into one of in.Personas. It never fails: any router problem yields Reason ReasonRouterError (or ReasonNoRouter when no router model is configured) with Err/ErrClass set.
func (*Engine) RouteWithPersona ¶
func (e *Engine) RouteWithPersona(ctx context.Context, policy config.ModelAutoModeConfig, in Input, p PersonaInput) (Decision, PersonaDecision)
RouteWithPersona takes the model-route decision and the persona decision from ONE /v1/systemone request carrying two questions ("task" and "persona"). The state comes from in; p.Prompt and p.History are ignored in that case. The request's CostUSD and InputTokens are attributed to the model Decision only and left nil/0 on the PersonaDecision so the cost is not counted twice; LatencyMs is reported on both. When either question has nothing to ask (no enabled routes, no personas, no router model) the decisions are taken separately, which sends at most one question.
type EngineOption ¶
type EngineOption func(*Engine)
EngineOption customises NewEngine.
func WithProvider ¶
func WithProvider(p systemone.DecisionProvider) EngineOption
WithProvider injects a decision provider (tests).
type Input ¶
type Input struct {
Prompt string
History []string // previous user prompts, newest last
AttachmentNames []string
HasAttachments bool
CoderModel models.ModelID
}
Input is what the engine needs to route one prompt.
type PersonaDecision ¶
type PersonaDecision struct {
Persona string // "" when not matched
Matched bool
Probability float64
Confidence float64
// Probabilities is keyed by persona NAME plus the reserved "none". A persona
// literally named "none" is reported under its criterion key (e.g. "none_2")
// so it never overwrites the real "none" entry.
Probabilities map[string]float64
Reason string
Err error
ErrClass string
RouterProvider string
RouterModel string
LatencyMs int64
CostUSD *float64 // 0/nil when the request was shared with the model decision (see RouteWithPersona)
InputTokens int
Offered int // personas actually sent
Dropped []string // personas left out by the 25 cap
}
PersonaDecision is the outcome of classifying one prompt into a persona.
type PersonaInput ¶
type PersonaInput struct {
Prompt string
History []string // previous user prompts, newest last (may be empty)
Personas []PersonaOption
// HistoryPrompts caps how many History entries reach the state (the
// auto mode policy value; 0 sends none).
HistoryPrompts int
}
PersonaInput is what the engine needs to pick a persona for one prompt.
type PersonaOption ¶
type PersonaOption struct {
Name string // persona name as known to the persona manager
Description string // when this persona applies (<= 500 chars)
}
PersonaOption is one persona the router may pick.
type RelevanceOptions ¶
type RelevanceOptions struct {
// Decision returns the shared decision model. Default: config.Get().DecisionModel.
Decision func() config.DecisionModelConfig
// Threshold is the minimum p(useful) to keep a candidate. Default: remembrances config.
Threshold func() float64
// MaxCandidates caps how many candidates are judged per call. Default: remembrances config.
MaxCandidates func() int
// MaxCandidateChars caps the text of one candidate. Default: remembrances config.
MaxCandidateChars func() int
// LocalOnly skips hosted (typesafe/custom) providers. Default: remembrances config.
LocalOnly func() bool
// HistoryPrompts returns previous user prompts (oldest first) added to the state.
// Default: none.
HistoryPrompts func() []string
// HistoryCount is how many previous prompts to include. Default: config
// ModelAutoMode.HistoryPrompts.
HistoryCount func() int
}
RelevanceOptions configures NewRelevanceFilter. Every func is optional and is evaluated on each Filter call, so hot reload needs no rebuild; a nil func reads the current global config (config.Get()).