advisor

package
v0.1.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 9, 2026 License: MIT Imports: 13 Imported by: 0

README

advisor

Extension. Ejectable — the loop never names it. Without it, the answer the agent produces is the answer the run returns.

A reviewer reads the finished work before the run is allowed to end, and a note that matters sends the agent back to resolve it. Same seam as finishguard; what changes is that the verdict comes from a reviewer instead of a rule. It is the capability finishguard's README lists as its own limitation: "the guard sees only passive evidence — it cannot run a check of its own, call a tool, or consult a model."

The plugin never learns the reviewer is a model. It calls a Reviewer func the host supplies, which is what keeps it testable without a provider and keeps the model, tier, and credential choices where those decisions already live (internal/runtime/advisor.go supplies agentray's).

Ported from oh-my-pi's advisor subsystem.

Model Experience

The reviewer raises a concern or a blocker
What the model sees

The advisories, injected as one synthetic user message, and the run continues instead of returning:

[System: an advisor reviewed the answer you were about to give and raised the
following. It is a second opinion from a reviewer that did NOT do the work and
may be wrong.

<advisory severity="concern" guidance="weigh, don&#39;t blindly obey">
the revenue figure sums a filtered and an unfiltered query
</advisory>

Check each point against what you actually did. Fix what is right; for anything
you judge wrong or already handled, say so in one line with the reason. This
note is not visible to the user and is not a question to answer: your next
message is what the user reads, so reply with the COMPLETE answer they should
see, never a comment about this note. If nothing needs to change, send the
answer again unchanged.]

Two parts of that wording are load-bearing.

guidance="weigh, don't blindly obey" is the agent's only cue for how to treat an advisory — the primary system prompt never mentions them. A tool-less reviewer is wrong often enough that a note which reads like an order gets obeyed when it should have been argued with.

The closing instruction exists because the injection arrives as a user message, so whatever the model says next is what the owner reads. Without it, a model that judges itself already compliant replies with its verdict on the note ("the concern does not apply") and that verdict silently replaces the answer. The owner asked a question and gets back a self-audit about a reply they never saw. Same failure the evidence guard's nudge documents; the repair path must always terminate in the full answer.

The user sees a curated progress note (reviewing the answer with the advisor) — never the raw injection.

Token effect

One reviewer call per round (input: a bounded window of the transcript; output: capped and terse), plus the injection and the extra agent turn it buys. Bounded by MaxRounds (default 2). This is the most expensive plugin in the composition per run, which is why the host is expected to make it opt-in.

KV cache effect

Append-only for the agent. The advisory lands after the rejected answer; nothing earlier moves. The reviewer's own call shares no prefix with the run and is not cached across rounds.

The reviewer raises only nits
What the model sees

Nothing. The finish is accepted and the nits go to OnNotes.

At a finish there is no next step boundary for an aside to ride, so a nit has nowhere to go except a whole extra turn — which costs more than the nit is worth. omp delivers nits as batched asides because its advisor watches a running session; that channel does not exist here.

Token effect

Zero-direct beyond the reviewer call itself.

The reviewer is silent, errors, or the cap is spent
What the model sees

Nothing. The run returns the answer it produced.

Silence is the expected outcome of a run that went fine — the shipped prompt says so in as many words, because a reviewer's default failure is not missing bugs, it is finding something to say.

Impact on the agent

  • Consulted only on a normal finish. A budget wrap-up, a tool-budget stop, a MaxTurns stop, an abort, or a terminal tool never re-opens the run.
  • A failed reviewer accepts the finish. A provider outage, a timeout, or an unparseable response leaves the run exactly as it would have been with no advisor configured. The advisor is a second opinion, never a dependency.
  • The cap is enforced against StopInfo.Attempt, a count the loop keeps — so a reviewer that never runs out of objections is still bounded.
  • The injection is persisted to the durable log like a steer, so a resumed run replays the conversation the model actually saw.
  • The plugin needs no kernel change: StopInfo carries the answer and the tool trace but not the history, so the extension implements RunObserver and keeps the last PhaseRequest snapshot — exactly the context the answer came out of.
  • A PhaseRebase (compaction, a context edit) replaces the history the earlier notes were about, so the dedupe history is dropped with it.
  • Ordering matters. Interceptors are consulted in registration order and the first to say Continue wins. Register the advisor after the goal gate and the finish guard: an unmet goal makes any review of that answer moot, and a rule that costs no tokens should not be rediscovered by a pro-tier call.

The emission guard

Between what the reviewer says and what the agent reads sits emission.go, and it is the part that decides whether the feature is usable.

Reviewer models do not obey the prose rules in their own prompt. oh-my-pi issue #3520 recorded one session with 309 advise calls covering 92 unique notes — 114× "Stop.", 52× "No issue; continue.", 41× "Done." The fix is to make the rules executable:

  1. Clamp each note to MaxNoteRunes (2000). The reviewer read a transcript that may contain attacker-controlled text and its note becomes a user message in the agent's conversation, so an unbounded note is an unbounded injection.
  2. Normalize — lowercase, NFKC, non-alphanumeric runs to one space, trim. "Stop.", "*Stop*", " stop " all key to stop.
  3. Drop content-free phrases — stop, done, lgtm, no issue continue, and similar. Silence is how "no concerns" is expressed. A genuine blocker that merely opens with such a word ("Stop: the revenue query sums a filtered and an unfiltered CTE") does not match, because normalization does not truncate.
  4. Dedupe by normalized text, by escalation rank. A repeat passes only when its severity strictly exceeds what was already delivered for that text, so a reviewer cannot buy a second turn by re-raising the same point — but a real nit → concern → blocker escalation gets through.
  5. Bound breadth at MaxNotesPerReview (3). Suppressed noise never spends the budget: a junk note must not burn the slot for the real concern behind it in the same review.

The gate is deliberately invisible to the reviewer. Telling a model its note was suppressed teaches it to rephrase the same useless note to get past the filter, which defeats the dedupe and costs a round trip to do it.

Rendering escapes the note text, so a quoted </advisory> cannot close the element and let what follows read as the host's own instructions.

Known limitations and deferred work

  • The reviewer has no tools here. It reviews what is in the transcript. In agentray that is usually enough (the SQL and the rows it returned are both there), but a reviewer that could check a claim against the source would catch a class this one cannot. omp grants read/grep/glob; the agentray analog is a nested run with its own scope gating and budget — a second governed agent, not a config change.
  • Delivered is per run, not per session. A resumed run gets a fresh dedupe history, so a note it raised before its crash can be raised again.
  • First interceptor to say Continue wins. If the goal gate or the finish guard re-opens the run that turn, the advisor is not consulted at all.
  • The reviewer sees a window, not the whole run. The host bounds the transcript; the elision is marked so the reviewer knows not to conclude that an unseen step never happened, but it can still miss what fell in the gap.
  • A different tier may mean a different provider. The reviewer reads the run's transcript, so turning it on can send that conversation to a vendor the run itself did not use. That is a workspace configuration question, not something the plugin can decide.

Periodic and multiple reviewers

Plugin.IntervalTurns enables review at existing turn boundaries, with MaxPeriodicReviews (default 8), CooldownTurns after errors (default 3), and Timeout (default 30 seconds). Finish review retains its separate two-round cap. There is no background scheduler: reviews are synchronous, cancellable callbacks before the next primary request. Custom callbacks must respect their context.

Plugin.Reviewers accepts up to four uniquely named ReviewerConfig values, each with its own reviewer, cadence, caps, cooldown and emission guard. Evidence is bounded to 32 text messages / 32 tool traces, and delivered-note history to 32 notes. Cadence, budgets, cooldown and dedupe are checkpointed. A fresh user request after finish review starts a new review allowance. Compaction resets note dedupe but does not replenish review quotas.

NativeWithOptions(provider, options) applies these scheduling controls to the native reviewer; all fallback attempts are charged to the owning agent. Reviewer errors accept the primary result and defer periodic retries through cooldown.

Documentation

Overview

Package advisor reviews completed work and, optionally, work in progress at turn boundaries. Reviewer callbacks contribute bounded, deduplicated notes; material concerns reopen a finish or enter the next request. Each reviewer owns independent review caps, cooldown and checkpointed delivery history. Native bindings account AI usage to the owning run. No background scheduler or in-flight turn interruption is introduced.

Index

Constants

View Source
const DefaultMaxNotesPerReview = 3

DefaultMaxNotesPerReview bounds how many notes one review may deliver. It is breadth control, not rate limiting — a reviewer handed a whole run will happily list nine things, and an injection that long stops being advice and becomes a second task.

View Source
const DefaultMaxRounds = 2

DefaultMaxRounds caps reviewer consultations per run. Two is the same allowance finishguard gives, for the same reason: a reviewer that is never satisfied must not be able to loop the run against MaxTurns.

View Source
const MaxNoteRunes = 2000

MaxNoteRunes clamps one note's length. The reviewer reads a transcript that may contain attacker-controlled text and its note becomes a user message in the primary conversation, so an unbounded note is an unbounded injection.

Variables

This section is empty.

Functions

func FormatAdvisories

func FormatAdvisories(notes []Note) string

FormatAdvisories renders notes as the agent-facing advisory elements: one element per note, severity as an attribute, text XML-escaped.

Escaping is not cosmetic. The reviewer read a transcript that may contain text an attacker controls (a fetched page, a row from the event store), and its note becomes a user message in the primary conversation. Escaping keeps a quoted `</advisory>` inside the note from closing the element and letting whatever follows read as the host's own instructions.

func FormatInjection

func FormatInjection(notes []Note) string

FormatInjection is the complete message injected when a note re-opens the run: the advisories, then what the agent is expected to do about them.

func Native

func Native(provider *ai.FallbackProvider) agentcore.Plugin

Native binds the reviewer to native AI fallback. It has no tools and accounts every attempt's usage to the actual owning agent, including on child forks.

func NativeWithOptions

func NativeWithOptions(provider *ai.FallbackProvider, options Plugin) agentcore.Plugin

NativeWithOptions binds the native reviewer with periodic scheduling and caps.

func NormalizeNote

func NormalizeNote(note string) string

NormalizeNote folds a note to its identity key: lowercase, NFKC-normalized, every run of non-letter/non-digit characters collapsed to a single space, trimmed. "Stop.", "*Stop*", and " stop " all key to "stop", while "No issue; continue." keys to "no issue continue".

Exported because the dedupe key is part of the plugin's observable contract: a caller recording notes on a run needs the same identity the guard used.

Types

type EmissionGuard

type EmissionGuard struct {
	// contains filtered or unexported fields
}

EmissionGuard decides which reviewer notes reach the primary agent.

It enforces, in order: the length clamp, the content-free phrase filter, run-scoped dedupe by normalized text (FIFO-evicted at noteCapacity), escalation-rank dedupe (a repeat passes only when its severity strictly exceeds the rank already delivered for that text), and a per-review breadth budget. Suppressed noise never consumes the budget — a junk note must not burn the slot for a real concern behind it in the same review.

The zero value is not usable; construct with NewEmissionGuard.

func NewEmissionGuard

func NewEmissionGuard(maxPerReview int) *EmissionGuard

NewEmissionGuard builds a guard accepting at most maxPerReview notes per review round. A non-positive maxPerReview means unbounded breadth, which is only ever right for a test.

func (*EmissionGuard) Accept

func (g *EmissionGuard) Accept(n Note) (Note, bool)

Accept reports whether a note should reach the primary, and returns the note as it should be delivered (length-clamped). On true the guard has already recorded it: the budget is spent and the text is in the dedupe history.

func (*EmissionGuard) BeginReview

func (g *EmissionGuard) BeginReview()

BeginReview clears the per-review budget. Called once before each reviewer consultation; the dedupe history deliberately survives, because "you already said that last round" is the whole point of a second round.

func (*EmissionGuard) Reset

func (g *EmissionGuard) Reset()

Reset drops all state. Called when the conversation the notes were about is rewritten (compaction, a rebase), so a re-primed reviewer may re-raise an issue it already raised against a transcript that no longer exists.

type Note

type Note struct {
	// Text is the advice itself: concrete, terse, actionable.
	Text string `json:"text"`
	// Severity decides delivery. An unrecognized or empty value is treated as
	// a nit, so a reviewer that omits the field cannot accidentally interrupt.
	Severity Severity `json:"severity,omitempty"`
}

Note is one piece of advice from the reviewer.

type Plugin

type Plugin struct {
	// Reviewer is consulted at each normal finish.
	Reviewer Reviewer
	// Reviewers adds independently scheduled reviewers (at most four).
	Reviewers []ReviewerConfig
	// IntervalTurns enables periodic review after this many completed turns.
	IntervalTurns      int
	MaxPeriodicReviews int
	CooldownTurns      int
	Timeout            time.Duration
	// MaxRounds bounds consultations per run. 0 uses DefaultMaxRounds. The cap
	// lives here rather than in the loop because it is this capability's
	// property: the loop only knows that SOMETHING asked to continue.
	MaxRounds int
	// MaxNotesPerReview bounds notes delivered per review. 0 uses
	// DefaultMaxNotesPerReview.
	MaxNotesPerReview int
	// OnNotes, when set, receives every note the guard accepted — including
	// the nits the agent never sees — so the host can record what the reviewer
	// said. delivered reports whether these notes were actually put in front of
	// the agent: a review is injected whole or not at all, so a nit riding
	// alongside a blocker IS delivered, and severity alone cannot tell you that.
	OnNotes func(ctx context.Context, notes []Note, delivered bool)
}

Plugin installs the advisor. A nil Reviewer registers nothing, so a composition that wires the plugin for a disabled agent is inert rather than broken — which is what makes "advisor: off" a config value rather than a different composition.

func Of

func Of(r Reviewer) Plugin

Of wraps a reviewer with the default bounds.

func (Plugin) BeginRun

func (p Plugin) BeginRun(ctx context.Context, info agentcore.RunInfo) (agentcore.Extension, error)

BeginRun starts one run's advisor state: its round budget, its emission guard, and the message snapshot the reviewer will read.

func (Plugin) Name

func (Plugin) Name() string

Name identifies the plugin and the extension it installs.

func (Plugin) Register

func (p Plugin) Register(r *agentcore.Registry) error

Register adds the plugin as a run extension. A nil reviewer declines.

type Review

type Review struct {
	// Final is the answer the run would return.
	Final string
	// Turns is the number of reasoning turns consumed, including this one.
	Turns int
	// Tools is the run's tool trace — what was called, what was blocked, what
	// errored. Read-only; it aliases the run's live slice.
	Tools []agentcore.ToolTrace
	// Messages is the conversation the final answer came out of, captured at
	// the last provider request. It does NOT include the final assistant
	// message itself, which is Final.
	Messages []agentcore.Message
	// Round is 0 on the first review of a run and increments per re-opening,
	// so a reviewer can tell "look at this work" from "look at whether the
	// agent dealt with what you already said".
	Round int
	// Delivered is every note this run has already put in front of the agent,
	// oldest first. On a second round it is what the reviewer checks the new
	// answer against; repeating one of these verbatim is dropped by the
	// emission guard, so a reviewer that still objects must escalate.
	Delivered []Note
}

Review is the evidence handed to a Reviewer.

type Reviewer

type Reviewer func(ctx context.Context, r Review) ([]Note, error)

Reviewer reads a finished run and returns the notes worth making. Returning no notes accepts the finish, and silence is the expected outcome of a run that went fine.

An error accepts the finish too. A reviewer is a second opinion, not a dependency: a provider outage, a timeout, or an unparseable response must leave the agent's answer exactly as it would have been with no advisor configured, never wedge or fail the run.

type ReviewerConfig

type ReviewerConfig struct {
	Name               string
	Reviewer           Reviewer
	IntervalTurns      int
	MaxRounds          int
	MaxPeriodicReviews int
	CooldownTurns      int
}

ReviewerConfig names an independent review policy. No goroutines or second scheduler: periodic reviews run at the core's existing turn boundary.

type Severity

type Severity string

Severity is how strongly one note should be weighed. It also decides delivery: a nit is recorded and the finish is accepted, while a concern or a blocker re-opens the run.

const (
	// SeverityNit is cleanup, simplification, a low-risk edge case. It never
	// re-opens the run: at a finish there is no next step boundary for an aside
	// to ride, so spending a whole turn on a nit costs more than the nit is
	// worth. It is reported to the host instead.
	SeverityNit Severity = "nit"
	// SeverityConcern is material risk: a likely-wrong direction, a missing
	// constraint, a figure that does not follow from the evidence. Re-opens the
	// run so the agent can weigh it.
	SeverityConcern Severity = "concern"
	// SeverityBlocker is work that would be wrong to hand over: a claim of
	// completion over sampled scope, an answer never exercised against what was
	// asked. Re-opens the run.
	SeverityBlocker Severity = "blocker"
)

func (Severity) Interrupting

func (s Severity) Interrupting() bool

Interrupting reports whether a note at this severity re-opens the run.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL