Documentation
¶
Overview ¶
Package goal installs the run-level completion contract: a goal-gated run may only stop when it says so.
The condition is stated in the system prompt up front, and the model must end its final answer with STATUS: DONE once the goal is met, or STATUS: BLOCKED plus the reason when only the user can unblock it. A finish with neither sentinel re-opens the run with a keep-going nudge.
Unlike the finishguard plugin (a capped verify nudge), the gate is UNCAPPED by design: it holds the line until the model declares done or blocked. It cannot wedge a run — MaxTurns, MaxToolCalls, and the budget gate still bound the loop, a budget wrap-up never consults a stop interceptor at all, and the two stall breakers below stop a model that has nothing new to say.
Stalling, and why one breaker is not enough ¶
Being uncapped is only affordable if "the model has nothing left to give" is detected cheaply. The original breaker compared the finish to the previous one verbatim, which is a good proxy for a strong model — asked the same question against a near-identical context, it tends to emit the same tokens — and a useless one for a weaker model, which paraphrases instead. Measured: a model that answers every nudge with the same claim in different words trips the verbatim breaker never, and burns the entire turn budget re-asserting that it is finished. So the second breaker asks the behavioral question rather than the lexical one: did the model DO anything between one nudge and the next? A run of finishes with no tool call in between is a stall whatever the wording, and unlike text comparison it does not care how the model phrases itself.
The nudge escalates for the same reason. A model that ignored an instruction is unlikely to obey a byte-identical repeat of it, so the second and later nudges state the contract far more mechanically than the first.
The goal STRING is durable state the loop owns: it writes the EntryGoal record and recovers it on resume, then hands it to this plugin as RunInfo.Goal. So a crashed run comes back gated even when the resuming caller could not re-supply the condition, without this package ever writing to the log.
Composition order ¶
The loop consults stop interceptors in registration order and the first one to continue wins, so this plugin must be registered BEFORE finishguard: an unmet goal makes any verify pass on that answer moot.
Index ¶
Constants ¶
const ( Done = "STATUS: DONE" Blocked = "STATUS: BLOCKED" )
Done and Blocked are the completion sentinels the gate looks for in the final answer's closing line (case-insensitive, so "Status: done" also counts).
const StopReasonStalled = "goal_stalled"
StopReasonStalled is recorded when a stall breaker accepts a finish because the model has stopped making progress toward the goal — it either repeated its previous answer verbatim, or kept declaring itself finished without doing any work between one nudge and the next.
const ToolName = "update_goal"
ToolName is the model-facing name of the goal-revision tool.
Variables ¶
This section is empty.
Functions ¶
func PublicText ¶
PublicText removes a standalone closing protocol line from customer prose. A status mentioned within an explanation is retained verbatim.
func StripStatusLine ¶
StripStatusLine also works on a buffered stream tail: it preserves every preceding byte, so hosts can hide the completion protocol without joining words or dropping whitespace between preview chunks.
func UntilRevisable ¶
UntilRevisable gates the run on a condition the AGENT may restate, with the reason recorded, when the work shows that condition to be wrong. The returned Store is the audit trail: Revisions() holds every condition the run has held, in order, each with the justification the model gave for changing it.
Types ¶
type Command ¶
type Command struct {
Action string `json:"action"`
Reason string `json:"reason,omitempty"`
TokenBudget *int `json:"token_budget,omitempty"`
TimeBudgetMS *int64 `json:"time_budget_ms,omitempty"`
}
Command is a host-authorized lifecycle operation. Budgets are absolute totals, not additional grants. Resume never resets usage or bypasses an exhausted cap. The model can inspect state; pause/resume/drop belong to the caller.
type Plugin ¶
type Plugin struct {
Goal string
// Lifecycle enables checkpointed pause/resume, accounting and host controls.
Lifecycle bool
NewTaskOnInput bool
TokenBudget int
TimeBudget time.Duration
// Revisable offers the model the update_goal tool, letting it restate what
// finishing means when the work shows the condition to be wrong.
//
// Off by default, and that default is the honest one: it hands the model the
// key to its own gate, so a composition takes it deliberately rather than
// inheriting it. See updateGoalTool for what the trade actually is.
Revisable bool
// Store, when set, is the condition's live home — share it to read the
// revision trail after the run. Left nil, the plugin builds its own per run.
Store *Store
}
Plugin installs the goal gate. An empty Goal leaves the run ungated, so a composition can wire the plugin unconditionally and let configuration decide.
func (Plugin) BeginRun ¶
BeginRun arms the gate for this run. It declines an ungated run — the run has no condition to hold it to, so there is nothing to consult and no contract to put in the prompt.
The condition comes from RunInfo, not from p.Goal: on a resume the loop recovers it from the durable log, and a plugin that trusted its own field would come back ungated exactly when gating matters most.
func (Plugin) Register ¶
Register claims the goal seam (the loop persists the condition) and adds the gate itself as a run extension. Both halves are one call because a recorded goal nobody enforces, or an enforced goal nobody recorded, is a bug either way — the first resumes ungated, the second gates a run whose condition is invisible to the log.
type Revision ¶
Revision is one change to the run's objective: what it became, why, and when. The reason is the part a human reads later — "the goal changed" is not a finding, "the goal changed because CLEARING-7742 turned out to span three corpora" is.
type State ¶
type State struct {
Objective string `json:"objective"`
Status Status `json:"status"`
TokenBudget int `json:"token_budget,omitempty"`
TokensUsed int `json:"tokens_used"`
TimeBudgetMS int64 `json:"time_budget_ms,omitempty"`
ElapsedMS int64 `json:"elapsed_ms"`
Reason string `json:"reason,omitempty"`
}
State is retained with the transcript, never reconstructed from a summary. ElapsedMS counts active execution only; time spent paused/offline is excluded.
type Store ¶
type Store struct {
// contains filtered or unexported fields
}
Store holds the run's live completion condition and the record of how it got there.
It exists for the same reason todo.Store does: the state has to outlive the transcript. A condition that lived only in the system message would be rewritten by compaction like anything else, and one that lived only in the gate's own field could not be read by the tool that changes it. So the condition is held here, the gate reads it, the tool writes it, and the loop drains it into the durable log.
A nil *Store is not valid; the plugin always builds one.
func (*Store) ApplyCommand ¶
ApplyCommand lets a host pause a live run at its next settled boundary. It is safe to call concurrently; resume across processes uses NativeRun.Commands.
func (*Store) Update ¶
Update replaces the condition and marks it for the loop to persist. A revision to the identical condition is dropped: it is not a change, and letting it through would write an EntryGoal and rebuild the system prompt — invalidating the provider's KV cache for the whole prefix — in exchange for nothing.