goal

package
v0.1.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 8, 2026 License: MIT Imports: 7 Imported by: 0

README

goal

Seam + extension. The gate mechanism lives here, in this package: the completion contract, the sentinel match, the nudge, and the stall breaker are all in goal.go, reaching the loop through PromptContributor and StopInterceptor. Delete this folder and a run stops whenever the model likes.

Core owns the checkpoint's current goal and revision trail. RunNative restores these before constructing the plugin and supplies the current condition through RunInfo.Goal. A tool commits a revision through the native run owner before changing its store; durable session hosts supply their journal recorder at that same boundary. The plugin owns the completion policy, not persistence.

Claude Code's /goal analog.

Composition

preset.New(agentcore.Config{Goal: "all tests pass", /* … */})   // wired for you
agentcore.Build(/* … */, goal.Until("all tests pass"))          // à la carte

plugin, trail := goal.UntilRevisable("all tests pass")          // + the update_goal tool
agentcore.Build(/* … */, plugin)
trail.Revisions()                                               // the audit trail, oldest first

agentcore.New(Config) alone records the goal but does not enforce it — core cannot import this package. Use preset.New, or pass goal.Until(…) in Config.Extensions.

The plugin is installed even when no goal is configured, and declines the run at BeginRun. That is not defensive coding: a resumed run gets its condition from the log, so an empty Goal field is not proof the run is ungated.

Model Experience

A goal is set
What the model sees

The completion contract appended to the system prompt.

Token effect

Fixed and retained — it is in the prefix of every provider call.

KV cache effect

Prefix-stable.

The model finishes without a sentinel
What the model sees

A keep-going nudge as a synthetic user message, and the run continues.

Token effect

Fixed per nudge. Uncapped in count — but the run is still bounded by turn/tool/budget limits.

KV cache effect

Append-only.

The model revises the goal (UntilRevisable only)
What the model sees

An update_goal(goal, reason) tool. Calling it replaces the completion condition the gate enforces, from the next turn on.

Token effect

Fixed and retained — one schema in the prefix of every provider call. This is why it is opt-in: an ungated run, or a run whose objective is not in question, should not pay it.

KV cache effect

Prefix-invalidating, once per accepted revision. The condition is in the system prompt, so changing it rewrites the whole cached prefix. That is the real price of a revision, and it is why an identical restatement is rejected rather than written through.

Impact on the agent

  • The run may only stop by ending its answer with STATUS: DONE (goal met) or STATUS: BLOCKED (+ reason). The sentinel is matched on the closing line only, so mentioning it mid-prose does not count.

  • The gate must be registered before other stop interceptors: the first to re-open the run wins, and an unmet goal makes any verification pass on that same answer moot. preset.Plugins and internal/runtime both encode this.

  • It never wedges a run: a budget wrap-up bypasses it, and two stall breakers stop it as goal_stalled rather than burning turns to MaxTurns — a verbatim-repeated answer (immediately), or three consecutive finishes with no tool call in between. The second exists because the first only works on a model deterministic enough to repeat itself; a weaker model paraphrases its way past a text comparison, so progress is measured by what the model did, not by how it worded itself. A model that goes back to work between nudges resets the count and is never capped.

  • The nudge escalates. The first explains the contract; the second and later ones dictate the literal line to emit and rule out a trailing sign-off after it — the most common way a weaker model misses a contract it is trying to follow. Repeating identical prose to a model that already ignored it is the one approach known not to work.

  • The native checkpoint retains the goal, so a resumed run is still gated even when the caller cannot re-supply the condition.

  • UntilRevisable transfers real authority, and should be read as one. The gate exists to stop a model ending a run it has not finished; a tool that rewrites the gate's condition can end any run by redefining success. It is here because the alternative is worse in the case it is for: a long autonomous run discovers things, and a requirement resting on an assumption the work disproved leaves the agent only two honest moves — grind to MaxTurns, or declare BLOCKED and hand back nothing.

    What keeps it accountable is the record, not a restriction. Every revision requires a reason and is recorded with its previous condition before the store changes. It stays in the checkpoint and the store's trail — so "the agent narrowed its goal until it could pass" is something you can see afterwards, in order, with the model's own justification on each step. The pinned user requirement is not touched: a revision changes what finishing means, never what was asked for.

    GoalReviser publishes a committed condition between turns. The native host verifies it matches the recorded condition before rebuilding the prompt. Restoring a checkpoint does not append another revision.

Known limitations and deferred work

  • The sentinel is prose. A model that writes STATUS: DONE without having done anything satisfies the gate; this is a protocol, not a verifier. Pair it with a finishguard for evidence.
  • One goal per run. No sub-goals, no partial completion. UntilRevisable lets the single goal change; it does not let a run hold two.
  • Nothing reviews a revision while the run is live. The trail is written for a human reading afterwards, and the tool's description is the only thing arguing against a self-serving narrowing. A composition that wants a revision approved before it takes effect has to build that; the seam for it is GoalReviser, which the loop drains and could gate.
  • Ordering is a convention, not a constraint. Nothing rejects a composition that registers a verify guard ahead of the gate; it would merely verify answers the gate is about to throw away.

Structured lifecycle (opt-in)

Plugin{Lifecycle: true, TokenBudget: ..., TimeBudget: ...} retains active, paused, budget_limited, complete or dropped state in the native checkpoint. Accounting includes provider, cache, child and secondary-model tokens, plus active wall time; paused/offline time is excluded. Budgets are checked at settled turn boundaries, so one in-flight turn can overshoot a cap.

The host sends goal.Command JSON in NativeRun.Commands["goal"]: get, pause, resume, drop, or create (a new task after a terminal goal). ControlOnly: true applies controls without requesting the model; the result is in RunResult.CommandResults. Resume retains usage; exhausted absolute caps must be raised via token_budget / time_budget_ms. Store.ApplyCommand also supports concurrent host pause of a live run. Current tool effects settle before pause takes effect; pause does not roll them back.

get_goal is the model-facing inspection tool. Lifecycle authority stays with the host; model revision still requires Revisable. Only accepted final work with STATUS: DONE marks complete. NewTaskOnInput lets conversation hosts start fresh accounting when new input follows a complete/dropped task; it never implicitly resumes paused/budget-limited work. Forks get independent stores.

Documentation

Overview

Package goal installs the run-level completion contract: a goal-gated run may only stop when it says so.

The condition is stated in the system prompt up front, and the model must end its final answer with STATUS: DONE once the goal is met, or STATUS: BLOCKED plus the reason when only the user can unblock it. A finish with neither sentinel re-opens the run with a keep-going nudge.

Unlike the finishguard plugin (a capped verify nudge), the gate is UNCAPPED by design: it holds the line until the model declares done or blocked. It cannot wedge a run — MaxTurns, MaxToolCalls, and the budget gate still bound the loop, a budget wrap-up never consults a stop interceptor at all, and the two stall breakers below stop a model that has nothing new to say.

Stalling, and why one breaker is not enough

Being uncapped is only affordable if "the model has nothing left to give" is detected cheaply. The original breaker compared the finish to the previous one verbatim, which is a good proxy for a strong model — asked the same question against a near-identical context, it tends to emit the same tokens — and a useless one for a weaker model, which paraphrases instead. Measured: a model that answers every nudge with the same claim in different words trips the verbatim breaker never, and burns the entire turn budget re-asserting that it is finished. So the second breaker asks the behavioral question rather than the lexical one: did the model DO anything between one nudge and the next? A run of finishes with no tool call in between is a stall whatever the wording, and unlike text comparison it does not care how the model phrases itself.

The nudge escalates for the same reason. A model that ignored an instruction is unlikely to obey a byte-identical repeat of it, so the second and later nudges state the contract far more mechanically than the first.

The goal STRING is durable state the loop owns: it writes the EntryGoal record and recovers it on resume, then hands it to this plugin as RunInfo.Goal. So a crashed run comes back gated even when the resuming caller could not re-supply the condition, without this package ever writing to the log.

Composition order

The loop consults stop interceptors in registration order and the first one to continue wins, so this plugin must be registered BEFORE finishguard: an unmet goal makes any verify pass on that answer moot.

Index

Constants

View Source
const (
	Done    = "STATUS: DONE"
	Blocked = "STATUS: BLOCKED"
)

Done and Blocked are the completion sentinels the gate looks for in the final answer's closing line (case-insensitive, so "Status: done" also counts).

View Source
const StopReasonStalled = "goal_stalled"

StopReasonStalled is recorded when a stall breaker accepts a finish because the model has stopped making progress toward the goal — it either repeated its previous answer verbatim, or kept declaring itself finished without doing any work between one nudge and the next.

View Source
const ToolName = "update_goal"

ToolName is the model-facing name of the goal-revision tool.

Variables

This section is empty.

Functions

func PublicText

func PublicText(text string) string

PublicText removes a standalone closing protocol line from customer prose. A status mentioned within an explanation is retained verbatim.

func StripStatusLine

func StripStatusLine(text string) string

StripStatusLine also works on a buffered stream tail: it preserves every preceding byte, so hosts can hide the completion protocol without joining words or dropping whitespace between preview chunks.

func UntilRevisable

func UntilRevisable(condition string) (Plugin, *Store)

UntilRevisable gates the run on a condition the AGENT may restate, with the reason recorded, when the work shows that condition to be wrong. The returned Store is the audit trail: Revisions() holds every condition the run has held, in order, each with the justification the model gave for changing it.

Types

type Command

type Command struct {
	Action       string `json:"action"`
	Reason       string `json:"reason,omitempty"`
	TokenBudget  *int   `json:"token_budget,omitempty"`
	TimeBudgetMS *int64 `json:"time_budget_ms,omitempty"`
}

Command is a host-authorized lifecycle operation. Budgets are absolute totals, not additional grants. Resume never resets usage or bypasses an exhausted cap. The model can inspect state; pause/resume/drop belong to the caller.

type Plugin

type Plugin struct {
	Goal string
	// Lifecycle enables checkpointed pause/resume, accounting and host controls.
	Lifecycle      bool
	NewTaskOnInput bool
	TokenBudget    int
	TimeBudget     time.Duration
	// Revisable offers the model the update_goal tool, letting it restate what
	// finishing means when the work shows the condition to be wrong.
	//
	// Off by default, and that default is the honest one: it hands the model the
	// key to its own gate, so a composition takes it deliberately rather than
	// inheriting it. See updateGoalTool for what the trade actually is.
	Revisable bool
	// Store, when set, is the condition's live home — share it to read the
	// revision trail after the run. Left nil, the plugin builds its own per run.
	Store *Store
}

Plugin installs the goal gate. An empty Goal leaves the run ungated, so a composition can wire the plugin unconditionally and let configuration decide.

func Until

func Until(condition string) Plugin

Until gates the run on a condition.

func (Plugin) BeginRun

BeginRun arms the gate for this run. It declines an ungated run — the run has no condition to hold it to, so there is nothing to consult and no contract to put in the prompt.

The condition comes from RunInfo, not from p.Goal: on a resume the loop recovers it from the durable log, and a plugin that trusted its own field would come back ungated exactly when gating matters most.

func (Plugin) Name

func (Plugin) Name() string

Name identifies the plugin and the extension it installs.

func (Plugin) Register

func (p Plugin) Register(r *agentcore.Registry) error

Register claims the goal seam (the loop persists the condition) and adds the gate itself as a run extension. Both halves are one call because a recorded goal nobody enforces, or an enforced goal nobody recorded, is a bug either way — the first resumes ungated, the second gates a run whose condition is invisible to the log.

type Revision

type Revision struct {
	Goal   string
	Reason string
	At     time.Time
}

Revision is one change to the run's objective: what it became, why, and when. The reason is the part a human reads later — "the goal changed" is not a finding, "the goal changed because CLEARING-7742 turned out to span three corpora" is.

type State

type State struct {
	Objective    string `json:"objective"`
	Status       Status `json:"status"`
	TokenBudget  int    `json:"token_budget,omitempty"`
	TokensUsed   int    `json:"tokens_used"`
	TimeBudgetMS int64  `json:"time_budget_ms,omitempty"`
	ElapsedMS    int64  `json:"elapsed_ms"`
	Reason       string `json:"reason,omitempty"`
}

State is retained with the transcript, never reconstructed from a summary. ElapsedMS counts active execution only; time spent paused/offline is excluded.

type Status

type Status string
const (
	Active        Status = "active"
	Paused        Status = "paused"
	BudgetLimited Status = "budget_limited"
	Complete      Status = "complete"
	Dropped       Status = "dropped"
)

type Store

type Store struct {
	// contains filtered or unexported fields
}

Store holds the run's live completion condition and the record of how it got there.

It exists for the same reason todo.Store does: the state has to outlive the transcript. A condition that lived only in the system message would be rewritten by compaction like anything else, and one that lived only in the gate's own field could not be read by the tool that changes it. So the condition is held here, the gate reads it, the tool writes it, and the loop drains it into the durable log.

A nil *Store is not valid; the plugin always builds one.

func NewStore

func NewStore(initial string) *Store

NewStore builds a store holding the run's starting condition.

func (*Store) ApplyCommand

func (s *Store) ApplyCommand(ctx context.Context, c Command) (State, error)

ApplyCommand lets a host pause a live run at its next settled boundary. It is safe to call concurrently; resume across processes uses NativeRun.Commands.

func (*Store) Goal

func (s *Store) Goal() string

Goal returns the condition currently in force.

func (*Store) Revisions

func (s *Store) Revisions() []Revision

Revisions returns the audit trail, oldest first.

func (*Store) State

func (s *Store) State() State

func (*Store) Update

func (s *Store) Update(goal, reason string) bool

Update replaces the condition and marks it for the loop to persist. A revision to the identical condition is dropped: it is not a change, and letting it through would write an EntryGoal and rebuild the system prompt — invalidating the provider's KV cache for the whole prefix — in exchange for nothing.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL