spill

package
v0.1.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 9, 2026 License: MIT Imports: 10 Imported by: 0

README

spill

Extension. Ejectable — the loop never names it. Without it, an oversized tool result is head+tail truncated and the omitted middle is gone from the run for good.

Model Experience

A tool returns more bytes than the inline cap
What the model sees

A bounded preview — head and tail, because the end of a long result usually carries the signal — followed by an omission notice, and a new tool it can call to read any part of what was cut.

Verbatim text for this field
(Omitted 41827 of 42394 bytes. Full result saved at: spill_9f3c1a…. Call
read_spill with this locator (and an offset/limit in bytes) to read any part of it.)

A backend may replace the trailing sentence with its own RetrievalHint.

Token effect

Capped. The replacement is provably <= MaxInlineBytes: the notice's byte cost is reserved out of the cap before the preview is sized. Since the original was above the cap, spilling can never make a result larger. What it changes is not the size but the recoverability — the same budget now buys a pointer to everything instead of a hole.

KV cache effect

Append-only. The replacement is written once, at the position the raw result would have occupied, and never rewritten. Nothing earlier in the prefix moves.

The model calls read_spill
What the model sees

A byte window with an explicit position header, so paging is unambiguous:

[bytes 0-8192 of 42394]
<content>

The tail read appends , end of artifact.

Token effect

Capped. One call returns at most 8 KB (defaultSpillReadLimit), regardless of what the model asks for. A large artifact is paged, never dumped.

KV cache effect

Append-only.

The run has no spill store, or the save fails
What the model sees

Exactly what it would see without this plugin installed: the loop's own head+tail truncation. The plugin returns "no opinion" rather than an error.

Token effect

Zero-direct. No change against the baseline.

KV cache effect

Independent.

Impact on the agent

  • Adds one tool (read_spill) that bypasses the permission gate. It is SelfGated: it can only return text this session produced and had truncated away. The store is fenced by session id, and a store implementing SpillOwner is asked directly before every read, so a locator minted by another run reads back as not-found — deliberately indistinguishable from a locator that never existed.
  • Takes over result bounding by declaring Replace, which is why interceptors receive the raw result: a lossless bound cannot be built on top of an already-truncated string.
  • A failed tool call is left alone — its "result" is an error string the loop is about to overwrite, so persisting it would spend storage on nothing.

Known limitations and deferred work

  • No eviction. MemorySpillStore holds every artifact for the lifetime of the store. A long-lived process that spills heavily grows without bound; supply a store with a retention policy.
  • No cross-run retrieval. A locator is fenced to its session, so a resumed run under the same session id can read its own spills, but a different run investigating the same task cannot. There is no operation to grant access.
  • The preview split is fixed at two-thirds head / one-third tail. A result whose signal sits in the middle is previewed badly, and there is no way for a tool to declare a better shape.
  • The provider still sees a text notice. The opaque locator also rides ToolTrace.SpillLocator and durable Message.ResultRef, so context pruning and fallback compaction can keep it reachable, but provider wire formats do not gain a vendor-specific artifact block; the active model reads the notice.

Documentation

Overview

Package spill keeps an oversized tool result out of the model's context WITHOUT destroying it.

The loop's default bound is head+tail truncation: the omitted middle is gone from the run for good, and the model is told a byte count it can never cash in. This plugin takes the bounding job over — it saves the full text and hands back a preview plus a locator — so the omitted span is one read_spill call away instead of a re-run of the tool.

Ported from deepseek-harness's spill capability family (packages/spill: dsh-spill / dsh-spill-local / dsh-spill-policy, MIT). The seam is theirs; the Go shape, the session fence, and the retrieval tool are ours.

Index

Constants

This section is empty.

Variables

View Source
var ErrSpillNotFound = errors.New("agentcore: spill artifact not found")

ErrSpillNotFound is returned by ReadText for an unknown or out-of-scope locator. It is deliberately indistinguishable from "belongs to another session": a locator is not an existence oracle.

Functions

This section is empty.

Types

type MemorySpillStore

type MemorySpillStore struct {
	// contains filtered or unexported fields
}

MemorySpillStore is an in-process SpillStore: artifacts live in a map for the lifetime of the store. It is the default when a run enables spill without supplying storage, and it is what the tests use. A consumer that needs spilled output to survive the process (or to be visible to another node) supplies its own store.

func NewMemorySpillStore

func NewMemorySpillStore() *MemorySpillStore

NewMemorySpillStore builds an empty in-process store.

func (*MemorySpillStore) OwnsSpill

func (m *MemorySpillStore) OwnsSpill(locator, sessionID string) bool

OwnsSpill reports whether a locator was minted for the given session. The loop checks this before serving a read, so the fence holds even for a store whose locators are guessable.

func (*MemorySpillStore) ReadText

func (m *MemorySpillStore) ReadText(_ context.Context, locator string, offset, limit int) (SpillSlice, error)

ReadText returns a bounded, rune-aligned slice of a stored artifact.

func (*MemorySpillStore) SaveText

SaveText stores content under a locator branded with the session, so a locator from another session cannot be read back through this store.

type Plugin

type Plugin struct {
	// Store persists the oversized text. Required — without it the plugin is
	// inert, so a partially wired composition degrades rather than fails.
	Store SpillStore
	// MaxInlineBytes is the model-facing cap for a plain-text tool result. A
	// result above it is spilled and replaced by a preview + notice sized to
	// stay within this cap. 0 uses the run's limits.MaxToolResultLen.
	MaxInlineBytes int
	// ExcludeTools are tool names whose results are never spilled. read_spill is
	// always excluded (a read of a spill must not spill again, which would
	// loop); add tools whose output is already bounded or must stay verbatim.
	ExcludeTools []string
}

Plugin installs oversized-output storage. It is both an agentcore.Plugin (so a composition can register it) and an agentcore.ExtensionFactory (so the loop can drive it), and it declines a run rather than failing one: a nil Store, or a run with truncation disabled, leaves the loop's default bounding in place.

func To

func To(store SpillStore) Plugin

To builds a plugin that spills into the given store, using the run's MaxToolResultLen as the inline cap.

func (Plugin) BeginRun

BeginRun resolves the policy against this run's session and limits. A nil Extension declines the run — normal, not an error.

func (Plugin) Name

func (Plugin) Name() string

Name identifies the plugin and the extension it installs.

func (Plugin) Register

func (p Plugin) Register(r *agentcore.Registry) error

Register adds the plugin as a run extension. It claims no seam, so it never conflicts with another plugin and unloading it removes the capability entirely.

type SpillOwner

type SpillOwner interface {
	OwnsSpill(locator, sessionID string) bool
}

SpillOwner is an optional SpillStore capability: a store that can answer whether a locator was minted for a given session lets the loop enforce the fence before it ever calls ReadText. A store that does not implement it is trusted to scope locators itself (it was handed the session id at save time).

type SpillRecord

type SpillRecord struct {
	// Locator is the opaque model-facing handle passed back to read_spill.
	Locator string
	// RetrievalHint is backend-supplied prose telling the model how to get the
	// rest (appended to the notice). Empty uses the default hint.
	RetrievalHint string
	// Bytes is the exact size of the persisted content.
	Bytes int
}

SpillRecord is a saved artifact's handle.

type SpillRequest

type SpillRequest struct {
	// SessionID is the save-time storage namespace AND the fence: read_spill
	// only serves a locator minted for the reading run's session, so one agent
	// can never read another's spilled output by guessing a locator.
	SessionID string
	// ToolName and CallID identify the call that produced the text. Descriptive
	// (naming, audit) — never interpreted for access control.
	ToolName string
	CallID   string
	// Label is a short human tag for the artifact ("result").
	Label string
	// SuggestedName is a naming hint (e.g. "run_sql.txt"), never a path: the
	// backend sanitizes it to a single safe segment before use.
	SuggestedName string
	// Content is the full text to persist (UTF-8).
	Content string
}

SpillRequest is one request to persist an oversized tool result.

type SpillSlice

type SpillSlice struct {
	Content string
	// Offset is the byte offset the returned content actually starts at (after
	// clamping and rune-boundary snapping).
	Offset int
	// Total is the artifact's full size in bytes.
	Total int
	// EOF reports whether this slice reaches the end of the artifact.
	EOF bool
}

SpillSlice is a bounded read of a saved artifact.

type SpillStore

type SpillStore interface {
	// SaveText persists content verbatim and returns its locator. An error
	// makes the policy fall back to plain truncation — a spill failure must
	// never turn a successful tool call into a failed one.
	SaveText(ctx context.Context, req SpillRequest) (SpillRecord, error)
	// ReadText returns a byte slice of a previously saved artifact. offset and
	// limit are in bytes; the implementation clamps them to the artifact and
	// snaps the returned slice to UTF-8 rune boundaries.
	ReadText(ctx context.Context, locator string, offset, limit int) (SpillSlice, error)
}

SpillStore persists a tool result too large to sit inline in the model's context and hands back a locator the model can read later.

The problem it solves: without it, an oversized result is cut down to limits.MaxToolResultLen by truncateMiddle and the omitted bytes are GONE — the model is told a number and can never recover the content. A 40 MB query export, a long build log, a fetched page: the agent sees a head and a tail and has to re-run the tool (paying for it twice) to get at the middle, or simply reasons over a hole. With a store configured, the full text is saved verbatim and the inline result becomes a bounded preview plus a locator, so the omitted span is one read_spill call away.

Implementations are consumer-supplied (agentray backs it with object storage keyed by session); MemorySpillStore is the in-process default used by tests and by runs that opt in without wiring durable storage.

Ported from deepseek-harness's spill capability family (packages/spill: dsh-spill / dsh-spill-local / dsh-spill-policy, MIT). The seam is theirs; the Go shape, the fencing, and the retrieval tool are ours.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL