Documentation
¶
Overview ¶
Package lifecycle implements workspace lifecycle automation: the Reaper loop finds RUNNING agent runs that have been idle past their policy's AutoStopAfterSec threshold and stops them, emitting a "run.autostop" audit event for each.
Idleness definition — PARTIAL, and honestly so ¶
"Idleness" here is wall-clock age of agent_runs.updated_at. The clock resets only when something writes that row through the store — state transitions, sandbox_ref updates, or an explicit store.TouchRun keepalive. Two callers touch it today: the interactive-attach handler (a human is watching) and the proxy's decision ingest (the agent made an egress call), so a run doing either is measured on real activity. The ingest touch is coalesced to one UPDATE per TouchDebounce window, so thresholdFor adds TouchDebounce as slack — a run stops once idle past policy + TouchDebounce, never earlier.
What still does NOT reset the clock is activity that never leaves the sandbox — CPU, file writes, a long local build. Such a run is stopped once updated_at ages past its policy threshold even while busy. The remaining upgrade is runner-reported liveness; until then, operators who need an unbounded session should use the never-reap escape hatch (policy AutoStopAfterSec < 0). This residual risk is documented here and should be mirrored in threatmodel/ if appropriate.
AutoStopAfterSec semantics (policy auto_stop_after_sec) ¶
> 0 idle timeout in seconds (stop after this much wall-clock idleness)
0 DISABLED — the run is never reaped (matches docs/POLICIES.md; this is the
default for a run with no auto-stop configured, since the store COALESCEs
an absent policy field to 0)
< 0 never reaped (explicit interactive escape hatch; equivalent to 0 for the
reaper, but kept distinct so an operator can express intent loudly)
Index ¶
Constants ¶
const TouchDebounce = 30 * time.Second
TouchDebounce is how stale a run's activity signal may legitimately be: the decision-ingest touch (internal/api) coalesces bursts to one updated_at UPDATE per window, so an ACTIVE run's row can lag real activity by up to this much. thresholdFor adds it as slack, so the debounce can never make an active run look idle — a genuinely idle run stops at most this much later.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Config ¶
type Config struct {
// Interval is how often the reaper scans. Default: 1 minute.
Interval time.Duration
// Now overrides the wall clock. Nil means use real time.
// (overridable in tests, mirrors embedded.Provider's now func idiom)
Now func() time.Time
// TickLock, when non-nil, makes each tick single-flight across control
// planes: it must TRY to take a cluster-wide lock and return a release func,
// or (nil, false) when someone else holds it — in which case the tick is
// skipped, not queued. wardynd wires a Postgres try-advisory-lock
// (db.TryAdvisoryLock, key db.ReaperAdvisoryLockKey); it stays a func here so
// lifecycle keeps no database dependency, like Store/Stopper/Recorder.
//
// Nil = ungated, the single-process default (and what tests drive Tick with).
TickLock func(ctx context.Context) (release func(), ok bool)
}
Config holds optional overrides for Reaper behaviour. Zero value is valid: defaults are applied in New.
type Reaper ¶
type Reaper struct {
// contains filtered or unexported fields
}
Reaper is the idle-workspace garbage collector. It runs a periodic loop that finds RUNNING workspaces idle past their policy threshold and stops them.
Reaper.Run is designed to be called in a dedicated goroutine and blocks until ctx is cancelled. A stopper error is logged and skipped so that one broken sandbox never prevents the rest from being reaped on the same tick.
func New ¶
New constructs a Reaper. store, stopper, and recorder are required and must be non-nil. cfg may be zero-valued.
func (*Reaper) Run ¶
Run starts the reap loop. It returns when ctx is cancelled (typically at process shutdown). The first tick fires after one full Interval so that the control plane can finish starting up before the first scan.
func (*Reaper) Tick ¶
Tick is one reap scan, gated by Config.TickLock when one is wired. It is exported so that integration callers and tests can drive it directly; in production code always use Run instead.
The whole tick runs under an Interval+defaultStopTimeout deadline: the advisory-lock connection is held for the tick's duration, so a single wedged StopRun must not pin the lock (and with it, reaping cluster-wide) past that bound. defaultStopTimeout is added on top of Interval, not folded into it — a child context.WithTimeout can only SHORTEN its parent's deadline, so budgeting the tick at bare Interval silently capped every per-stop deadline below at whatever was left of Interval, never the full defaultStopTimeout the constant promises (a stop past ~1 minute failed on every tick and was retried forever, never actually stopping the run).
type Recorder ¶
type Recorder interface {
Record(ctx context.Context, ev types.AuditEvent) error
}
Recorder is the minimal audit interface the Reaper needs. It matches audit.Recorder exactly so the integrator can pass a store.Recorder directly.
type RunSummary ¶
type RunSummary struct {
ID uuid.UUID
UpdatedAt time.Time
// PolicyAutoStopAfterSec is the value from the run's attached policy spec.
// Semantics (see package doc): 0 means DISABLED (never reap), a NEGATIVE value
// also means "never reap", and a POSITIVE value is the idle timeout in seconds.
// The Store must JOIN to the policy table and surface this field; a zero value
// is intentional (a run with no auto-stop configured is never reaped, not a
// missing join).
PolicyAutoStopAfterSec int
}
RunSummary is the minimal projection a Store must return for idle detection. It carries everything the Reaper needs without leaking the full run row.
type StopOutcome ¶
type StopOutcome struct {
// Applied reports whether the stop actually transitioned the run from RUNNING
// to STOPPED. It is false when a concurrent kill/complete had already moved the
// run terminal, OR when the idleness guard no-op'd because the run's updated_at
// advanced past the reaper's snapshot (an active `wardyn attach` touched it
// after the scan — finding N3). Either way the reaper must NOT emit a spurious
// run.autostop (finding #1).
Applied bool
// Errors, when non-empty, carries the teardown/revocation failures that
// occurred AFTER the stop transition won (Applied=true) but which left the run
// not fully contained: "identity_error" / "broker_error" (the run token or
// minted broker creds may still be live — finding N1), and "teardown_error"
// (the sandbox may still be routable). The reaper emits these as a distinct
// run.revoke/failure audit event so the live-credential/live-sandbox window is
// visible, mirroring handleKillRun/revokeRunCascade. Nil on a clean stop.
Errors map[string]string
}
StopOutcome is what StopRun reports back to the Reaper beyond the raw error.
type Stopper ¶
type Stopper interface {
// StopRun gracefully stops the run identified by runID. Implementations must
// be idempotent: stopping an already-stopped run returns ({Applied:false}, nil).
//
// notAfter is the run's updated_at from the reaper's tick snapshot: the stop
// transition must be conditional on updated_at not having advanced past it, so
// a run an active attach touched after the snapshot is NOT stopped (finding N3).
//
// The returned StopOutcome reports whether the transition applied and any
// post-transition teardown/revocation failures. A non-nil error means the stop
// failed outright; the outcome is meaningless and the reaper logs and skips.
StopRun(ctx context.Context, runID uuid.UUID, notAfter time.Time) (StopOutcome, error)
}
Stopper stops a single run. The implementation is expected to conditionally transition the run's state to STOPPED (RUNNING->STOPPED only, guarded on the snapshot's idleness) and then call runner.StopSandbox + the revoke cascade; all concerns are hidden behind this interface so the lifecycle package stays target-agnostic.
type Store ¶
type Store interface {
// ListRunningWithPolicy returns all runs currently in state RUNNING
// together with the auto_stop_after_sec value from their attached policy
// (0 when no policy is attached or the policy field is unset).
ListRunningWithPolicy(ctx context.Context) ([]RunSummary, error)
}
Store is the narrow persistence interface the Reaper requires.
The real adapter is trivial: store.ListRunningWithPolicy wraps the pgxpool and maps from (agent_runs JOIN run_policies). Tests supply a fake.
Method shapes deliberately mirror the existing store package naming conventions (List*, returning a slice) so the integrator's adapter is mechanical.