run

package
v0.16.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 29, 2026 License: Apache-2.0 Imports: 23 Imported by: 0

Documentation

Overview

Package run reconciles Run entities: one sandbox, one command, one exit code, then teardown.

The controller is the only writer of Run.Status. Everything else -- the exit code the sandbox reports, a cancellation request from the CLI, a deadline passing -- arrives as an input it reads, never as a status another component writes. Splitting that ownership would be a race with no lock available to fix it, since the entity store has no cross-entity transaction.

Index

Constants

View Source
const (
	// DefaultTimeout bounds a run whose task didn't set one. Unbounded is the
	// wrong default: a forgotten run holds a concurrency slot indefinitely.
	DefaultTimeout = 8 * time.Hour

	// StartDeadline bounds how long a run may sit pending before its sandbox is
	// running. It covers image pulls and an unschedulable cluster, and replaces
	// the client-side two-minute wait the exec proxy used to impose -- which
	// only applied while a client was connected.
	StartDeadline = 5 * time.Minute

	// SweepInterval is how often deadlines are checked. The sweep does two
	// indexed lookups and enqueues; it transitions nothing.
	SweepInterval = 10 * time.Second
)
View Source
const SchedulerInterval = 15 * time.Second

SchedulerInterval is how often the scheduler looks for ticks that have come due. It bounds how late a run can start, not how precisely ticks are computed: tick times come from the calendar expression, so a slow poll delays a run without shifting the schedule.

Variables

This section is empty.

Functions

This section is empty.

Types

type Controller

type Controller struct {
	Log *slog.Logger
	EC  *entityserver.Client
	EAC *entityserver_v1alpha.EntityAccessClient

	// RC is this controller's own reconcile controller, used by the deadline
	// sweep and the sandbox bridge to enqueue work. Set after construction
	// because the reconcile controller needs the handler that wraps this.
	RC *controller.ReconcileController
}

Controller reconciles Run entities.

func (*Controller) Init

func (c *Controller) Init(ctx context.Context) error

func (*Controller) Reconcile

func (c *Controller) Reconcile(ctx context.Context, r *run_v1alpha.Run, meta *entity.Meta) error

Reconcile drives a run through its lifecycle. Every step is idempotent and re-entrant so framework retries, later watch events, and the deadline sweep can all safely resume from the current stored state.

func (*Controller) SweepDeadlines

func (c *Controller) SweepDeadlines(ctx context.Context) error

SweepDeadlines re-enqueues runs whose start deadline or timeout has passed.

It transitions nothing. Every status change stays inside Reconcile, which the framework serializes per entity -- a sweep that wrote statuses directly could race a watch-driven reconcile of the same run with no lock to prevent it.

type GCConfig

type GCConfig struct {
	// CheckInterval is how often a sweep happens.
	CheckInterval time.Duration

	// RetentionCount is how many terminal runs to keep per app regardless of
	// age. It is per app rather than per task on purpose: an app with two
	// hundred tasks would otherwise accumulate two hundred times this many.
	RetentionCount int

	// RetentionPeriod keeps runs newer than this regardless of count.
	RetentionPeriod time.Duration

	// ConsoleRetention is the shorter window for console runs. Their value is
	// an audit record and a recent scrollback, not a permanent log, and on an
	// app people debug regularly they would otherwise crowd out everything
	// worth reading.
	ConsoleRetention time.Duration

	// ScheduledFloor is the age below which a schedule-triggered run is never
	// deleted.
	//
	// This is a correctness constraint, not a tuning knob. The dedup guard for
	// a scheduled tick *is* the run entity's existence, so removing one while
	// any replica could still evaluate that tick -- one that was partitioned,
	// or restarted behind the others -- lets create-if-absent succeed a second
	// time and the job double-fires. The floor has to sit well beyond any
	// plausible partition, and scheduled runs are exempt from the count cap
	// entirely: a count applied naively to a busy app would evict same-day
	// ticks and silently reintroduce double-firing.
	ScheduledFloor time.Duration

	// OrphanSandbox is how long a terminal run's sandbox may linger before it
	// is torn down.
	OrphanSandbox time.Duration
}

GCConfig bounds how long runs are kept.

func DefaultGCConfig

func DefaultGCConfig() GCConfig

type GCController

type GCController struct {
	Log    *slog.Logger
	EC     *entityserver.Client
	EAC    *entityserver_v1alpha.EntityAccessClient
	Config GCConfig
	// contains filtered or unexported fields
}

GCController retires finished runs and the sandboxes they leave behind.

func (*GCController) RunGC

func (c *GCController) RunGC(ctx context.Context, now time.Time) (GCResult, error)

RunGC performs one sweep. now is a parameter so tests can drive it.

func (*GCController) Start

func (c *GCController) Start(ctx context.Context)

func (*GCController) Stop

func (c *GCController) Stop()

type GCResult

type GCResult struct {
	DeletedRuns    int
	FailedRuns     int
	RetainedRuns   int
	StoppedSandbox int
	ReapedOrphans  int
	TotalScanned   int
}

GCResult reports what one sweep did.

type SandboxWatchController

type SandboxWatchController struct {
	RunController *controller.ReconcileController
}

SandboxWatchController wakes the run controller when a run's sandbox changes.

It is needed because a reconcile controller watches exactly one index, and the run controller watches runs. A sandbox reaching STOPPED with an exit code produces no event on that index, so without this bridge a finished run would sit in running until the deadline sweep happened to look at it -- turning every completion into a delay of up to the sweep interval.

func NewSandboxWatchController

func NewSandboxWatchController(runController *controller.ReconcileController) *SandboxWatchController

func (*SandboxWatchController) Create

func (*SandboxWatchController) Delete

Delete matters as much as Update. A sandbox deleted out from under a running run means nothing will ever report its exit, so the run has to be told rather than left waiting for a timeout that may be hours away.

func (*SandboxWatchController) Init

func (*SandboxWatchController) Update

type Scheduler

type Scheduler struct {
	Log *slog.Logger
	EC  *entityserver.Client
	EAC *entityserver_v1alpha.EntityAccessClient

	Interval time.Duration
	// contains filtered or unexported fields
}

Scheduler creates runs for tasks whose calendar expression has come due.

It holds no leader election and no external coordination. A tick is a pure function of the stored calendar expression, so every replica derives exactly the same firing times; the run's entity id is derived from the tick, which makes creating it a put-if-absent that etcd resolves to a single winner. The losers see the run already exists and move on.

The consequence worth stating: the dedup guard is the run entity's existence. Deleting a tick's run resurrects that tick, which is why retention treats scheduled runs as a correctness constraint rather than a preference.

func (*Scheduler) Start

func (s *Scheduler) Start(ctx context.Context)

func (*Scheduler) Stop

func (s *Scheduler) Stop()

Stop cancels the scheduler and waits for an in-flight sweep to finish, so a caller shutting down in order can rely on no further runs being created.

func (*Scheduler) Sweep

func (s *Scheduler) Sweep(ctx context.Context, now time.Time) error

Sweep fires every tick that has come due since the last sweep.

now is a parameter so tests can drive it; in production it is time.Now.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL