Documentation
¶
Overview ¶
Package run reconciles Run entities: one sandbox, one command, one exit code, then teardown.
The controller is the only writer of Run.Status. Everything else -- the exit code the sandbox reports, a cancellation request from the CLI, a deadline passing -- arrives as an input it reads, never as a status another component writes. Splitting that ownership would be a race with no lock available to fix it, since the entity store has no cross-entity transaction.
Index ¶
- Constants
- type Controller
- type GCConfig
- type GCController
- type GCResult
- type SandboxWatchController
- func (w *SandboxWatchController) Create(ctx context.Context, sb *compute.Sandbox, meta *entity.Meta) error
- func (w *SandboxWatchController) Delete(ctx context.Context, id entity.Id, sb *compute.Sandbox) error
- func (w *SandboxWatchController) Init(ctx context.Context) error
- func (w *SandboxWatchController) Update(ctx context.Context, sb *compute.Sandbox, meta *entity.Meta) error
- type Scheduler
Constants ¶
const ( // DefaultTimeout bounds a run whose task didn't set one. Unbounded is the // wrong default: a forgotten run holds a concurrency slot indefinitely. DefaultTimeout = 8 * time.Hour // StartDeadline bounds how long a run may sit pending before its sandbox is // running. It covers image pulls and an unschedulable cluster, and replaces // the client-side two-minute wait the exec proxy used to impose -- which // only applied while a client was connected. StartDeadline = 5 * time.Minute // SweepInterval is how often deadlines are checked. The sweep does two // indexed lookups and enqueues; it transitions nothing. SweepInterval = 10 * time.Second )
const SchedulerInterval = 15 * time.Second
SchedulerInterval is how often the scheduler looks for ticks that have come due. It bounds how late a run can start, not how precisely ticks are computed: tick times come from the calendar expression, so a slow poll delays a run without shifting the schedule.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Controller ¶
type Controller struct {
Log *slog.Logger
EC *entityserver.Client
EAC *entityserver_v1alpha.EntityAccessClient
// RC is this controller's own reconcile controller, used by the deadline
// sweep and the sandbox bridge to enqueue work. Set after construction
// because the reconcile controller needs the handler that wraps this.
RC *controller.ReconcileController
}
Controller reconciles Run entities.
func NewController ¶
func NewController(log *slog.Logger, ec *entityserver.Client, eac *entityserver_v1alpha.EntityAccessClient) *Controller
func (*Controller) Reconcile ¶
func (c *Controller) Reconcile(ctx context.Context, r *run_v1alpha.Run, meta *entity.Meta) error
Reconcile drives a run through its lifecycle. Every step is idempotent and re-entrant so framework retries, later watch events, and the deadline sweep can all safely resume from the current stored state.
func (*Controller) SweepDeadlines ¶
func (c *Controller) SweepDeadlines(ctx context.Context) error
SweepDeadlines re-enqueues runs whose start deadline or timeout has passed.
It transitions nothing. Every status change stays inside Reconcile, which the framework serializes per entity -- a sweep that wrote statuses directly could race a watch-driven reconcile of the same run with no lock to prevent it.
type GCConfig ¶
type GCConfig struct {
// CheckInterval is how often a sweep happens.
CheckInterval time.Duration
// RetentionCount is how many terminal runs to keep per app regardless of
// age. It is per app rather than per task on purpose: an app with two
// hundred tasks would otherwise accumulate two hundred times this many.
RetentionCount int
// RetentionPeriod keeps runs newer than this regardless of count.
RetentionPeriod time.Duration
// ConsoleRetention is the shorter window for console runs. Their value is
// an audit record and a recent scrollback, not a permanent log, and on an
// app people debug regularly they would otherwise crowd out everything
// worth reading.
ConsoleRetention time.Duration
// ScheduledFloor is the age below which a schedule-triggered run is never
// deleted.
//
// This is a correctness constraint, not a tuning knob. The dedup guard for
// a scheduled tick *is* the run entity's existence, so removing one while
// any replica could still evaluate that tick -- one that was partitioned,
// or restarted behind the others -- lets create-if-absent succeed a second
// time and the job double-fires. The floor has to sit well beyond any
// plausible partition, and scheduled runs are exempt from the count cap
// entirely: a count applied naively to a busy app would evict same-day
// ticks and silently reintroduce double-firing.
ScheduledFloor time.Duration
// OrphanSandbox is how long a terminal run's sandbox may linger before it
// is torn down.
OrphanSandbox time.Duration
}
GCConfig bounds how long runs are kept.
func DefaultGCConfig ¶
func DefaultGCConfig() GCConfig
type GCController ¶
type GCController struct {
Log *slog.Logger
EC *entityserver.Client
EAC *entityserver_v1alpha.EntityAccessClient
Config GCConfig
// contains filtered or unexported fields
}
GCController retires finished runs and the sandboxes they leave behind.
func NewGCController ¶
func NewGCController(log *slog.Logger, ec *entityserver.Client, eac *entityserver_v1alpha.EntityAccessClient) *GCController
func (*GCController) Start ¶
func (c *GCController) Start(ctx context.Context)
func (*GCController) Stop ¶
func (c *GCController) Stop()
type GCResult ¶
type GCResult struct {
DeletedRuns int
FailedRuns int
RetainedRuns int
StoppedSandbox int
ReapedOrphans int
TotalScanned int
}
GCResult reports what one sweep did.
type SandboxWatchController ¶
type SandboxWatchController struct {
RunController *controller.ReconcileController
}
SandboxWatchController wakes the run controller when a run's sandbox changes.
It is needed because a reconcile controller watches exactly one index, and the run controller watches runs. A sandbox reaching STOPPED with an exit code produces no event on that index, so without this bridge a finished run would sit in running until the deadline sweep happened to look at it -- turning every completion into a delay of up to the sweep interval.
func NewSandboxWatchController ¶
func NewSandboxWatchController(runController *controller.ReconcileController) *SandboxWatchController
func (*SandboxWatchController) Delete ¶
func (w *SandboxWatchController) Delete(ctx context.Context, id entity.Id, sb *compute.Sandbox) error
Delete matters as much as Update. A sandbox deleted out from under a running run means nothing will ever report its exit, so the run has to be told rather than left waiting for a timeout that may be hours away.
type Scheduler ¶
type Scheduler struct {
Log *slog.Logger
EC *entityserver.Client
EAC *entityserver_v1alpha.EntityAccessClient
Interval time.Duration
// contains filtered or unexported fields
}
Scheduler creates runs for tasks whose calendar expression has come due.
It holds no leader election and no external coordination. A tick is a pure function of the stored calendar expression, so every replica derives exactly the same firing times; the run's entity id is derived from the tick, which makes creating it a put-if-absent that etcd resolves to a single winner. The losers see the run already exists and move on.
The consequence worth stating: the dedup guard is the run entity's existence. Deleting a tick's run resurrects that tick, which is why retention treats scheduled runs as a correctness constraint rather than a preference.
func NewScheduler ¶
func NewScheduler(log *slog.Logger, ec *entityserver.Client, eac *entityserver_v1alpha.EntityAccessClient) *Scheduler