runstate

package
v1.137.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 27, 2026 License: Apache-2.0 Imports: 2 Imported by: 0

Documentation

Overview

Package runstate is the vocabulary of a managed-script run's queue history (#1859, #1860): how each attempt ended, why a failed run failed and whether running it again is expected to help, and what a running run's worker is doing. It is its own package so the run record in pkg/script can carry it without growing that package's public surface, and so the store, the worker, the tool layer and the portal all spell it one way.

Index

Constants

View Source
const (
	// AttemptFinished is an attempt that recorded the run's verdict:
	// succeeded, failed or canceled.
	AttemptFinished = "finished"
	// AttemptRetried is an attempt that failed on a platform fault and put the
	// run back on the queue under the attempt budget.
	AttemptRetried = "retried"
	// AttemptReleased is an attempt its worker handed back at shutdown.
	AttemptReleased = "released"
	// AttemptShed is an attempt its worker stopped and requeued to relieve
	// memory while other runs executed beside it.
	AttemptShed = "shed"
	// AttemptLeaseExpired is an attempt whose worker stopped without reporting
	// a result: its lease ran out with the run still marked running, which is
	// what an OOM kill, a crash or a lost node leaves behind.
	AttemptLeaseExpired = "lease_expired"
	// AttemptUnresponsive is an attempt a cancel ended because its worker had
	// stopped reporting while its lease still ran.
	AttemptUnresponsive = "unresponsive"
)

How an attempt of a run ended (#1860), as its history records it.

View Source
const (
	// CauseScript is a failure the script produced: an evaluation error, a
	// fail(), an argument a binding refused, a limit it exceeded. A script
	// that reads the outside world can fail once and succeed on its next run,
	// so this names where the failure was raised, not that it will repeat
	// (#1935).
	CauseScript = "script"
	// CauseUpstream is a failure whose cause is outside the script and
	// usually temporary: an upstream that timed out, dropped the connection,
	// kept refusing with 429 or 503 past the host's retries, or answered the
	// call the script made last with a 5xx or 429 before the script failed
	// (#1935).
	CauseUpstream = "upstream"
	// CauseTransient is a failure the script itself declared temporary with
	// fail(..., retryable=True) (#1935).
	CauseTransient = "transient"
	// CauseMemory is a run stopped for holding more memory than its budget,
	// or more than its replica could give it (#1861).
	CauseMemory = "memory"
	// CauseWorkerLost is a run whose workers kept stopping without reporting
	// a result until its reclaims were spent (#1860).
	CauseWorkerLost = "worker_lost"
	// CausePlatform is a platform fault that outlasted the attempt budget:
	// the run's session or its script could not be opened or read.
	CausePlatform = "platform"
	// CauseStateConflict is a run whose platform.save_state was refused
	// because another run of the script saved first (#1537). Its outputs
	// stand; run again, it reads the state the other run saved.
	CauseStateConflict = "state_conflict"
)

Why a failed run failed (#1859), which decides whether running it again is expected to succeed.

View Source
const (
	// LivenessExecuting is a run whose worker holds the lease and reports.
	LivenessExecuting = "executing"
	// LivenessUnresponsive is a run whose worker holds the lease but has not
	// reported for HeartbeatStaleAfter: most often a replica that was killed,
	// whose lease has not yet run out.
	LivenessUnresponsive = "unresponsive"
	// LivenessLeaseExpired is a run whose lease ran out with no worker
	// holding it; the next claim takes it over, or fails it once its
	// reclaims are spent.
	LivenessLeaseExpired = "lease_expired"
)

What a running run's holder is doing, as Liveness reports it.

View Source
const DefaultMaxReclaims = 2

DefaultMaxReclaims is how many times a run is taken over from a worker whose lease expired before it is failed instead (scripts.worker.max_reclaims). A run that kills its worker kills the next one too, so the bound is small.

View Source
const HeartbeatStaleAfter = 30 * time.Second

HeartbeatStaleAfter is how long a running run's worker may go without reporting before the run reads as unresponsive and a cancel ends it directly (#1860). A worker reports every two seconds whatever the script is doing, so this is fifteen missed reports.

Variables

This section is empty.

Functions

func CauseRetryable

func CauseRetryable(cause string) bool

CauseRetryable reports whether a run that failed for cause is expected to succeed when it runs again unchanged: an upstream that was unavailable, a failure the script declared temporary, or a state another run moved first. The platform never runs one again on its own; this is what the owner and an agent are told.

Types

type Attempt

type Attempt struct {
	Attempt   int        `json:"attempt"`
	Worker    string     `json:"worker"`
	ClaimedAt *time.Time `json:"claimed_at,omitempty"`
	EndedAt   *time.Time `json:"ended_at,omitempty"`
	// Outcome is one of the Attempt* values.
	Outcome string `json:"outcome"`
	// Error is the failure or reason the attempt ended on, when it had one.
	Error string `json:"error,omitempty"`
}

Attempt is one ended attempt of a run.

type FailureStreak added in v1.137.1

type FailureStreak struct {
	// Failed is how many of the most recent finished runs failed in a row,
	// newest first, up to the window the store reads; zero when the newest
	// finished run succeeded.
	Failed int
	// SameError is how many of those, newest first, ended on the same last
	// error line as the newest one. A failure that repeats word for word is
	// one the next run is unlikely to get past on its own.
	SameError int
	// LastError is the newest failure's last error line, which names what
	// went wrong: "Error in fail: fail: NWS returned 500".
	LastError string
	// LastFailedRunID, LastFailedVersion, LastFailedAt and LastCause describe
	// the newest failed run, when the streak holds one: the run a reader opens
	// to see why, the version it ran, and the cause it was recorded under.
	LastFailedRunID   string
	LastFailedVersion int
	LastFailedAt      *time.Time
	LastCause         string
	// LastSuccessAt is when the script's most recent successful run finished,
	// nil when none has.
	LastSuccessAt *time.Time
}

FailureStreak is how a script's recent finished runs stand (#1934, #1935): how many in a row failed, how many of those failed the same way, and when it last succeeded. A pending, running or canceled run is not counted either way.

type FailureStreakReader added in v1.137.1

type FailureStreakReader interface {
	FailureStreaks(ctx context.Context, scriptIDs []string) (map[string]FailureStreak, error)
}

FailureStreakReader reads the failure streak of each of a set of scripts. A script with no finished run is absent from the map.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL