Documentation
¶
Overview ¶
Package runstate is the vocabulary of a managed-script run's queue history (#1859, #1860): how each attempt ended, why a failed run failed and whether running it again is expected to help, and what a running run's worker is doing. It is its own package so the run record in pkg/script can carry it without growing that package's public surface, and so the store, the worker, the tool layer and the portal all spell it one way.
Index ¶
Constants ¶
const ( // AttemptFinished is an attempt that recorded the run's verdict: // succeeded, failed or canceled. AttemptFinished = "finished" // AttemptRetried is an attempt that failed on a platform fault and put the // run back on the queue under the attempt budget. AttemptRetried = "retried" // AttemptReleased is an attempt its worker handed back at shutdown. AttemptReleased = "released" // AttemptShed is an attempt its worker stopped and requeued to relieve // memory while other runs executed beside it. AttemptShed = "shed" // AttemptLeaseExpired is an attempt whose worker stopped without reporting // a result: its lease ran out with the run still marked running, which is // what an OOM kill, a crash or a lost node leaves behind. AttemptLeaseExpired = "lease_expired" // AttemptUnresponsive is an attempt a cancel ended because its worker had // stopped reporting while its lease still ran. AttemptUnresponsive = "unresponsive" )
How an attempt of a run ended (#1860), as its history records it.
const ( // CauseScript is a failure the script produced: an evaluation error, a // fail(), an argument a binding refused, a limit it exceeded. A script // that reads the outside world can fail once and succeed on its next run, // so this names where the failure was raised, not that it will repeat // (#1935). CauseScript = "script" // CauseUpstream is a failure whose cause is outside the script and // usually temporary: an upstream that timed out, dropped the connection, // kept refusing with 429 or 503 past the host's retries, or answered the // call the script made last with a 5xx or 429 before the script failed // (#1935). CauseUpstream = "upstream" // CauseTransient is a failure the script itself declared temporary with // fail(..., retryable=True) (#1935). CauseTransient = "transient" // CauseMemory is a run stopped for holding more memory than its budget, // or more than its replica could give it (#1861). CauseMemory = "memory" // CauseWorkerLost is a run whose workers kept stopping without reporting // a result until its reclaims were spent (#1860). CauseWorkerLost = "worker_lost" // CausePlatform is a platform fault that outlasted the attempt budget: // the run's session or its script could not be opened or read. CausePlatform = "platform" // CauseStateConflict is a run whose platform.save_state was refused // because another run of the script saved first (#1537). Its outputs // stand; run again, it reads the state the other run saved. CauseStateConflict = "state_conflict" )
Why a failed run failed (#1859), which decides whether running it again is expected to succeed.
const ( // LivenessExecuting is a run whose worker holds the lease and reports. LivenessExecuting = "executing" // LivenessUnresponsive is a run whose worker holds the lease but has not // reported for HeartbeatStaleAfter: most often a replica that was killed, // whose lease has not yet run out. LivenessUnresponsive = "unresponsive" // LivenessLeaseExpired is a run whose lease ran out with no worker // holding it; the next claim takes it over, or fails it once its // reclaims are spent. LivenessLeaseExpired = "lease_expired" )
What a running run's holder is doing, as Liveness reports it.
const DefaultMaxReclaims = 2
DefaultMaxReclaims is how many times a run is taken over from a worker whose lease expired before it is failed instead (scripts.worker.max_reclaims). A run that kills its worker kills the next one too, so the bound is small.
const HeartbeatStaleAfter = 30 * time.Second
HeartbeatStaleAfter is how long a running run's worker may go without reporting before the run reads as unresponsive and a cancel ends it directly (#1860). A worker reports every two seconds whatever the script is doing, so this is fifteen missed reports.
Variables ¶
This section is empty.
Functions ¶
func CauseRetryable ¶
CauseRetryable reports whether a run that failed for cause is expected to succeed when it runs again unchanged: an upstream that was unavailable, a failure the script declared temporary, or a state another run moved first. The platform never runs one again on its own; this is what the owner and an agent are told.
Types ¶
type Attempt ¶
type Attempt struct {
Attempt int `json:"attempt"`
Worker string `json:"worker"`
ClaimedAt *time.Time `json:"claimed_at,omitempty"`
EndedAt *time.Time `json:"ended_at,omitempty"`
// Outcome is one of the Attempt* values.
Outcome string `json:"outcome"`
// Error is the failure or reason the attempt ended on, when it had one.
Error string `json:"error,omitempty"`
}
Attempt is one ended attempt of a run.
type FailureStreak ¶ added in v1.137.1
type FailureStreak struct {
// Failed is how many of the most recent finished runs failed in a row,
// newest first, up to the window the store reads; zero when the newest
// finished run succeeded.
Failed int
// SameError is how many of those, newest first, ended on the same last
// error line as the newest one. A failure that repeats word for word is
// one the next run is unlikely to get past on its own.
SameError int
// LastError is the newest failure's last error line, which names what
// went wrong: "Error in fail: fail: NWS returned 500".
LastError string
// LastFailedRunID, LastFailedVersion, LastFailedAt and LastCause describe
// the newest failed run, when the streak holds one: the run a reader opens
// to see why, the version it ran, and the cause it was recorded under.
LastFailedRunID string
LastFailedVersion int
LastFailedAt *time.Time
LastCause string
// LastSuccessAt is when the script's most recent successful run finished,
// nil when none has.
LastSuccessAt *time.Time
}
FailureStreak is how a script's recent finished runs stand (#1934, #1935): how many in a row failed, how many of those failed the same way, and when it last succeeded. A pending, running or canceled run is not counted either way.
type FailureStreakReader ¶ added in v1.137.1
type FailureStreakReader interface {
FailureStreaks(ctx context.Context, scriptIDs []string) (map[string]FailureStreak, error)
}
FailureStreakReader reads the failure streak of each of a set of scripts. A script with no finished run is absent from the map.