serverlifecycle

package
v0.16.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 22, 2026 License: Apache-2.0 Imports: 20 Imported by: 0

Documentation

Overview

Package serverlifecycle restarts and upgrades the server as durable, idempotent operations: JSON records on disk that outlive the process they act on, so a caller can hand off an operation id and ask about it later.

Index

Constants

View Source
const (
	DefaultDir = "/var/lib/miren/server/lifecycle"
	RunnerDir  = "/var/lib/miren/runner/lifecycle"
)

DefaultDir is the server's ledger; RunnerDir is the runner's. They are separate ledgers with separate busy slots: a host that runs both daemons (the coordinator in standalone mode does not, but nothing forbids it) can have one operation on each.

View Source
const AdoptionTimeout = 2 * time.Minute

AdoptionTimeout is how long a handed-off operation may wait for a server to take it over before the hand-off is treated as failed.

View Source
const DefaultExitAfter = 8 * time.Minute

DefaultExitAfter is past the server's own shutdown timeout, so it only ever fires on a stop that is stuck rather than slow.

View Source
const DefaultHealthURL = "https://127.0.0.1:443/.well-known/miren/health"

Variables

View Source
var (
	ServerExecutorCommand = []string{"server", "operations", "run"}
	RunnerExecutorCommand = []string{"runner", "operations", "run"}
)

The executor entrypoints, one per ledger. Each opens its own daemon's ledger and probes its own daemon, so an operation id alone does not say which one it belongs to; the launcher has to.

View Source
var ErrBusy = errors.New("another lifecycle operation is in progress")
View Source
var ErrInvalidID = errors.New("lifecycle operation id must be a ULID")

ErrInvalidID is returned for an operation id that is not a ULID. The id is a file name and the ledger's sort key, so the format is not negotiable.

View Source
var ErrLauncherStopping = errors.New("server is shutting down; try again once it is back")

ErrLauncherStopping is returned by Launch once the server is going down.

View Source
var ErrLocked = errors.New("lifecycle operation is being executed by another process")

ErrLocked is returned when another executor already holds an operation.

View Source
var ErrNotFound = errors.New("lifecycle operation not found")
View Source
var ErrNothingToAbandon = errors.New("operation is finished and left nothing pending")

ErrNothingToAbandon is returned by Abandon for an operation that is finished and left nothing pending.

View Source
var ErrUnsupervised = errors.New("this server is not supervised in a way that can restart it: install it as a systemd service, or run the container image through `miren server container install` so its restart policy brings it back")

ErrUnsupervised is what a restart or upgrade gets on a server nothing can bring back: not a systemd unit, and not a container that booted through the image's entrypoint.

Functions

func ExecutorBinary

func ExecutorBinary() (string, error)

ExecutorBinary is the binary a launcher should run the executor from: this process's own, resolved through any symlink so the transient unit captures the real file. Running our own build rather than the installed server's matters when they differ: an older server may predate the executor entirely.

func HealthURL

func HealthURL(mode, address string) string

HealthURL is where a server with this ingress mode and address answers /.well-known/miren/health from its own host.

func NewID

func NewID() string

NewID mints a ULID so a directory listing sorts by creation time.

func SameBuild

func SameBuild(versionA, commitA, versionB, commitB string) bool

SameBuild reports whether two builds are the same: commits decide when both are known, otherwise version strings.

func UnitActive

func UnitActive(ctx context.Context, unit string) bool

UnitActive reports whether a unit is running or starting.

func UnitName

func UnitName(opID string) string

func ValidateID

func ValidateID(id string) error

ValidateID checks that id is a ULID as NewID would mint one.

Types

type Action

type Action string
const (
	ActionRestart Action = "restart"
	ActionUpgrade Action = "upgrade"
)

type ContainerLauncher

type ContainerLauncher struct {
	// NewExecutor builds an executor over the store; it is called once per
	// launch or resume so each run has its own downloaded-artifact state.
	NewExecutor func() (*Executor, error)
	Log         *slog.Logger
	// contains filtered or unexported fields
}

ContainerLauncher runs the executor inside the server process. Nothing outlives the server in a container, so the executor dies with it at the restart it asked for. That is fine: the operation record in the volume is the checkpoint, and the next instance's Resume picks it up where it was, verifying itself as the new instance.

func NewContainerLauncher

func NewContainerLauncher(ctx context.Context, log *slog.Logger, newExecutor func() (*Executor, error)) *ContainerLauncher

NewContainerLauncher binds runs to ctx. Wait returns once every run launched under it has returned.

func (*ContainerLauncher) Launch

func (l *ContainerLauncher) Launch(_ context.Context, opID string) error

func (*ContainerLauncher) Resume

func (l *ContainerLauncher) Resume(ctx context.Context, store *Store) error

Resume relaunches the operation the previous instance left unfinished, if there is one. It answers the question a systemd install never has to ask: the transient unit there outlives the restart, but here the executor was the process that just exited.

func (*ContainerLauncher) Wait

func (l *ContainerLauncher) Wait()

Wait refuses further launches and blocks until every launched run has returned. Runs return promptly once the launcher's context is cancelled.

type ContainerRestarter

type ContainerRestarter struct {
	// Shutdown asks this process to stop. The default sends SIGTERM to
	// itself, which the server handles the same way as one from outside: the
	// boot graph's stop path takes the nested stack down cleanly.
	Shutdown func() error
	// ExitAfter bounds the graceful stop: a restart that never finishes
	// stopping would leave the container running the build it was meant to
	// replace, with no supervisor to notice. 0 means DefaultExitAfter.
	ExitAfter time.Duration
}

ContainerRestarter restarts a container install by ending the process. There is no supervisor inside the container to ask; the one outside it (the container runtime's restart policy) starts a new container when this one exits, and container-boot in the image execs whatever the volume's release directory holds by then.

func (ContainerRestarter) Restart

type DataBackup

type DataBackup interface {
	// Backup takes a snapshot for the operation and returns its reference.
	Backup(ctx context.Context, opID string) (string, error)
}

DataBackup snapshots the server's data before an upgrade. The executor does not know what the data is: it stores the reference it gets back on the operation, and on rollback hands it to the restarted server, which does know how to put it back before serving.

type DataRestore

type DataRestore struct {
	BackupRef string `json:"backup_ref"`
	// ForVersion and ForCommit name the build the request is for: the one
	// being rolled back to. The request is persisted before the binary is
	// swapped, and the build being rolled back from may still be crash
	// looping under systemd at that point. If it honored the request it
	// would restore, migrate the data again, and fail, leaving a settled
	// request and migrated data for the build that actually needed it.
	ForVersion string `json:"for_version,omitempty"`
	ForCommit  string `json:"for_commit,omitempty"`

	RestoredAt *time.Time `json:"restored_at,omitempty"`
	Error      string     `json:"error,omitempty"`
}

DataRestore is the restore request a rollback records on the operation, and its outcome.

func (*DataRestore) MeantFor

func (r *DataRestore) MeantFor(version, commit string) bool

MeantFor reports whether the build identified by version and commit is the one this request is for. A request that could not name a build (the server was unreachable when the operation began) is for whoever boots.

type Executor

type Executor struct {
	// contains filtered or unexported fields
}

Executor drives an operation through its phases, persisting each transition. Everything needed to resume is in the Operation record.

func NewExecutor

func NewExecutor(store *Store, opts Options, log *slog.Logger) *Executor

func (*Executor) Run

func (e *Executor) Run(ctx context.Context, id string) (*Operation, error)

Run drives the operation to a terminal phase, to the hand-off to the server, or until ctx ends. A finished operation is returned unchanged; an interrupted one resumes where it was.

func (*Executor) WithDataBackup

func (e *Executor) WithDataBackup(b DataBackup) *Executor

func (*Executor) WithDownloader

func (e *Executor) WithDownloader(d release.Downloader) *Executor

func (*Executor) WithInstaller

func (e *Executor) WithInstaller(i release.Installer) *Executor

func (*Executor) WithProber

func (e *Executor) WithProber(p Prober) *Executor

func (*Executor) WithRestarter

func (e *Executor) WithRestarter(r Restarter) *Executor

type HealthProber

type HealthProber struct {
	URL string
	// ServerName is the SNI to send. Empty works with the default autocert
	// ingress, which answers SNI-less handshakes with its self-signed cert.
	ServerName string
	// contains filtered or unexported fields
}

HealthProber reads the "server" block of /.well-known/miren/health.

func NewHealthProber

func NewHealthProber(url, serverName string) *HealthProber

NewHealthProber accepts any certificate: the point is to reach the process on this host, not to authenticate it.

func (*HealthProber) Probe

func (p *HealthProber) Probe(ctx context.Context) (Snapshot, error)

type InstanceProber

type InstanceProber struct {
	Source *serverinfo.Source
}

InstanceProber reads this process's own serverinfo instead of asking the health endpoint. Inside the container the executor and the server are the same process, so there is nothing to reach over the network, and the answer is available before ingress is listening.

func (InstanceProber) Probe

type LaunchChecker

type LaunchChecker interface {
	Launched(ctx context.Context, opID string) bool
}

LaunchChecker is an optional Launcher capability: after Launch reported an error, it says whether the executor is running anyway. systemd-run can fail after the unit was submitted, and an executor that is running owns the record, so the caller must not mark it failed on top of it.

type Launcher

type Launcher interface {
	Launch(ctx context.Context, opID string) error
}

Launcher starts an executor somewhere that outlives the caller and the server.

type NodeStep

type NodeStep struct {
	Name     string `json:"name"`
	RunnerID string `json:"runner_id"`
	// OperationID is the runner-side operation, minted before the request
	// so a repeated request finds it rather than starting another.
	OperationID string `json:"operation_id,omitempty"`
	// Phase is the runner operation's phase as last observed, PhasePending
	// before the runner was asked, or StepSkipped.
	Phase           Phase  `json:"phase"`
	Error           string `json:"error,omitempty"`
	Progress        string `json:"progress,omitempty"`
	PreviousVersion string `json:"previous_version,omitempty"`
	NewVersion      string `json:"new_version,omitempty"`
	// Cordoned records that the walk cordoned the node for this step, so a
	// resumed walk knows to uncordon it. An operator's cordon is left alone.
	Cordoned bool `json:"cordoned,omitempty"`

	StartedAt  *time.Time `json:"started_at,omitempty"`
	FinishedAt *time.Time `json:"finished_at,omitempty"`
}

NodeStep is one runner's part of an upgrade. The runner runs its own operation in its own ledger; the step mirrors what the server last saw of it, so the cluster record stands on its own.

func (*NodeStep) Done

func (s *NodeStep) Done() bool

func (*NodeStep) Succeeded

func (s *NodeStep) Succeeded() bool

type Operation

type Operation struct {
	ID          string `json:"id"`
	Action      Action `json:"action"`
	RequestedBy string `json:"requested_by,omitempty"`

	// TargetVersion is as requested ("latest", "main", a tag); ResolvedVersion
	// and ResolvedCommit are what it became.
	TargetVersion   string `json:"target_version,omitempty"`
	ResolvedVersion string `json:"resolved_version,omitempty"`
	ResolvedCommit  string `json:"resolved_commit,omitempty"`
	// ArtifactType ("base" or "release"), NoRollback, and ReadyTimeoutSeconds
	// override the executor defaults for one upgrade.
	ArtifactType        string `json:"artifact_type,omitempty"`
	NoRollback          bool   `json:"no_rollback,omitempty"`
	ReadyTimeoutSeconds int    `json:"ready_timeout_seconds,omitempty"`

	Phase Phase  `json:"phase"`
	Error string `json:"error,omitempty"`

	// BackupRef names the data snapshot taken before an upgrade; empty when
	// none was taken. DataRestore is set on rollback and asks the restarted
	// server to put BackupRef back before it serves data; the server's answer
	// lands in it once the executor has read the RestoreResult.
	BackupRef   string       `json:"backup_ref,omitempty"`
	DataRestore *DataRestore `json:"data_restore,omitempty"`
	// Progress is a short note for the current phase, e.g. a download percentage.
	Progress string `json:"progress,omitempty"`

	// DrivenBy is the server instance that took the operation over for the
	// runner phase, and Nodes is one step per runner it found. Both are set
	// together, so DrivenBy on a record with no Nodes means a cluster with
	// nothing to walk.
	DrivenBy string      `json:"driven_by,omitempty"`
	Nodes    []*NodeStep `json:"nodes,omitempty"`

	// A successful operation ends with NewInstanceID != PreviousInstanceID;
	// that change is how we know the restart actually happened.
	PreviousInstanceID string `json:"previous_instance_id,omitempty"`
	PreviousVersion    string `json:"previous_version,omitempty"`
	PreviousCommit     string `json:"previous_commit,omitempty"`
	NewInstanceID      string `json:"new_instance_id,omitempty"`
	NewVersion         string `json:"new_version,omitempty"`
	// Components are the runtime versions (containerd, runc, ...) the server
	// reported once it was up. A base upgrade replaces those binaries next to
	// miren, and this is the record that the restarted server is on them.
	Components map[string]string `json:"components,omitempty"`

	// RollbackFrom names the instance that asked for the rollback's restart,
	// so a resumed rollback can tell a restart that already took (the
	// instance answering now is a different one) from one still to do. The
	// systemd executor outlives the restart and never resumes; the container
	// executor is that instance and dies at the restart it asks for.
	// container-boot, rolling back a build that never answered, writes its
	// own name.
	RollbackFrom string `json:"rollback_from,omitempty"`
	// BootAttempts counts container boots of the new build since the restart
	// phase, kept by container-boot; it is what catches a build that crashes
	// before the executor inside it can run.
	BootAttempts int `json:"boot_attempts,omitempty"`

	CreatedAt  time.Time  `json:"created_at"`
	UpdatedAt  time.Time  `json:"updated_at"`
	FinishedAt *time.Time `json:"finished_at,omitempty"`
}

Operation is the durable record of one restart or upgrade.

func Abandon

func Abandon(store *Store, id string) (*Operation, error)

Abandon is the operator's way out of an operation nobody will finish: an executor that died mid-operation, or a rollback whose data restore keeps failing and keeps the server from booting. It marks an unfinished operation failed and settles a pending restore request as abandoned, so the next boot starts on the data as it is. It refuses while an executor still holds the operation.

func AwaitAdoption

func AwaitAdoption(ctx context.Context, store *Store, id string, timeout, interval time.Duration) (*Operation, error)

AwaitAdoption waits for the server to take over a handed-off operation. The executor cannot know whether the build it just verified drives runner upgrades: a target older than that support would leave the record in upgrading_runners with nobody to finish it. If nothing has claimed it by the deadline, the operation is failed with that diagnosis. The wait polls rather than holds the operation lock, since the server needs the lock to adopt it.

func NewOperation

func NewOperation(action Action, requestedBy string) *Operation

func Start

func Start(ctx context.Context, store *Store, launcher Launcher, op *Operation) (*Operation, bool, error)

Start records op and launches its executor. It is closed over the id: a second Start with an id already on disk returns that record untouched, so a caller that lost the first reply (the reply may well be lost, since the operation restarts the server answering it) can retry without starting a second operation. Two such calls arriving together are settled under the store's lock, so exactly one creates and the other gets the record. A record that exists with a different action or target is a caller bug and is refused rather than silently returned. The bool reports whether this call is the one that started it.

The launch runs on a context cut loose from the caller's. The caller is an RPC or an uplink session about to be torn down by the very restart it asked for, and a launch cancelled midway is the worst outcome: the unit may be running while the caller believes it is not.

func (*Operation) Adopted

func (o *Operation) Adopted() bool

Adopted reports whether a server has taken over a handed-off operation.

func (*Operation) Done

func (o *Operation) Done() bool

func (*Operation) HandedOff

func (o *Operation) HandedOff() bool

HandedOff reports whether the executor is finished with the operation but the operation is not: the server drives it from here.

func (*Operation) Succeeded

func (o *Operation) Succeeded() bool

type Options

type Options struct {
	ServiceName string
	// Daemon names what is being restarted in progress notes and errors:
	// "server" or "runner".
	Daemon string
	// StateDir is the daemon's state directory, passed through to the
	// resource-limit refresh that precedes a restart.
	StateDir     string
	InstallPath  string
	TempDir      string
	ArtifactType release.ArtifactType
	// ReadyTimeout bounds how long a restarted server gets to report ready
	// before the operation fails (and, for upgrades, rolls back).
	ReadyTimeout  time.Duration
	ProbeInterval time.Duration
	AutoRollback  bool
	// PathSymlink, when set, is kept pointing at InstallPath after a
	// successful upgrade so the CLI on $PATH tracks the server.
	PathSymlink string
	// UpgradeRunners hands a verified upgrade to the daemon for the
	// upgrading_runners phase instead of finishing it. On for the server,
	// which has runners; off for a runner, which is one.
	UpgradeRunners bool
}

Options tunes an Executor. Start from DefaultOptions; a zero Options is not a working configuration.

func DefaultOptions

func DefaultOptions() Options

func RunnerOptions

func RunnerOptions() Options

RunnerOptions is DefaultOptions for the runner daemon: its unit, its state directory, and no data backup, since a runner keeps no data of its own that a rollback would need to put back. The prober is the caller's to set; a runner has no health URL.

type Phase

type Phase string

Phase is a checkpoint: an executor resumes from the recorded phase, and each phase's work is safe to repeat. The list is meant to grow: the pre-upgrade backup today is an etcd snapshot, and the full RFD-75 bundle slots into the same phase and BackupRef.

const (
	PhasePending     Phase = "pending"
	PhaseDownloading Phase = "downloading"
	// PhaseBackingUp sits after the download so the snapshot is as fresh as
	// possible when the restart happens, and so a failed download or an
	// already-installed target never costs a snapshot.
	PhaseBackingUp  Phase = "backing_up"
	PhaseInstalling Phase = "installing"
	PhaseRestarting Phase = "restarting"
	PhaseVerifying  Phase = "verifying"
	// PhaseUpgradingRunners follows a verified server upgrade on a cluster
	// with runners. The executor hands the operation to the server here:
	// the server has the node inventory and the RPC to each runner, and it
	// is the new build, which is the one that knows the runner protocol.
	PhaseUpgradingRunners Phase = "upgrading_runners"
	PhaseRollingBack      Phase = "rolling_back"

	PhaseSucceeded  Phase = "succeeded"
	PhaseFailed     Phase = "failed"
	PhaseRolledBack Phase = "rolled_back"

	// StepSkipped is terminal and step-only: the runner was not asked to
	// upgrade, and the step's Error says why (not ready, predates managed
	// upgrades). An operation never has it.
	StepSkipped Phase = "skipped"
)

func (Phase) Terminal

func (p Phase) Terminal() bool

type Prober

type Prober interface {
	Probe(ctx context.Context) (Snapshot, error)
}

Prober observes the running server: which process is there before acting, and when the new one is up.

type Restarter

type Restarter interface {
	Restart(ctx context.Context) error
}

type RestoreResult

type RestoreResult struct {
	OperationID string    `json:"operation_id"`
	BackupRef   string    `json:"backup_ref"`
	RestoredAt  time.Time `json:"restored_at"`
	Error       string    `json:"error,omitempty"`
	// Abandoned records that an operator gave up on the restore and told the
	// server to start on the data as it is.
	Abandoned bool `json:"abandoned,omitempty"`
}

RestoreResult is what the server writes after acting on a DataRestore request during boot. It is a file beside the operation record rather than a field in it because the executor keeps its own copy of the record and rewrites it while the server boots; a second writer would be overwritten.

A result with an Error does not settle the request: the server refuses to start on data it was asked to replace, and tries again on its next boot. Only a successful restore or an operator's Abandon ends that.

func (*RestoreResult) Settled

func (r *RestoreResult) Settled() bool

Settled reports whether the request this result answers is over, one way or the other.

type Snapshot

type Snapshot struct {
	InstanceID  string
	Version     string
	Commit      string
	Ready       bool
	InstallKind string
	// Components are the runtime versions the server reports driving
	// (containerd, runc, ...); nil from a server that predates the report.
	Components map[string]string
}

type Store

type Store struct {
	// contains filtered or unexported fields
}

Store keeps one JSON file per operation, written atomically so a process dying mid-write leaves the previous record intact.

func NewStore

func NewStore(dir string) (*Store, error)

func (*Store) Active

func (s *Store) Active() (*Operation, error)

func (*Store) Create

func (s *Store) Create(op *Operation) error

Create refuses while another operation is running: two executors racing on one binary and one systemd unit cannot both be right. The check and the write happen under a directory lock so two concurrent callers cannot both pass it.

func (*Store) CreateOrGet

func (s *Store) CreateOrGet(op *Operation) (existing *Operation, created bool, err error)

CreateOrGet records op unless a record with its id already exists, in which case that record is returned instead and created is false. The lookup and the write share the directory lock, so two callers carrying one id cannot both miss and cannot both create; exactly one of them creates. Like Create, it refuses to create while another operation is running.

func (*Store) Dir

func (s *Store) Dir() string

func (*Store) Get

func (s *Store) Get(id string) (*Operation, error)

func (*Store) List

func (s *Store) List() ([]*Operation, error)

List returns every operation, oldest first. A record that cannot be read is left out rather than hiding the rest; callers that must not miss one use listStrict.

func (*Store) LockOperation

func (s *Store) LockOperation(id string) (func(), error)

LockOperation claims id for one executor. It does not wait: a second executor on the same operation gets ErrLocked immediately.

func (*Store) PendingRestore

func (s *Store) PendingRestore() (*Operation, error)

PendingRestore returns the operation whose rollback asked the server to restore data and has not been answered, or nil. The server calls this early in boot, before the data it would restore is in use.

The request outlives its operation on purpose. An executor that waited on a server refusing to boot marks the operation failed and exits, and if that ended the request the very next boot would start on the data the rollback was meant to replace. Only a restore or Abandon settles it.

Records are read strictly: a request in a record that cannot be decoded is still a request, and the caller fails closed on the error.

func (*Store) ReadRestoreResult

func (s *Store) ReadRestoreResult(id string) (*RestoreResult, error)

func (*Store) RecordRestoreAttempt

func (s *Store) RecordRestoreAttempt(result *RestoreResult) (*RestoreResult, error)

RecordRestoreAttempt writes the outcome of a restore attempt, unless an operator abandoned the request while the attempt ran: a failed attempt must not reopen a request that was just settled, or the server would go back to refusing to boot right after being told not to. A successful attempt is always recorded. It returns the result now on disk.

func (*Store) Remove

func (s *Store) Remove(id string) error

Remove deletes a record. It exists for one case: a record that was created but whose executor never started and whose failure could not be written, which would otherwise hold the busy slot forever.

func (*Store) Update

func (s *Store) Update(op *Operation) error

func (*Store) WriteRestoreResult

func (s *Store) WriteRestoreResult(result *RestoreResult) error

WriteRestoreResult records the server's answer to a DataRestore request.

type SystemdLauncher

type SystemdLauncher struct {
	Binary string
	// Command is the subcommand that runs the executor, without the binary
	// and without the operation flag. Nil means ServerExecutorCommand.
	Command []string
}

SystemdLauncher runs the executor as a transient unit, independent of the daemon's own unit, so restarting the daemon does not take it down and journald keeps its output. The binary is captured by inode at exec, so the executor keeps running the build it started with after replacing the file on disk.

func (SystemdLauncher) Launch

func (l SystemdLauncher) Launch(ctx context.Context, opID string) error

func (SystemdLauncher) Launched

func (l SystemdLauncher) Launched(ctx context.Context, opID string) bool

Launched reports whether the operation's unit is running, for a Launch that errored after the unit may have been submitted.

type SystemdRestarter

type SystemdRestarter struct {
	Unit string
	// StateDir is the daemon's state directory; the resource-limit drop-in
	// refreshed before each restart records stop reasons there.
	StateDir string
}

func (SystemdRestarter) Restart

func (r SystemdRestarter) Restart(ctx context.Context) error

type UnsupervisedLauncher

type UnsupervisedLauncher struct{}

UnsupervisedLauncher refuses every launch with ErrUnsupervised. It stands in where neither systemd nor container-boot is present, so the refusal says why instead of a failed systemd-run.

func (UnsupervisedLauncher) Launch

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL