coordinate

package
v0.16.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 29, 2026 License: Apache-2.0 Imports: 122 Imported by: 0

Documentation

Index

Constants

View Source
const (
	DefaultProjectOwner = "miren.system@miren.dev"
	DefaultCloudURL     = "https://api.miren.cloud"
)
View Source
const APIServerName = "api.miren"

APIServerName is the stable name in-cluster clients verify the API certificate against.

A sandbox reaches the API through its bridge router address, which cannot be a certificate SAN: the subnet is leased after the certificate has been issued, and on a first boot there is no prior lease to anticipate. Clients therefore dial by address and pass this as the TLS server name. Nothing resolves it — no DNS record backs it, and none is needed — but it lets a sandbox verify the certificate against the cluster CA rather than skipping verification, which matters on a bridge shared with other sandboxes.

View Source
const ServerLifecycleService = "dev.miren.runtime/server-lifecycle"

ServerLifecycleService is the RPC service name for ServerLifecycle. Mirrored in cli/commands, which cannot import this linux-only package.

Variables

View Source
var ErrRunnerUnmanaged = errors.New("runner does not serve a lifecycle ledger")

ErrRunnerUnmanaged: the runner answers RPC but does not serve a lifecycle ledger, so it predates managed upgrades.

Functions

func EnsureCA added in v0.12.1

func EnsureCA(log *slog.Logger, dataPath string) (*caauth.Authority, error)

EnsureCA makes sure the server CA certificate and key exist on disk under <dataPath>/server, creating them if they're missing, and returns the loaded authority. It's safe to call before the coordinator starts.

This is what lets the server bootstrap its own etcd mTLS: SetupEtcdTLS reads the CA from disk and requires it to already exist, so the CA must be materialized before that runs. Historically the only create-if-missing path was Coordinator.LoadCA, which doesn't run until Coordinator.Start, well after the etcd TLS block, so a fresh install with distributed runners enabled had no CA yet. Calling EnsureCA early closes that gap.

func EtcdCertsDir added in v0.16.0

func EtcdCertsDir(dataPath string) string

SetupEtcdTLS loads the existing CA and ensures valid etcd mTLS certificates. Existing certificates are reused if their SANs match and they aren't near expiry; otherwise they are regenerated. The dataPath should be the same path used for CoordinatorConfig.DataPath. The CA must already exist (created by the coordinator's LoadCA). Additional DNS names and IPs are included in the server certificate SANs so that distributed runners can connect to etcd over the network. EtcdCertsDir is where SetupEtcdTLS keeps the etcd CA and server cert, for callers outside the server process that need to dial etcd.

Types

type AdvertiseCandidate added in v0.8.0

type AdvertiseCandidate struct {
	Source         string // "listen", "explicit", "discovered", "netcheck"
	HostPort       string
	IP             net.IP
	Interface      string // discovering interface, when known
	Classification string // tailnet / container-bridge / loopback / link-local / private / global-unicast / other
	Included       bool
	Reason         string
}

AdvertiseCandidate describes one candidate address the advertise logic considered, and whether it ended up in the final advertised set. Used by both production (building the final list) and debug tooling (explaining the decision for every IP).

func ComputeAdvertise added in v0.8.0

func ComputeAdvertise(in AdvertiseInput) ([]AdvertiseCandidate, []string)

ComputeAdvertise is the single source of truth for computing the addresses the server advertises. It returns the ordered list of candidates (including rejected ones, so callers can explain why) and the final list of advertised host:port strings.

The returned list is intended for StatusReport.APIAddresses, i.e. the addresses miren.cloud hands out to clients that want to reach this cluster. Loopback and unspecified (0.0.0.0, ::) addresses are never included — a client coming in through miren.cloud is by definition not running on the same host, so those entries would only produce failed connection attempts.

Filtering rules:

  1. Listen address: included if it parses as host:port with a literal, non-loopback, non-unspecified IP.

  2. Explicit IPs (user-configured): always included, except loopback and unspecified which are dropped with a reason.

  3. Discovered IPs (auto-scanned from interfaces): a. Loopback and unspecified are dropped. b. Addresses on a container bridge (docker0, flannel.1, rt0, …) are dropped — they exist only for workloads on this host. c. Addresses the internet can't route to (LAN, CGNAT, ULA, and so any overlay network) are kept, since they may be how this client reaches us. d. Internet-routable IPs are dropped if netcheck ran for that address family and proved the family unreachable or found reachable addresses (replaced by netcheck-confirmed ones). e. Otherwise kept as a fallback.

  4. Netcheck public addresses: included when reachable on at least one port.

type AdvertiseInput added in v0.8.0

type AdvertiseInput struct {
	// ListenAddr is the server's own listen address (e.g. "0.0.0.0:8443").
	// Included in the advertised list only if it has a literal,
	// non-loopback, non-unspecified IP.
	ListenAddr string

	// IPs is the unified list of all candidate IP addresses, each
	// tagged as explicit (user-configured) or discovered (interface scan).
	// Explicit IPs bypass all filtering except loopback / unspecified.
	// Discovered IPs are subject to container-bridge filtering, netcheck,
	// and other pruning.
	IPs []SourcedIP

	// Netcheck is the result of the dual-stack netcheck, if one has run.
	// A nil pointer means netcheck never ran / failed entirely.
	Netcheck *cloudauth.NetcheckDualStackResult

	// Port is the port to append to bare IPs (defaults to 8443).
	Port int
}

AdvertiseInput is the raw input for computing the set of API addresses the server should advertise to clients and to miren.cloud.

type ApplicationManagement added in v0.15.0

type ApplicationManagement struct {
	*Foundation
	// contains filtered or unexported fields
}

ApplicationManagement owns the durable app, addon, and run management surface. WorkloadControl realizes that intent as running workloads.

func NewApplicationManagement added in v0.15.0

func NewApplicationManagement(foundation *Foundation, secrets *SecretStore, diagnostics ...*entitysync.Diagnostics) *ApplicationManagement

NewApplicationManagement constructs the application management plane on top of cluster state and the secret store.

func (*ApplicationManagement) ExposeBuildAndDeployment added in v0.15.0

func (c *ApplicationManagement) ExposeBuildAndDeployment() error

func (*ApplicationManagement) RecoverBuildSagas added in v0.15.0

func (c *ApplicationManagement) RecoverBuildSagas(ctx context.Context)

func (*ApplicationManagement) Start added in v0.15.0

Start initializes the durable application-management objects and exposes their RPC handlers. Controllers that realize this intent belong to WorkloadControl.

func (*ApplicationManagement) Stop added in v0.15.0

func (c *ApplicationManagement) Stop()

Stop tears down the management RPC resources.

type CloudAuthConfig

type CloudAuthConfig struct {
	Enabled     bool              `json:"enabled" yaml:"enabled"`
	CloudURL    string            `json:"cloud_url" yaml:"cloud_url"`       // URL of miren.cloud (default: https://api.miren.cloud)
	PrivateKey  string            `json:"private_key" yaml:"private_key"`   // Required: Path to service account private key when enabled
	Tags        map[string]string `json:"tags" yaml:"tags"`                 // Tags from registration for RBAC evaluation
	ClusterID   string            `json:"cluster_id" yaml:"cluster_id"`     // Cluster ID for status reporting
	DNSHostname string            `json:"dns_hostname" yaml:"dns_hostname"` // Cloud-provisioned DNS hostname for the cluster

	// IdentityIssuerURL is the workload identity anchor cloud assigned this
	// cluster, when it has one. Empty means the cluster anchors identity at its
	// own hostname and serves its own discovery.
	IdentityIssuerURL string `json:"identity_issuer_url" yaml:"identity_issuer_url"`
}

CloudAuthConfig contains cloud authentication settings

type CloudControl added in v0.15.0

type CloudControl struct {
	*Foundation

	// Instance is optional; when nil reports omit the instance id.
	Instance *serverinfo.Source
	// Lifecycle, when set, lets cloud restart and upgrade this server over the
	// uplink.
	Lifecycle *ServerLifecycle
	// contains filtered or unexported fields
}

CloudControl owns cloud status, identity publication, and uplink integration.

func NewCloudControl added in v0.15.0

func NewCloudControl(foundation *Foundation, applications *ApplicationManagement, diagnostics ...*entitysync.Diagnostics) *CloudControl

NewCloudControl constructs cloud status, identity publication, and uplink integration on top of cluster state.

func (*CloudControl) EntitySyncDiagnostics added in v0.15.0

func (c *CloudControl) EntitySyncDiagnostics() *entitysync.Diagnostics

EntitySyncDiagnostics is the runtime-local view exposed by miren debug.

func (*CloudControl) NetworkFacts added in v0.16.0

func (c *CloudControl) NetworkFacts(ctx context.Context) clusternetwork.Report

NetworkFacts is what this cluster says about how it can be reached, in the shape both the status poll and the cluster-network capability send. It refreshes netcheck when the cached verdict is older than netcheckMaxAge (or was never taken), so whichever path asks first pays for the check and the other reads the cache.

func (*CloudControl) ReportStartupStatus added in v0.15.0

func (c *CloudControl) ReportStartupStatus(ctx context.Context) error

ReportStatus reports the current cluster status to miren.cloud

func (*CloudControl) ReportStatus added in v0.15.0

func (c *CloudControl) ReportStatus(ctx context.Context) error

ReportStatus sends the legacy status poll to miren.cloud. It is what a cloud that did not negotiate the cluster-network and cluster-resources capabilities still relies on; against one that did, the periodic loop skips it (see pollSuppressed).

func (*CloudControl) ResourceSample added in v0.16.0

func (c *CloudControl) ResourceSample() clusterresources.Sample

ResourceSample is one host reading in the cluster-resources wire shape, stamped with when it was taken.

The percentages are the ones the status poll carries. The totals are not: the poll's cpu_cores is a load average despite its name and its byte figures are used bytes, which cloud never read. The wire calls these capacities, so they come from the host's core count and total memory and storage.

func (c *CloudControl) RunCloudUplink(ctx context.Context, ingress *httpingress.Server, entitySyncReady <-chan struct{}) error

RunCloudUplink owns the shared cloud connection and all of its tenants. Entity sync alone waits for its source preparation; Anywhere and cloud RPC remain independent of that background migration.

func (*CloudControl) Start added in v0.15.0

func (c *CloudControl) Start(ctx context.Context) error

Start begins cloud-facing reporting. The uplink itself remains a separate long-running boot task because it owns the connection until shutdown.

func (*CloudControl) Stop added in v0.15.0

func (c *CloudControl) Stop()

Stop cancels and joins the background reporting loop started by Start.

type ControlPlane added in v0.15.0

type ControlPlane struct {
	*Foundation
	// contains filtered or unexported fields
}

ControlPlane is the reconstituted cluster-control handle. It is not a server boot phase; the graph owns and schedules each part above.

func NewControlPlane added in v0.15.0

func NewControlPlane(foundation *Foundation, supplied ...ControlPlaneParts) *ControlPlane

NewControlPlane reconstitutes the public cluster-control handle from independently owned parts. Direct callers may omit parts and let Start compose them; the server boot graph passes the exact instances it owns.

func (*ControlPlane) Start added in v0.15.0

func (c *ControlPlane) Start(ctx context.Context) (retErr error)

Start is the compatibility path for tests and embedded callers that need the cluster control plane in one call. Server-only migrations, build and deployment admission and recovery, the local runner, and ingress start separately in the server boot graph.

func (*ControlPlane) Stop added in v0.15.0

func (c *ControlPlane) Stop(ctx context.Context) error

Stop tears down the control plane and then its foundation for callers that construct both directly instead of using the server boot graph.

func (*ControlPlane) WorkloadControl added in v0.15.0

func (c *ControlPlane) WorkloadControl() *WorkloadControl

type ControlPlaneParts added in v0.15.0

type ControlPlaneParts struct {
	Secrets         *SecretStore
	RunnerEndpoints *RunnerEndpoints
	Workloads       *WorkloadControl
	Applications    *ApplicationManagement
	Maintenance     *EntityMaintenance
	Cloud           *CloudControl
}

type CoordinatorConfig

type CoordinatorConfig struct {
	Address         string              `json:"address" yaml:"address"`
	EtcdEndpoints   []string            `json:"etcd_endpoints" yaml:"etcd_endpoints"`
	Prefix          string              `json:"prefix" yaml:"prefix"`
	Resolver        netresolve.Resolver `json:"resolver" yaml:"resolver"`
	TempDir         string              `json:"temp_dir" yaml:"temp_dir"`
	DataPath        string              `json:"data_path" yaml:"data_path"`
	AdditionalNames []string            `json:"additional_names" yaml:"additional_names"`
	IPs             *IPSet              `json:"ips" yaml:"ips"`

	// ACME certificate configuration
	AcmeEmail       string `json:"acme_email" yaml:"acme_email"`
	AcmeDNSProvider string `json:"acme_dns_provider" yaml:"acme_dns_provider"`

	// Cloud authentication configuration
	CloudAuth CloudAuthConfig `json:"cloud_auth" yaml:"cloud_auth"`

	// NoAuth disables authentication entirely (for testing only)
	NoAuth bool `json:"no_auth" yaml:"no_auth"`

	// EtcdTLS holds mTLS configuration for etcd connections (optional).
	// When set, the coordinator will use mTLS to connect to etcd.
	EtcdTLS *EtcdTLSConfig `json:"etcd_tls" yaml:"etcd_tls"`

	Mem           *metrics.MemoryUsage
	Cpu           *metrics.CPUUsage
	HTTP          *metrics.HTTPMetrics
	MetricsWriter metrics.PointWriter

	// MetricsReader queries the cluster's metrics directly, rather than through
	// the app-shaped helpers on Cpu and Mem. The usage service needs arbitrary
	// groupings -- by sandbox, by node -- that those helpers do not express.
	MetricsReader *metrics.VictoriaMetricsReader

	Logs      *observability.LogReader
	LogWriter observability.LogWriter

	// Observability addresses for distributed runners
	VictoriametricsAddress string
	VictorialogsAddress    string

	// BuildKit is the persistent BuildKit component for container image builds
	BuildKit *buildkit.Component

	// AppVersionRetentionCount and AppVersionRetentionPeriod tune the version
	// retention GC. Values <= 0 fall back to the controller defaults.
	AppVersionRetentionCount  int
	AppVersionRetentionPeriod time.Duration

	// SecretKeyRotationPeriod is how old the cluster key may get before it
	// rotates on its own. Zero means rotate only when asked; negative means the
	// operator's value did not parse, so fall back to the default rather than
	// reading a typo as "never rotate".
	SecretKeyRotationPeriod time.Duration

	// SagaRetentionPeriod is how long a finished saga execution is kept.
	// Unlike the app-version knobs above, zero is meaningful here: it keeps
	// executions indefinitely, which is the escape hatch for an operator who
	// wants saga history frozen during an investigation. A negative value
	// falls back to the controller default.
	SagaRetentionPeriod time.Duration

	// DeploymentRetentionCount and DeploymentRetentionPeriod tune the
	// deployment record retention GC. The count follows the app-version
	// convention (<= 0 means the controller default); the period follows the
	// saga one, where zero keeps records indefinitely and a negative value
	// falls back to the default.
	DeploymentRetentionCount  int
	DeploymentRetentionPeriod time.Duration

	// WorkloadIssuer signs workload identity tokens for sandbox containers
	WorkloadIssuer *workloadidentity.Issuer

	// Secrets holds the registered secret backends. The caller builds it so the
	// runner sharing this process materializes through the same registry the
	// coordinator pins and serves with. Nil leaves the cluster without a secret
	// store, in which case a config referencing one fails rather than deploying
	// without it.
	Secrets *secret.Registry
}

type EntityMaintenance added in v0.15.0

type EntityMaintenance struct {
	*Foundation

	// DeploymentHistoryReady gates the deployment GC's first sweep on the
	// deployment-attempt migration, which must finish before any record is
	// pruned. Nil starts the sweep immediately.
	DeploymentHistoryReady <-chan struct{}

	// DeploymentExports lets the deployment GC hold a record until cloud has
	// landed it. Nil prunes on retention alone.
	DeploymentExports deploymentgcctrl.ExportProgress
	// contains filtered or unexported fields
}

EntityMaintenance owns background repair and garbage collection for durable cluster state. It does not reconcile live workloads.

func NewEntityMaintenance added in v0.15.0

func NewEntityMaintenance(foundation *Foundation) *EntityMaintenance

NewEntityMaintenance constructs the background repair and garbage-collection loops for cluster-owned entity state.

func (*EntityMaintenance) Start added in v0.15.0

func (c *EntityMaintenance) Start(ctx context.Context) error

func (*EntityMaintenance) Stop added in v0.15.0

func (c *EntityMaintenance) Stop()

type EtcdTLSConfig added in v0.4.0

type EtcdTLSConfig struct {
	CertPEM []byte // Client certificate PEM
	KeyPEM  []byte // Client private key PEM
	CACert  []byte // CA certificate PEM for verifying server
}

EtcdTLSConfig holds TLS configuration for connecting to etcd with mTLS.

type EtcdTLSSetupResult added in v0.4.0

type EtcdTLSSetupResult struct {
	// CertsDir is the directory containing etcd server certs (ca.crt, server.crt, server.key)
	CertsDir string
	// ClientTLS is the TLS config for clients connecting to etcd
	ClientTLS *EtcdTLSConfig
	// ClientCertFile is the path to the client certificate on disk
	ClientCertFile string
	// ClientKeyFile is the path to the client private key on disk
	ClientKeyFile string
	// CAFile is the path to the CA certificate on disk
	CAFile string
}

EtcdTLSSetupResult contains the results of setting up etcd TLS.

func SetupEtcdTLS added in v0.4.0

func SetupEtcdTLS(log *slog.Logger, dataPath string, extraDNSNames []string, extraIPs []net.IP) (*EtcdTLSSetupResult, error)

type Foundation added in v0.15.0

type Foundation struct {
	CoordinatorConfig

	Log *slog.Logger
	// contains filtered or unexported fields
}

Foundation owns the cluster authority, RPC state, and entity API. It is a real early readiness boundary: consumers can use it without access to the later control-plane controllers and handlers.

func NewFoundation added in v0.15.0

func NewFoundation(log *slog.Logger, cfg CoordinatorConfig) *Foundation

NewFoundation constructs the cluster authority and API foundation shared by the control plane and local runner-facing boot components.

func (*Foundation) BackfillCloudExportMarker added in v0.15.0

func (c *Foundation) BackfillCloudExportMarker(ctx context.Context) (entityexport.BackfillStats, error)

BackfillCloudExportMarker is the contract-wide safety net for exported kinds that no specialized migration covers. The deployment-attempt controller has already marked the apps, app versions, and deployments in its clean sweep; excluding deployments here preserves its ownership of validating old shapes.

func (*Foundation) CACertificate added in v0.15.0

func (c *Foundation) CACertificate() []byte

CACertificate returns the cluster CA in PEM form, or nil before LoadCA has run. Callers hand it to clients that must verify the cluster's certificates — sandboxes reaching the API with a workload identity token, for one.

func (*Foundation) IssueCertificate added in v0.15.0

func (c *Foundation) IssueCertificate(name string) (*caauth.ClientCertificate, error)

func (*Foundation) ListenAddress added in v0.15.0

func (c *Foundation) ListenAddress() string

func (*Foundation) LoadAPICert added in v0.15.0

func (c *Foundation) LoadAPICert(ctx context.Context) error

func (*Foundation) LoadCA added in v0.15.0

func (c *Foundation) LoadCA(ctx context.Context) error

func (*Foundation) LocalConfig added in v0.15.0

func (c *Foundation) LocalConfig() (*clientconfig.Config, error)

func (*Foundation) NamedConfig added in v0.15.0

func (c *Foundation) NamedConfig(name string) (*clientconfig.Config, error)

func (*Foundation) NewDeploymentAttemptController added in v0.15.0

func (c *Foundation) NewDeploymentAttemptController() (*deploymentattemptsctrl.Controller, error)

NewDeploymentAttemptController returns the migration controller after the coordinator has opened its entity store. The caller owns its lifecycle.

func (*Foundation) PrepareAppData added in v0.15.0

func (c *Foundation) PrepareAppData(ctx context.Context) error

PrepareAppData finishes the startup-time AppVersion rewrite that must not race consumers which update the same entities. Failures remain non-fatal, as they were when this work lived inside Start.

func (*Foundation) PublicIPs added in v0.15.0

func (c *Foundation) PublicIPs() []net.IP

PublicIPs returns the cluster's known public IP addresses, applying the same filtering rules as the advertised API addresses. Routes through ComputeAdvertise so the AutocertController's DNS sanity check honors per-family netcheck state (no leaking the source IP when its family has zero reachable ports) and the CGNAT filter (no advertising tailnet addresses as "public").

func (*Foundation) RunnerConfig added in v0.15.0

func (c *Foundation) RunnerConfig(listenAddress string) (*clientconfig.Config, error)

RunnerConfig returns a client config for a runner service with proper TLS certificate SANs. The certificate will be valid for localhost and the runner's listen address.

func (*Foundation) Server added in v0.15.0

func (c *Foundation) Server() *rpc.Server

func (*Foundation) ServiceConfig added in v0.15.0

func (c *Foundation) ServiceConfig() (*clientconfig.Config, error)

func (*Foundation) Start added in v0.15.0

func (c *Foundation) Start(ctx context.Context) (retErr error)

Start prepares the cluster CA, RPC state, entity store, and entity API.

func (*Foundation) Stop added in v0.15.0

func (c *Foundation) Stop(ctx context.Context) error

Stop drains RPC after all dependent components stop, then releases etcd.

type IPSet added in v0.8.0

type IPSet struct {
	// contains filtered or unexported fields
}

IPSet is an ordered, de-duplicated collection of SourcedIP entries. When a duplicate IP is added, the Explicit flag is sticky: adding an IP as explicit promotes a previously-discovered entry, but adding it as discovered never demotes an explicit one. Iteration order matches first-insertion order.

func NewIPSet added in v0.8.0

func NewIPSet() *IPSet

NewIPSet creates an empty IPSet.

func (*IPSet) Add added in v0.8.0

func (s *IPSet) Add(sip SourcedIP)

Add inserts an IP. If the IP already exists and the new entry is explicit, it promotes the existing entry. Discovered duplicates are silently ignored.

func (*IPSet) AddDiscovered added in v0.8.0

func (s *IPSet) AddDiscovered(ip net.IP)

AddDiscovered is a convenience for Add(SourcedIP{IP: ip, Explicit: false}).

func (*IPSet) AddDiscoveredFrom added in v0.13.0

func (s *IPSet) AddDiscoveredFrom(ip net.IP, iface string)

AddDiscoveredFrom records a discovered IP along with the interface it was found on, so bridge filtering can tell a real NIC from a container bridge.

func (*IPSet) AddExplicit added in v0.8.0

func (s *IPSet) AddExplicit(ip net.IP)

AddExplicit is a convenience for Add(SourcedIP{IP: ip, Explicit: true}).

func (*IPSet) All added in v0.8.0

func (s *IPSet) All() []SourcedIP

All returns the entries in insertion order. The returned slice is a copy — callers may not modify it. Safe to call on a nil receiver.

func (*IPSet) Len added in v0.8.0

func (s *IPSet) Len() int

Len returns the number of unique IPs in the set. Safe to call on a nil receiver.

func (*IPSet) RawIPs added in v0.8.0

func (s *IPSet) RawIPs() []net.IP

RawIPs extracts just the net.IP values in insertion order. Safe to call on a nil receiver.

type NodeInventory added in v0.16.0

type NodeInventory interface {
	// Runners lists every node that is not the coordinator, by name.
	Runners(ctx context.Context) ([]*compute_v1alpha.Node, error)
	// Runner finds one node by runner id; nil when it is gone.
	Runner(ctx context.Context, runnerID string) (*compute_v1alpha.Node, error)
	SetScheduling(ctx context.Context, node *compute_v1alpha.Node, scheduling entity.Id) error
}

NodeInventory is the walker's view of the cluster's nodes.

type ResourceUsage added in v0.15.0

type ResourceUsage struct {
	*Foundation
}

ResourceUsage answers what the cluster is spending its CPU and memory on.

This lives only on the coordinator because it is the only process holding both the entity store and a reader for every node's metrics; a runner can answer for itself but not for the cluster.

func NewResourceUsage added in v0.15.0

func NewResourceUsage(foundation *Foundation) *ResourceUsage

NewResourceUsage constructs the cluster-wide usage service on top of an already-created foundation.

func (*ResourceUsage) Start added in v0.15.0

func (c *ResourceUsage) Start(ctx context.Context) error

Start exposes usage queries to clients. A nil MetricsReader is not an error: the service still answers "what is running where" from the entity store, and reports the figures it cannot measure as warnings.

type Resumer added in v0.16.0

type Resumer interface {
	Resume(ctx context.Context, store *serverlifecycle.Store) error
}

Resumer is a Launcher whose executors do not outlive the server, so an operation the previous instance left unfinished has to be picked up here.

type RunnerDialer added in v0.16.0

type RunnerDialer func(ctx context.Context, address string) (RunnerLifecycleClient, error)

RunnerDialer reaches a runner's ledger at its API address.

type RunnerEndpoints added in v0.15.0

type RunnerEndpoints struct {
	*Foundation
}

RunnerEndpoints owns coordinator endpoints needed by runners before workload reconciliation begins.

func NewRunnerEndpoints added in v0.15.0

func NewRunnerEndpoints(foundation *Foundation) *RunnerEndpoints

NewRunnerEndpoints constructs the coordinator endpoints a runner needs in order to join and preserve local state before the rest of the control plane boots.

func (*RunnerEndpoints) Start added in v0.15.0

Start exposes the coordinator capabilities a runner needs to initialize. Keeping these ahead of workload control lets a restarting host recover disks and rejoin without depending on unrelated controllers.

func (*RunnerEndpoints) Stop added in v0.15.0

func (*RunnerEndpoints) Stop()

Stop intentionally does nothing. RunnerEndpoints only registers handlers on the foundation-owned RPC server; declaring the method prevents Foundation.Stop from being promoted as this component's lifecycle.

type RunnerLifecycleClient added in v0.16.0

type RunnerLifecycleClient interface {
	Start(ctx context.Context, op *serverlifecycle.Operation) (*serverlifecycle.Operation, error)
	Get(ctx context.Context, id string) (*serverlifecycle.Operation, error)
	Close()
}

RunnerLifecycleClient is one runner's ledger as the walker drives it.

type RunnerUpgradeOptions added in v0.16.0

type RunnerUpgradeOptions struct {
	// StepTimeout bounds one runner's operation, download included. The
	// runner's own ready timeout is inside it.
	StepTimeout time.Duration
	// ReadyTimeout is how long a runner that finished its operation gets to
	// come back READY on the new build in the node inventory.
	ReadyTimeout time.Duration
	// NotReadyGrace is how long a runner that is not READY when its turn
	// comes gets to become so before it is skipped. Right after the server
	// restart that precedes the walk, every runner is re-establishing its
	// session, and that must not read as a fleet of dead runners.
	NotReadyGrace time.Duration
	PollInterval  time.Duration
	// AdoptRetry is how long to keep trying for the operation lock while the
	// executor that handed off is still letting go of it.
	AdoptRetry time.Duration
	// RetryFor bounds how long a walk keeps retrying after an error that is
	// not a step outcome (the inventory or the ledger not answering) before
	// the operation is failed for good. RetryBackoff is the first wait
	// between attempts; it doubles up to RetryBackoffMax.
	RetryFor        time.Duration
	RetryBackoff    time.Duration
	RetryBackoffMax time.Duration
}

RunnerUpgradeOptions tunes the walk.

func DefaultRunnerUpgradeOptions added in v0.16.0

func DefaultRunnerUpgradeOptions() RunnerUpgradeOptions

type RunnerUpgrader added in v0.16.0

type RunnerUpgrader struct {
	// contains filtered or unexported fields
}

RunnerUpgrader finishes upgrade operations the executor handed off. The executor stops once the server it restarted reports ready; the server is then the one with the node inventory, the RPC to every runner, and the build that knows the runner protocol. It walks the runners one at a time: a bad build shows up on the first runner and stops there, and each step is gated on the runner coming back READY on the new build before the next begins. Steps are checkpointed on the record, so a server that restarts mid-walk resumes where it was.

func NewRunnerUpgrader added in v0.16.0

func NewRunnerUpgrader(log *slog.Logger, store *serverlifecycle.Store, watcher *lifecyclesync.Watcher, instanceID string, nodes NodeInventory, dial RunnerDialer, opts RunnerUpgradeOptions) *RunnerUpgrader

func (*RunnerUpgrader) Run added in v0.16.0

func (u *RunnerUpgrader) Run(ctx context.Context) error

Run adopts every handed-off operation in the ledger, now and as they appear, until ctx ends. Subscribing before the initial scan is what keeps the two from missing one between them.

type SecretStore added in v0.15.0

type SecretStore struct {
	*Foundation
	// contains filtered or unexported fields
}

SecretStore owns the cluster keyring, secret backend, rotation loop, and runner-facing secret endpoint.

func NewSecretStore added in v0.15.0

func NewSecretStore(foundation *Foundation) *SecretStore

NewSecretStore constructs the cluster-owned secret backend and rotation controller on top of an already-created foundation.

func (*SecretStore) Registry added in v0.15.0

func (c *SecretStore) Registry() *secret.Registry

func (*SecretStore) Start added in v0.15.0

func (c *SecretStore) Start(ctx context.Context) error

Start opens the cluster secret backend, starts rotation, and exposes secret resolution to distributed runners.

func (*SecretStore) Stop added in v0.15.0

func (c *SecretStore) Stop()

type ServerLifecycle added in v0.16.0

type ServerLifecycle struct {
	*Foundation
	// contains filtered or unexported fields
}

ServerLifecycle owns the read side of pkg/serverlifecycle within the server. Writes still go through the file store; every reader here follows it.

func NewServerLifecycle added in v0.16.0

func NewServerLifecycle(foundation *Foundation, instance *serverinfo.Source, dir string, launcher serverlifecycle.Launcher) *ServerLifecycle

NewServerLifecycle serves the restart and upgrade ledger over RPC and to cloud, and finishes upgrades across the runners. dir is the ledger directory. launcher is where the executor that writes the server's own phases runs: outside this process under systemd, inside it in a container.

func (*ServerLifecycle) Register added in v0.16.0

func (c *ServerLifecycle) Register(ctx context.Context, link lifecyclesync.Link) error

Register adds the cloud capability to link.

func (*ServerLifecycle) Start added in v0.16.0

func (c *ServerLifecycle) Start(ctx context.Context) error

func (*ServerLifecycle) Stop added in v0.16.0

func (c *ServerLifecycle) Stop()

type SourcedIP added in v0.8.0

type SourcedIP struct {
	IP       net.IP
	Explicit bool // true = user-configured, false = auto-discovered

	// Interface is the name of the link the IP was discovered on, when
	// known. Empty for explicit IPs and for callers that don't track it.
	// Used to tell a host's real NICs apart from the container bridges
	// Miren and Docker create, which are never reachable from a client.
	Interface string
}

SourcedIP is an IP address tagged with how it was obtained. Explicit IPs (user-configured via AdditionalIPs or the server config) always pass through to the advertised list. Discovered IPs (auto-scanned from local interfaces) are subject to netcheck pruning, bridge filtering, etc.

type WorkloadControl added in v0.15.0

type WorkloadControl struct {
	*Foundation
	// contains filtered or unexported fields
}

WorkloadControl owns workload realization and operational routing: app and addon reconciliation, one-shot runs, placement, exec routing, route defaults, and certificates. The HTTP listener itself remains data plane.

func NewWorkloadControl added in v0.15.0

func NewWorkloadControl(foundation *Foundation, applications *ApplicationManagement) *WorkloadControl

NewWorkloadControl constructs the controllers that turn management intent into running and routable sandbox capacity.

func (*WorkloadControl) Activator added in v0.15.0

func (c *WorkloadControl) Activator() activator.AppActivator

func (*WorkloadControl) AutocertReadySignal added in v0.15.0

func (c *WorkloadControl) AutocertReadySignal() func()

AutocertReadySignal returns a function that signals the autocert controller that the port-80 ACME challenge server is ready. Returns nil when the DNS-01 path is used (which doesn't need port 80).

func (*WorkloadControl) CertificateProvider added in v0.15.0

func (c *WorkloadControl) CertificateProvider() autotls.CertificateProvider

CertificateProvider returns the certificate provider for use by autotls.

func (*WorkloadControl) SandboxPoolManager added in v0.15.0

func (c *WorkloadControl) SandboxPoolManager() *sandboxpool.Manager

func (*WorkloadControl) SetCertificateHostChecker added in v0.16.0

func (c *WorkloadControl) SetCertificateHostChecker(fn certctrl.HostChecker)

SetCertificateHostChecker hands the autocert controller the function that vouches for on-demand names under a route. Without one, only names that routes name exactly get certificates. A no-op on the DNS-01 path.

func (*WorkloadControl) Start added in v0.15.0

func (c *WorkloadControl) Start(ctx context.Context) error

Start recovers the workload and routing view. The server boot graph waits for durable management state and the restored local sandbox host before calling it.

func (*WorkloadControl) Stop added in v0.15.0

func (c *WorkloadControl) Stop()

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL