sandbox

package
v0.16.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 29, 2026 License: Apache-2.0 Imports: 76 Imported by: 0

Documentation

Index

Constants

View Source
const (

	// SandboxEntityLabel is the container label key used to associate containers with sandbox entities.
	SandboxEntityLabel = "runtime.computer/entity-id"
)

Variables

This section is empty.

Functions

func PauseContainerID added in v0.4.0

func PauseContainerID(id entity.Id) string

PauseContainerID returns the containerd container ID for a sandbox's pause container.

Types

type BlobGCResult added in v0.4.0

type BlobGCResult struct {
	DeletedBlobs  []string
	FailedBlobs   map[string]error
	TotalBlobs    int
	RetainedBlobs int
}

BlobGCResult contains information about blobs cleaned up during GC.

type Cgroups

type Cgroups struct {
	// contains filtered or unexported fields
}

type CleanupResult

type CleanupResult struct {
	// DeletedContainers contains IDs of containers successfully removed
	DeletedContainers []string
	// FailedContainers contains IDs and errors for containers that failed to be removed
	FailedContainers map[string]error
}

CleanupResult contains information about containers cleaned up during orphan removal

type ContainerWatchdog

type ContainerWatchdog struct {
	Log *slog.Logger
	CC  *containerd.Client
	EAC *entityserver_v1alpha.EntityAccessClient

	Namespace string
	// NodeId scopes sandbox lookups to this node so we only consider
	// sandboxes that are scheduled here when building the valid set.
	NodeId compute.NodeId
	// CheckInterval is how often to check for orphaned containers
	CheckInterval time.Duration
	// GraceWindow is how long to wait before removing containers from non-running sandboxes
	GraceWindow time.Duration
	// Subnet is used to release IP addresses when removing orphaned containers
	Subnet *netdb.Subnet
	// contains filtered or unexported fields
}

ContainerWatchdog periodically checks that containers in containerd match what is expected by sandbox entities. It removes orphaned containers that shouldn't exist, acting as a safety mechanism to keep the container runtime clean.

func (*ContainerWatchdog) CleanupOrphanedContainers added in v0.4.0

func (w *ContainerWatchdog) CleanupOrphanedContainers(ctx context.Context) (*CleanupResult, error)

CleanupOrphanedContainers removes containers not associated with Running sandboxes. Returns a CleanupResult containing lists of successfully deleted and failed containers.

func (*ContainerWatchdog) Start

func (w *ContainerWatchdog) Start(ctx context.Context)

Start begins the periodic container cleanup process

func (*ContainerWatchdog) Stop

func (w *ContainerWatchdog) Stop()

Stop gracefully stops the watchdog

type Hub added in v0.14.0

type Hub struct {
	// contains filtered or unexported fields
}

Hub fans a container's output out to zero or more attached clients and funnels their input back in.

It exists so that attaching to a running container is a subscription rather than a new process. Executing a second process to serve an attach — which is what `miren app run` did before — means the attached command dies with the client that started it, and its exit code dies with it. A Hub decouples the two: the command is the container's own primary process, clients come and go around it, and nothing about their comings and goings reaches the workload.

The zero value is not usable; call NewHub.

func NewHub added in v0.14.0

func NewHub() *Hub

func (*Hub) Close added in v0.14.0

func (h *Hub) Close()

Close releases the Hub and closes the container's stdin. Only teardown may call it: closing stdin ends an interactive shell, so a client disconnecting must not.

func (*Hub) Done added in v0.14.0

func (h *Hub) Done() <-chan struct{}

Done is closed once the container is gone.

An attached client selects on this to learn that the workload ended. Without it the only event an attach can see is its own stdin reaching EOF, which is unrelated: stdin is /dev/null in any non-interactive caller, so an attach keyed on it returns immediately and reports a still-running task as finished.

func (*Hub) Dropped added in v0.14.0

func (h *Hub) Dropped() uint64

Dropped reports how many output chunks were discarded because attached clients could not keep up. Non-zero means some client saw a gap; the log stream still has everything.

func (*Hub) Resize added in v0.14.0

func (h *Hub) Resize(ctx context.Context, w, height uint32) error

Resize propagates a window size change to the container's terminal. It is a no-op before the task exists or on a container without a TTY.

func (*Hub) SetResizer added in v0.14.0

func (h *Hub) SetResizer(r Resizer)

SetResizer records the task whose terminal should follow attached clients' window sizes. It is called once the task exists, which is after the Hub is built and handed to containerd.

func (*Hub) Stdin added in v0.14.0

func (h *Hub) Stdin() io.Reader

Stdin returns the reader handed to containerd as the container's stdin.

It must never reach EOF while the container is alive. containerd's copyIO closes the container's stdin FIFO as soon as this reader returns EOF, which for an interactive shell means the shell exits. That is precisely the behavior a Hub exists to prevent, so detaching a client never closes this pipe — only Close does, at teardown.

func (*Hub) Subscribe added in v0.14.0

func (h *Hub) Subscribe(w io.Writer) (unsubscribe func())

Subscribe delivers container output to w until the returned function is called. Multiple clients may subscribe at once.

A new subscriber first receives the recent output the container has already produced, then everything written from here on. The replay is what makes a command that finishes before the client connects still print something; the snapshot and the registration happen under one lock, so no chunk is delivered twice or skipped between the two.

The returned function is idempotent and must be called to release the subscriber's goroutine.

func (*Hub) Write added in v0.14.0

func (h *Hub) Write(p []byte) (int, error)

Write fans one chunk of container output out to every attached client.

It is an io.Writer so it can sit in the same cio.WithStreams call as the log consumer. It never returns an error: it shares an io.MultiWriter with the log consumer, and a MultiWriter abandons the remaining writers on the first error, so a failing Hub would silently stop the container's output reaching the logs.

func (*Hub) WriteStdin added in v0.14.0

func (h *Hub) WriteStdin(p []byte) (int, error)

WriteStdin forwards input from an attached client to the container.

It blocks until the container reads, which is the correct backpressure: it is the calling client's own goroutine that waits, and a client typing faster than the workload reads should be slowed rather than have its keystrokes dropped.

type HubRegistry added in v0.14.0

type HubRegistry struct {
	// contains filtered or unexported fields
}

HubRegistry holds the live Hubs for a runner's containers, so the exec server can find the Hub for a sandbox the sandbox controller booted. Both live in the runner process.

func NewHubRegistry added in v0.14.0

func NewHubRegistry() *HubRegistry

func (*HubRegistry) Get added in v0.14.0

func (r *HubRegistry) Get(sandboxID entity.Id, container string) *Hub

Get returns the Hub for a container, or nil if the container is not attachable and never ran here.

A Hub whose container has already gone is still returned for a short window, so a client that arrives just after teardown gets the output rather than nothing. It is closed, so the attach ends as soon as it has replayed.

func (*HubRegistry) GetOrCreate added in v0.14.0

func (r *HubRegistry) GetOrCreate(sandboxID entity.Id, container string) *Hub

GetOrCreate returns the Hub for a container, creating it if the container is attachable. Callers hold the result for the life of the container.

func (*HubRegistry) Remove added in v0.14.0

func (r *HubRegistry) Remove(sandboxID entity.Id, container string)

Remove tears down a container's Hub, leaving it reachable for the linger window.

func (*HubRegistry) RemoveAll added in v0.14.0

func (r *HubRegistry) RemoveAll(sandboxID entity.Id)

RemoveAll tears down every Hub belonging to a sandbox, whatever its containers are called. Used when the sandbox itself goes away.

type IPReconciler added in v0.11.0

type IPReconciler struct {
	Log    *slog.Logger
	Subnet *netdb.Subnet

	// LiveIPs returns the set of addresses currently assigned to running sandbox
	// containers, determined independently of netdb.
	LiveIPs func(ctx context.Context) (map[netip.Addr]bool, error)

	// CheckInterval is how often a reconcile cycle runs.
	CheckInterval time.Duration
	// ReleaseAfterMisses is how many consecutive cycles an address must be
	// reserved-but-not-live before it is released. With the default interval this
	// is a multi-minute grace window, comfortably longer than any normal sandbox
	// create (whose saga releases the reservation on failure anyway).
	ReleaseAfterMisses int
	// contains filtered or unexported fields
}

IPReconciler keeps the netdb IP lease bookkeeping in agreement with the addresses actually live on the bridge. The two can diverge — a release path that ran while a container kept running, a lease lost across a restart — and a divergence where netdb believes a live address is free leads to that address being handed to a second sandbox, a duplicate assignment (MIR-1238).

Each cycle it:

  • re-reserves any address that is live on the bridge but not reserved in netdb (the repair direction — always safe, only ever adds a reservation for an address already in use); and
  • releases any address reserved in netdb with no live owner, but only after it has been absent for several consecutive cycles (the reap direction — conservative, to avoid racing a sandbox that is mid-create or mid-teardown).

func (*IPReconciler) Start added in v0.11.0

func (r *IPReconciler) Start(ctx context.Context)

Start begins the periodic reconcile loop.

func (*IPReconciler) Stop added in v0.11.0

func (r *IPReconciler) Stop()

Stop gracefully stops the reconciler.

type ImageGCConfig added in v0.3.0

type ImageGCConfig struct {
	// ScheduledGCInterval is how often to run scheduled GC regardless of pressure (default: 168h/weekly)
	ScheduledGCInterval time.Duration
	// PressureCheckInterval is how often to check disk pressure (default: 1h)
	PressureCheckInterval time.Duration
	// DiskPressureThreshold is the disk usage percentage that triggers immediate GC (default: 80%)
	DiskPressureThreshold float64
	// OrphanGracePeriod is how old a miren-managed image must be before it is
	// reclaimed when no Artifact entity exists for it (default: 24h). Zero or
	// negative means the default.
	OrphanGracePeriod time.Duration
}

ImageGCConfig holds configuration for the image garbage collector.

func DefaultImageGCConfig added in v0.3.0

func DefaultImageGCConfig() ImageGCConfig

DefaultImageGCConfig returns the default configuration for image GC.

type ImageGCResult added in v0.3.0

type ImageGCResult struct {
	// DeletedImages contains names of images successfully removed
	DeletedImages []string
	// FailedImages contains names and errors for images that failed to be removed
	FailedImages map[string]error
	// TotalImages is the total number of images before GC
	TotalImages int
	// RetainedImages is the number of images kept
	RetainedImages int
}

ImageGCResult contains information about images cleaned up during GC.

type ImageWatchdog added in v0.3.0

type ImageWatchdog struct {
	Log *slog.Logger
	CC  *containerd.Client
	EAC *entityserver_v1alpha.EntityAccessClient

	Namespace string
	DataPath  string
	Config    ImageGCConfig
	// contains filtered or unexported fields
}

ImageWatchdog periodically garbage collects container images from containerd. It uses Artifact entity status to determine which images to remove:

  • Images that are not miren-managed (infrastructure images, etc.) are kept
  • Images with Artifact status "active" or empty are kept
  • Images with Artifact status "archived" are deleted
  • Miren-managed images with no Artifact entity are deleted once they are older than OrphanGracePeriod. The Artifact is created when the manifest is pushed, before containerd ever pulls the image, so a missing Artifact means the app was deleted out from under it.
  • Images an AppVersion still references are kept regardless of what the Artifact says, since the version is what a deploy or rollback needs.

func (*ImageWatchdog) ParseArtifactID added in v0.4.0

func (w *ImageWatchdog) ParseArtifactID(imageName string) string

ParseArtifactID extracts the artifact ID from an image name. Image format: cluster.local:5000/{app}:{artifact-name} Artifact ID format: artifact/{artifact-name} Returns empty string if the image doesn't match the expected format.

func (*ImageWatchdog) RunBlobGC added in v0.4.0

func (w *ImageWatchdog) RunBlobGC(ctx context.Context) (*BlobGCResult, error)

RunBlobGC performs garbage collection of unreferenced registry blobs. It compares blob files on disk against digests referenced by non-archived artifacts or app versions and deletes any that are no longer needed.

func (*ImageWatchdog) RunGC added in v0.3.0

func (w *ImageWatchdog) RunGC(ctx context.Context) (*ImageGCResult, error)

RunGC performs garbage collection of unused images.

func (*ImageWatchdog) Start added in v0.3.0

func (w *ImageWatchdog) Start(ctx context.Context)

Start begins the periodic image cleanup process.

func (*ImageWatchdog) Stop added in v0.3.0

func (w *ImageWatchdog) Stop()

Stop gracefully stops the watchdog.

type Metrics

type Metrics struct {
	Log      *slog.Logger
	CPUUsage *metrics.CPUUsage
	MemUsage *metrics.MemoryUsage
	// contains filtered or unexported fields
}

func NewMetrics added in v0.3.0

func NewMetrics() *Metrics

NewMetrics creates a new Metrics.

func (*Metrics) Add

func (m *Metrics) Add(name string, pathes map[string]string, attributes map[string]string) error

func (*Metrics) AddIfAbsent added in v0.13.0

func (m *Metrics) AddIfAbsent(name string, pathes map[string]string, attributes map[string]string) (bool, error)

AddIfAbsent registers cgroups for name only when it isn't already monitored, reporting whether it did. The check and the insert happen under one lock, so two callers racing to re-register the same sandbox can't both win and reset the CPU counters Add would otherwise discard.

func (*Metrics) Gather

func (m *Metrics) Gather(name string) ([]*metric_v1alpha.ContainerSnapshot, error)

func (*Metrics) Has added in v0.13.0

func (m *Metrics) Has(name string) bool

Has reports whether a log entity is already being monitored. Re-registration paths that run on every reconcile check this before doing the work of collecting cgroup paths; AddIfAbsent is what actually makes the registration safe against a concurrent one.

func (*Metrics) Monitor

func (m *Metrics) Monitor(ctx context.Context)

func (*Metrics) Remove

func (m *Metrics) Remove(name string) error

func (*Metrics) Snapshot

func (*Metrics) Validate added in v0.7.0

func (m *Metrics) Validate() error

Validate checks that required fields are set so Monitor() won't nil-deref.

type PortMonitor

type PortMonitor struct {
	// contains filtered or unexported fields
}

PortMonitor monitors ports for containers using polling

func NewPortMonitor

func NewPortMonitor(log *slog.Logger, ports observability.PortTracker) *PortMonitor

NewPortMonitor creates a new port monitor

func (*PortMonitor) Close

func (pm *PortMonitor) Close() error

Close stops all monitoring

func (*PortMonitor) DiagnoseListening added in v0.11.0

func (pm *PortMonitor) DiagnoseListening(containerID string) (routable []int, loopback []int, ok bool)

DiagnoseListening reports the ports a container is listening on inside its netns, split into routable (reachable from the host) and loopback-only sets. ok is false when the container is not being monitored (its pid is unknown). It is used on the port-wait timeout path: when the configured port never bound, this reveals what the app actually listened on so we can route to it or explain the failure.

func (*PortMonitor) MonitorContainer

func (pm *PortMonitor) MonitorContainer(containerID string, ip string, pid int, ports []int)

MonitorContainer starts monitoring ports for a container. It checks port binding by reading /proc/<pid>/net/tcp from the container's network namespace (via the pause container's PID) rather than doing a TCP dial from the host, which can be interfered with by iptables DNAT rules.

func (*PortMonitor) StopMonitoring

func (pm *PortMonitor) StopMonitoring(containerID string)

StopMonitoring stops monitoring for a container

type Resizer added in v0.14.0

type Resizer interface {
	Resize(ctx context.Context, w, h uint32) error
}

Resizer is the part of containerd's task API a Hub needs to propagate window size changes. Narrowed to one method so tests don't need a real task.

type SandboxContainerRuntime added in v0.6.0

type SandboxContainerRuntime interface {
	BuildSpec(ctx context.Context, sb *compute.Sandbox, ep *network.EndpointConfig, meta *entity.Meta) ([]containerd.NewContainerOpts, error)
	CreateContainer(ctx context.Context, id string, opts ...containerd.NewContainerOpts) (string, error)
	LoadContainer(ctx context.Context, id string) (containerd.Container, error)
	// ContainerSpec fetches a loaded container's OCI spec. Callers must go
	// through here rather than calling cont.Spec directly: containerd resolves
	// the container out of the namespace on the context, and only this
	// implementation knows which namespace that is.
	ContainerSpec(ctx context.Context, cont containerd.Container) (*oci.Spec, error)
	CleanupContainer(ctx context.Context, cont containerd.Container)
	BootInitialTask(ctx context.Context, sb *compute.Sandbox, ep *network.EndpointConfig, container containerd.Container, shortID string) (containerd.Task, error)
	ConfigureVolumes(ctx context.Context, sb *compute.Sandbox, meta *entity.Meta) (map[string]string, error)
	BootContainers(ctx context.Context, sb *compute.Sandbox, ep *network.EndpointConfig, sbPid int, cgroups map[string]string, meta *entity.Meta, volumeMounts map[string]string) ([]WaitPort, error)
	DestroySubContainers(ctx context.Context, id entity.Id) error
	ReleaseDiskLeases(ctx context.Context, sandboxID entity.Id) error
	ReleaseHubs(id entity.Id)
	ReleaseSqliteDisks(ctx context.Context, sandboxID entity.Id)
	ReleaseTokenState(id entity.Id)
	UnconfigureFirewall(sb *compute.Sandbox)
	WaitForPort(ctx context.Context, id string, port int, timeout time.Duration) error
	// DiagnoseListening reports which ports a container is actually listening
	// on, split into routable (reachable from the host) and loopback-only sets.
	// Used on the port-wait timeout path to detect an app that bound a port
	// other than the one Miren configured. ok is false when the container is no
	// longer monitored and its pid is unknown.
	DiagnoseListening(id string) (routable []int, loopback []int, ok bool)
}

SandboxContainerRuntime provides containerd container operations.

type SandboxController

type SandboxController struct {
	Log *slog.Logger
	CC  *containerd.Client

	EAC *entityserver_v1alpha.EntityAccessClient

	Namespace string
	NodeId    compute.NodeId

	NetServ *network.ServiceManager

	Bridge string
	Subnet *netdb.Subnet

	DataPath string
	Tempdir  string

	LogsMaintainer *observability.LogsMaintainer
	LogWriter      observability.LogWriter

	StatusMon *observability.StatusMonitor

	Resolver       netresolve.Resolver
	Metrics        *Metrics
	WorkloadIssuer workloadidentity.TokenIssuer
	ApiAddress     string
	CACert         []byte

	// Secrets materializes the secret references a sandbox spec carries, at the
	// moment a container is created. Nil where no backend is reachable, in which
	// case a spec that references one fails rather than starting the container
	// with the reference in place of the value.
	Secrets secret.Resolver

	// SqliteDisks replicates sqlite-provider disks to the coordinator. Nil is
	// inert, so disks still mount when backups are unavailable.
	SqliteDisks *sqlitedisk.Manager
	// contains filtered or unexported fields
}

func NewSandboxController added in v0.3.0

func NewSandboxController(cfg SandboxControllerDeps, sagaStorage saga.Storage) (*SandboxController, error)

NewSandboxController creates a new SandboxController with validated dependencies. sagaStorage backs the create-sandbox saga; it is required, since sandbox creation has no non-saga implementation.

func (*SandboxController) AllocateNetwork added in v0.6.0

func (c *SandboxController) AllocateNetwork(
	ctx context.Context,
	co *compute.Sandbox,
) (*network.EndpointConfig, error)

AllocateNetwork gives a sandbox its bridge endpoint. A sandbox that already carries addresses is set up on those, so a re-run adopts what it had rather than leasing a second address; one that carries none is allocated a fresh address and has it recorded on the passed-in spec.

func (*SandboxController) BootContainers added in v0.6.0

func (c *SandboxController) BootContainers(
	ctx context.Context,
	sb *compute.Sandbox,
	ep *network.EndpointConfig,
	sbPid int,
	cgroups map[string]string,
	meta *entity.Meta,
	volumeMounts map[string]string,
) ([]WaitPort, error)

func (*SandboxController) BootInitialTask added in v0.6.0

func (c *SandboxController) BootInitialTask(
	ctx context.Context,
	sb *compute.Sandbox,
	ep *network.EndpointConfig,
	container containerd.Container,
	shortID string,
) (containerd.Task, error)

func (*SandboxController) BuildSpec added in v0.6.0

func (*SandboxController) CheckSandbox added in v0.6.0

func (c *SandboxController) CheckSandbox(ctx context.Context, co *compute.Sandbox, meta *entity.Meta) (int, error)

func (*SandboxController) CleanupContainer added in v0.6.0

func (c *SandboxController) CleanupContainer(ctx context.Context, cont containerd.Container)

CleanupContainer removes a container and its snapshot during failure scenarios

func (*SandboxController) Close

func (c *SandboxController) Close() error

func (*SandboxController) ConfigureVolumes added in v0.6.0

func (c *SandboxController) ConfigureVolumes(ctx context.Context, sb *compute.Sandbox, meta *entity.Meta) (map[string]string, error)

ConfigureVolumes prepares volumes and returns a map of volume name to actual mount path

func (*SandboxController) Create

func (c *SandboxController) Create(ctx context.Context, co *compute.Sandbox, meta *entity.Meta) error

Create handles sandbox create/update events, driving new sandboxes through the create-sandbox saga so a crash mid-create resumes or unwinds rather than stranding containers, addresses, and disk leases.

func (*SandboxController) Delete

func (c *SandboxController) Delete(ctx context.Context, id entity.Id, sb *compute.Sandbox) error

func (*SandboxController) DestroySubContainers added in v0.6.0

func (c *SandboxController) DestroySubContainers(ctx context.Context, id entity.Id) error

func (*SandboxController) EmitSandboxEvent added in v0.7.0

func (c *SandboxController) EmitSandboxEvent(sb *compute.Sandbox, shortID, line string)

EmitSandboxEvent writes a single runtime lifecycle line to the sandbox's log stream, using the same entity and attrs the container stdio pipeline uses. Events go on the Stderr stream so operators see them in `miren logs sandbox <id>` and can distinguish them from application output via the [miren] prefix.

func (*SandboxController) Hubs added in v0.14.0

func (c *SandboxController) Hubs() *HubRegistry

Hubs exposes the registry of attachable containers so the runner's exec server can join a client to a running container's terminal.

func (*SandboxController) Init

func (c *SandboxController) Init(ctx context.Context) error

func (*SandboxController) Periodic

func (c *SandboxController) Periodic(ctx context.Context, timeHorizon time.Duration) error

Periodic cleans up dead sandboxes that are older than the specified time horizon

func (*SandboxController) ReleaseDiskLeases added in v0.6.0

func (c *SandboxController) ReleaseDiskLeases(ctx context.Context, sandboxID entity.Id) error

ReleaseDiskLeases releases all disk leases owned by the given sandbox. This transitions leases to RELEASED status, which triggers the disk lease controller to unmount the volumes and release the underlying resources.

func (*SandboxController) ReleaseTokenState added in v0.13.0

func (c *SandboxController) ReleaseTokenState(id entity.Id)

ReleaseTokenState revokes every workload identity record for a sandbox, including its persisted token-request secret. Called from StopSandbox so the state goes away with the sandbox itself rather than with the entity, which outlives it by up to the periodic cleanup horizon, and from boot-failure paths that never reach StopSandbox.

func (*SandboxController) SetPortStatus

func (c *SandboxController) SetPortStatus(id string, port observability.BoundPort, status observability.PortStatus)

func (*SandboxController) SetWriteTracker

func (c *SandboxController) SetWriteTracker(wt controller.WriteTracker)

SetWriteTracker sets the write tracker for recording manual entity writes

func (*SandboxController) StopSandbox added in v0.6.0

func (c *SandboxController) StopSandbox(ctx context.Context, id entity.Id, sb *compute.Sandbox) error

StopSandbox tears down a sandbox and releases every resource it holds. sb is optional: callers that already have the entity pass it so cleanup still works once the entity has been deleted from the store, and it is fetched here otherwise.

func (*SandboxController) UnconfigureFirewall added in v0.6.0

func (c *SandboxController) UnconfigureFirewall(sb *compute.Sandbox)

func (*SandboxController) UpdateServices added in v0.6.0

func (c *SandboxController) UpdateServices(
	ctx context.Context,
	co *compute.Sandbox,
	meta *entity.Meta,
	ep *network.EndpointConfig,
) error

func (*SandboxController) WaitForPort added in v0.6.0

func (c *SandboxController) WaitForPort(ctx context.Context, id string, port int, timeout time.Duration) error

type SandboxControllerDeps added in v0.3.0

type SandboxControllerDeps struct {
	Log       *slog.Logger
	CC        *containerd.Client
	EAC       *entityserver_v1alpha.EntityAccessClient
	Namespace string
	NodeId    compute.NodeId
	NetServ   *network.ServiceManager
	Bridge    string
	Subnet    *netdb.Subnet
	DataPath  string
	Tempdir   string

	LogsMaintainer *observability.LogsMaintainer
	LogWriter      observability.LogWriter
	StatusMon      *observability.StatusMonitor
	Resolver       netresolve.Resolver
	Metrics        *Metrics
	WorkloadIssuer workloadidentity.TokenIssuer

	// ApiAddress is where a sandbox reaches the cluster API, as a literal
	// host:port. It must be an IP: sandbox DNS resolves only app.miren names,
	// so a hostname here would not resolve inside the container. On a
	// coordinator this is the local bridge router; on a distributed runner it
	// is the coordinator, which is a different machine — the local router would
	// be wrong there, since no API listens on it.
	//
	// Empty disables in-cluster API access.
	ApiAddress string

	// CACert is the cluster CA in PEM form, mounted into sandboxes so they can
	// verify the API certificate rather than skipping verification.
	CACert []byte

	// Secrets materializes the secret references a sandbox spec carries, at the
	// moment a container is created. Nil where no backend is reachable, in which
	// case a spec that references one fails rather than starting the container
	// with the reference in place of the value.
	Secrets secret.Resolver

	// Hubs is the stdio fan-out registry shared with the runner's exec server.
	// Optional; one is created if absent, which is what tests want.
	Hubs *HubRegistry

	// SqliteDisks replicates sqlite-provider disks to the coordinator. Nil
	// disables replication; the disks still mount.
	SqliteDisks *sqlitedisk.Manager
}

SandboxControllerDeps holds required dependencies for SandboxController.

type SandboxEntityStore added in v0.6.0

type SandboxEntityStore interface {
	GetSandbox(ctx context.Context, id string) (*compute.Sandbox, *entity.Meta, error)
	PatchSandbox(ctx context.Context, attrs []entity.Attr, revision int64) (int64, error)
}

SandboxEntityStore provides entity read/write operations with write tracking.

type SandboxLogs

type SandboxLogs struct {
	// contains filtered or unexported fields
}

func NewSandboxLogs

func NewSandboxLogs(
	log *slog.Logger,
	entity string,
	attrs map[string]string,
	lw observability.LogWriter,
) *SandboxLogs

func (*SandboxLogs) Stderr

func (s *SandboxLogs) Stderr() *SandboxLogs

func (*SandboxLogs) Write

func (s *SandboxLogs) Write(p []byte) (n int, err error)

type SandboxNetworking added in v0.6.0

type SandboxNetworking interface {
	AllocateNetwork(ctx context.Context, sb *compute.Sandbox) (*network.EndpointConfig, error)
	ReleaseAddr(addr netip.Addr) error
	RebuildEndpointConfig(addresses []string) (*network.EndpointConfig, error)
	BridgeName() string
}

SandboxNetworking provides network allocation and configuration.

type SandboxObservability added in v0.6.0

type SandboxObservability interface {
	AddMetrics(logEntity string, cgroups map[string]string, attrs map[string]string) error
	RemoveMetrics(logEntity string)
	UpdateServices(ctx context.Context, co *compute.Sandbox, meta *entity.Meta, ep *network.EndpointConfig) error
	// LogSandboxEvent writes a runtime lifecycle message to the
	// sandbox's normal log stream, so `miren logs sandbox <id>`
	// surfaces it alongside container output. Intended for startup
	// or teardown events where a container never produced logs of
	// its own (e.g. volume mount failures, image pull failures).
	LogSandboxEvent(sb *compute.Sandbox, shortID, line string)
}

SandboxObservability provides metrics and service management.

type WaitPort added in v0.6.0

type WaitPort struct {
	ID   string
	Port int
}

WaitPort describes a container port to wait for during sandbox creation.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL