monitor

package
v0.14.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 25, 2026 License: Apache-2.0 Imports: 15 Imported by: 0

Documentation

Overview

Package monitor manages the lifecycle of per-workload resctrl monitoring groups (mon_groups). It provides runtime-agnostic operations for creating, assigning PIDs to, and removing mon_groups — independent of the config-driven pkg/rdt allocation model.

The primary use case is assigning per-pod (or per-container) RMIDs so that downstream energy-monitoring tools like Kepler can attribute hardware energy counters (e.g. Intel AET) to individual workloads.

pkg/monitor vs pkg/rdt

Both packages can create and read resctrl monitoring groups, but they target different models:

  • Use pkg/rdt when you want config-driven RDT allocation (cache/memory bandwidth partitioning via CtrlGroups) with bundled L3 (CMT/MBM) monitoring and built-in Prometheus/OpenTelemetry export. Its monitoring surface (CreateMonGroup, GetMonData, NewCollector) is coupled to the Initialize + SetConfig lifecycle and is scoped to the L3 resource.

  • Use pkg/monitor when you need standalone per-workload mon_group lifecycle management decoupled from allocation: no Initialize/SetConfig required, group placement under any (optionally pkg/rdt-managed) ctrl_group, typed counter readings (gauge vs cumulative, with units) across all mon_data domains including Intel AET energy and PERF_PKG counters, key validation and canonicalization (e.g. pod UIDs), pre-fork PID assignment, and crash recovery via Reconcile.

In short: pkg/rdt monitors the classes it allocates (L3, for export); pkg/monitor manages per-workload group lifecycle and reads arbitrary counters, independent of allocation.

Usage:

monitor.SetLogger(slog.Default().WithGroup("monitor"))

mgr, err := monitor.New(monitor.Options{
    KeyValidator:      monitor.PodUIDValidator,
    KeyCanonicalizer:  monitor.CanonicalizePodUID,
})
if err != nil {
    log.Fatal(err)
}

grp, err := mgr.EnsureGroup(podUID, rdtClass)
if err != nil {
    log.Fatal(err)
}

if err := mgr.AssignPID(podUID, containerPID); err != nil {
    log.Fatal(err)
}

// On teardown:
mgr.Remove(podUID)

Index

Constants

This section is empty.

Variables

View Source
var (
	// ErrNotTracked is returned when an operation references a key that has
	// no active mon_group.
	ErrNotTracked = errors.New("monitor: key not tracked")

	// ErrNoRMIDs is returned when the kernel has no available RMIDs (mkdir
	// returns ENOSPC on the resctrl filesystem).
	ErrNoRMIDs = errors.New("monitor: no RMIDs available")

	// ErrBadKey is returned when a key fails the configured KeyValidator.
	ErrBadKey = errors.New("monitor: invalid key")

	// ErrBadClass is returned when an rdtClass name is unsafe (path
	// traversal, empty, or contains separators).
	ErrBadClass = errors.New("monitor: invalid rdt class")

	// ErrClassMismatch is returned by AssignPID when the target PID already
	// belongs to a non-root control group that differs from the mon_group's
	// parent ctrl_group. Writing the PID would silently overwrite its CLOSID
	// (its CAT/MBA allocation), so the assignment is refused instead.
	ErrClassMismatch = errors.New("monitor: pid belongs to a different control group")
)

Typed errors returned by Manager methods.

Functions

func CanonicalizePodUID

func CanonicalizePodUID(key string) string

CanonicalizePodUID converts a pod UID to the canonical 8-4-4-4-12 lowercase UUID form. A compact 32-character hex input has dashes inserted; a 36- character input is lowercased and returned unchanged in shape. Any other input is returned lowercased but otherwise untouched, so it is safe to use as a KeyCanonicalizer guarded by PodUIDValidator. Suitable for use as Options.KeyCanonicalizer alongside PodUIDValidator.

func DefaultKeyValidator

func DefaultKeyValidator(key string) bool

DefaultKeyValidator accepts any non-empty key that does not contain path separators, NUL bytes, or dot-segments ("." and ".."), and rejects the specific resctrl directory names "mon_groups", "mon_data", and "info" that are kernel-managed and cannot be reconciled.

func DomainInstance

func DomainInstance(domain string) string

DomainInstance extracts the numeric instance from a resctrl domain directory name. E.g. "mon_L3_00" → "00", "mon_PERF_PKG_01" → "01".

func DomainPrefix

func DomainPrefix(domain string) string

DomainPrefix extracts the resource prefix from a resctrl domain directory name by taking the first component after "mon_" (lowercased). E.g. "mon_L3_00" → "l3", "mon_PERF_PKG_01" → "perf".

Multi-word qualifiers (e.g. PKG in PERF_PKG) describe the instance scope (package vs core) and are available via the domain.name attribute rather than the instrument name prefix.

func InstrumentName

func InstrumentName(domain, counter string) string

InstrumentName derives the OTel instrument name from a resctrl domain directory name and counter file name: the domain's resource prefix followed by the counter file name with "_" replaced by ".". The mapping is mechanical (the _bytes suffix is kept) so every instrument name identifies its resctrl file.

Prometheus names are chosen by the consumer's exporter, not by this package, and depend on its translation strategy. For the counter l3.mbm.total.bytes (unit By):

UnderscoreEscapingWithSuffixes → l3_mbm_bytes_total
NoUTF8EscapingWithSuffixes     → l3.mbm.total.bytes_total

The underscore strategy removes every "total" word from a counter name before appending _total, so pkg/rdt's l3.mbm.total renders the same way. Consumers should select a strategy explicitly, e.g. with prometheus.WithTranslationStrategy(otlptranslator.UnderscoreEscapingWithSuffixes).

Examples:

InstrumentName("mon_L3_00", "llc_occupancy")     = "l3.llc.occupancy"
InstrumentName("mon_L3_00", "mbm_local_bytes")   = "l3.mbm.local.bytes"
InstrumentName("mon_L3_00", "mbm_total_bytes")   = "l3.mbm.total.bytes"
InstrumentName("mon_PERF_PKG_01", "core_energy") = "perf.core.energy"
InstrumentName("mon_PERF_PKG_01", "c1_res")      = "perf.c1.res"

func PodUIDValidator

func PodUIDValidator(key string) bool

PodUIDValidator matches the Kubernetes pod-UID UUID shape in either of the two forms a container runtime may report:

  • the standard 8-4-4-4-12 dashed form, e.g. "a1b2c3d4-e5f6-7890-abcd-ef1234567890" (containerd), and
  • the compact 32-character hex form without dashes, e.g. "a1b2c3d4e5f67890abcdef1234567890" (some CRI-O versions).

Hex digits of either case are accepted. Pair this validator with CanonicalizePodUID (via Options.KeyCanonicalizer) so the on-disk mon_group directory name is always the canonical dashed form regardless of which form the runtime reported.

func SetLogger

func SetLogger(l *slog.Logger)

SetLogger sets the logger used by the package. Safe to call before New and concurrently with Manager operations. A nil argument resets to the default logger.

func Validate

func Validate(root string) ([]string, error)

Validate checks that the resctrl filesystem at root is mounted and has monitoring support. It returns the deduplicated, sorted union of counter file names found across all domain subdirectories under root/mon_data. An error is returned if the path does not exist or has no mon_data subdirectory.

Types

type AttributeFunc

type AttributeFunc func(key, path string) []attribute.KeyValue

AttributeFunc returns additional OTel attributes for a reading. Called for each group during collection. The key is the canonicalized tracking key (e.g. dashed pod UID), path is the absolute mon_group directory.

type FilterFunc

type FilterFunc func(Reading) bool

FilterFunc returns whether a Reading should be exported. Return false to suppress the reading (e.g. to gate perf counters behind a config flag).

type Group

type Group struct {
	// contains filtered or unexported fields
}

Group is a handle to one mon_group on the resctrl filesystem.

func (*Group) Class

func (g *Group) Class() string

Class returns the rdtClass (ctrl_group) the mon_group lives under. An empty string means the mon_group is under the root resctrl group.

func (*Group) Gen

func (g *Group) Gen() uint64

Gen returns the group's generation, which increments each time the mon_group for this key is (re)created. Consumers use it to detect RMID reuse across a remove/recreate of the same key.

func (*Group) Key

func (g *Group) Key() string

Key returns the canonicalized tracking key (e.g. dashed pod UID) for this group.

func (*Group) Path

func (g *Group) Path() string

Path returns the absolute filesystem path of the mon_group directory.

type Manager

type Manager struct {
	// contains filtered or unexported fields
}

Manager owns the lifecycle of per-workload resctrl mon_groups.

It is safe for concurrent use from multiple goroutines.

func New

func New(o Options) (*Manager, error)

New creates a Manager with the given options.

func (*Manager) AssignPID

func (m *Manager) AssignPID(key string, pid int) error

AssignPID writes pid to the group's tasks file. The kernel assigns the RMID to this PID and all future child processes. Call while the init process is created but paused (the pre-fork window) for race-free attribution.

func (*Manager) EnsureGroup

func (m *Manager) EnsureGroup(key, rdtClass string) (*Group, error)

EnsureGroup idempotently creates a mon_group for key under an optional pre-existing rdtClass ctrl_group ("" = root resctrl directory). It never creates the ctrl_group itself. Returns a Group handle on success.

func (*Manager) List

func (m *Manager) List() []string

List returns the keys currently tracked in memory.

func (*Manager) ReadCounters

func (m *Manager) ReadCounters(key string) ([]Reading, error)

ReadCounters walks <group>/mon_data/<domain>/* for the tracked key and returns every readable counter file. Missing files/dirs are skipped, not errors: not every domain exposes every counter. Returns ErrNotTracked if the key is unknown. The caller key is canonicalized once here.

func (*Manager) Reconcile

func (m *Manager) Reconcile(live []string) error

Reconcile removes on-disk mon_groups whose key is not present in live. Eligibility for removal requires all of:

  • the directory name satisfies KeyValidator (scopes reaping to recognized name shapes — e.g. UUID-shaped names with PodUIDValidator);
  • the canonicalized name is NOT in the live set;
  • the directory is not the authoritative path for a tracked entry.

Kernel metadata directories (mon_groups, mon_data, info) are always skipped.

Note: Reconcile identifies groups by their directory name (the key). If the same key were to appear under multiple ctrl_groups due to misconfiguration, only the in-memory tracked instance is authoritative; duplicates under other ctrl_groups are treated as orphans and removed.

Reconcile is best-effort but observable: it continues past per-directory failures (unreadable mon_groups/ directories, failed removals) rather than aborting, and returns all such failures aggregated with errors.Join. A nil return means every eligible orphan was reaped successfully.

func (*Manager) RegisterOTelInstruments

func (m *Manager) RegisterOTelInstruments(meter metric.Meter, opts ...OTelOption) (*Registration, error)

RegisterOTelInstruments registers observable OTel instruments for all counters readable by the Manager. A batch callback is registered with the provided Meter that calls Snapshot + ReadCounters on each collection cycle.

Instrument names are derived mechanically from the resctrl domain and counter file names:

mon_L3_00/llc_occupancy          → l3.llc.occupancy           (unit: By)
mon_L3_00/mbm_local_bytes        → l3.mbm.local.bytes         (unit: By)
mon_L3_00/mbm_total_bytes        → l3.mbm.total.bytes         (unit: By)
mon_PERF_PKG_00/core_energy      → perf.core.energy           (unit: J)
mon_PERF_PKG_00/activity         → perf.activity              (unit: farads; kernel nF × 1e-9)
mon_PERF_PKG_00/c1_res           → perf.c1.res
mon_PERF_PKG_00/c6_res           → perf.c6.res
mon_PERF_PKG_00/uops_retired     → perf.uops.retired
mon_PERF_PKG_00/unhalted_core_cycles → perf.unhalted.core.cycles
mon_PERF_PKG_00/unhalted_ref_cycles  → perf.unhalted.ref.cycles
mon_PERF_PKG_00/stalls_llc_miss  → perf.stalls.llc.miss
mon_PERF_PKG_00/stalls_llc_hit   → perf.stalls.llc.hit

The L3 counters and core_energy/activity are always present when the hardware supports them. The remaining PERF_PKG counters (c1_res, c6_res, uops_retired, unhalted_*_cycles, stalls_*) require the kernel command-line parameter "rdt=perf". Counter names are platform-dependent; the library registers instruments for whatever files the kernel exposes.

The domain prefix is the first component after "mon_" (lowercased): "mon_L3_00" → "l3", "mon_PERF_PKG_00" → "perf". Multi-word qualifiers like PKG describe the instance scope; the full domain directory name is available via the "domain.name" attribute (e.g. "mon_PERF_PKG_00").

L3 instrument names are similar to pkg/rdt's RegisterOpenTelemetryInstruments but preserve the _bytes counter suffix (e.g. l3.mbm.total.bytes vs pkg/rdt's l3.mbm.total). This maintains backward compatibility with the kernel counter file names as a mechanical derivation. See InstrumentName for how Prometheus exporters render these names.

Each metric carries a "domain.id" attribute with the numeric instance (e.g. "00") and a "domain.name" attribute with the full domain directory name (e.g. "mon_PERF_PKG_00"). Additional per-group attributes can be injected via WithAttributes.

Monotonic counters (ReadingKind == Cumulative) are smoothed via an internal accumulator that suppresses brief negative deltas caused by cross-CPU aggregation races in the kernel.

The caller owns the Meter and MeterProvider lifecycle (endpoint, push interval, shutdown). Call Unregister (or Close) on the returned Registration during shutdown to remove the batch callback and release its reference to the Manager.

func (*Manager) Remove

func (m *Manager) Remove(key string) error

Remove deletes the mon_group for key (kernel releases the RMID) and drops all in-memory state for that key. Removing a directory that is already gone on disk is not an error, but calling Remove for a key that is not tracked returns ErrNotTracked.

The lock is held across the on-disk removal so that a concurrent EnsureGroup cannot create a new entry that is immediately clobbered, and so that a failed rmdir leaves the entry intact (retryable) without risk of state corruption. Note: AssignPID/ReadCounters may still observe the entry (fetched before Remove acquires the lock) and attempt I/O on the now-removed directory. AssignPID will return an error wrapping ENOENT; ReadCounters will return empty readings (it treats a missing mon_data directory as a non-error). Both outcomes are benign transient races — the key will be gone from the map by the time the caller retries.

func (*Manager) Snapshot

func (m *Manager) Snapshot() map[string]*Group

Snapshot returns a point-in-time map of tracked key -> *Group handle. The returned Group values are copies safe to use after the lock is released.

type OTelOption

type OTelOption func(*otelConfig)

OTelOption configures RegisterOTelInstruments.

func WithAttributes

func WithAttributes(fn AttributeFunc) OTelOption

WithAttributes sets a function that provides extra attributes per group. The attributes are appended to the default domain.id and domain.name attributes on each observed metric value.

func WithFilter

func WithFilter(fn FilterFunc) OTelOption

WithFilter sets a function that gates which readings are exported. If the filter returns false, the reading is silently skipped.

type Options

type Options struct {
	// ResctrlRoot is the resctrl mount point. Default: "/sys/fs/resctrl".
	// Tests typically set this to a temp directory.
	ResctrlRoot string

	// KeyValidator, if set, rejects keys that do not satisfy it (e.g. the
	// pod-UID UUID shape). Default: accept any non-empty key without path
	// separators or NUL bytes.
	//
	// The validator also scopes Reconcile: only on-disk mon_group directories
	// whose name satisfies KeyValidator are eligible for orphan removal.
	// Directories that do not match the validator are unconditionally skipped,
	// so a narrow validator (e.g. PodUIDValidator) limits the blast radius of
	// reconciliation to the set of directory names it recognizes. Directories
	// that DO match but are not in the live set will still be removed.
	KeyValidator func(key string) bool

	// KeyCanonicalizer, if set, normalizes a caller key to a canonical form
	// before it is used as a mon_group directory name and as the in-memory
	// tracking key. It is applied after KeyValidator. Default: identity.
	//
	// Pair it with a matching KeyValidator (e.g. CanonicalizePodUID alongside
	// PodUIDValidator) so that keys reported in different-but-equivalent forms
	// (e.g. a pod UID with or without dashes) map to a single, predictable
	// on-disk directory name.
	//
	// It must be idempotent: canonicalizing an already-canonical key returns it
	// unchanged (canon(canon(k)) == canon(k)). The bundled CanonicalizePodUID
	// satisfies this; internal callers pass already-canonical keys and rely on
	// it to avoid re-normalizing.
	KeyCanonicalizer func(key string) string
}

Options configures a Manager.

type Reading

type Reading struct {
	Domain string      // resctrl mon_data subdir, e.g. "mon_L3_00", "mon_PERF_PKG_01"
	Name   string      // counter file name, e.g. "llc_occupancy", "core_energy"
	Value  float64     // parsed value (float to cover core_energy/activity)
	Kind   ReadingKind // Gauge or Cumulative
	Unit   string      // UCUM unit of Value where known (e.g. "By", "J", "nF"), otherwise ""
}

Reading is one raw counter sample from one monitoring domain.

type ReadingKind

type ReadingKind int

ReadingKind distinguishes gauge from cumulative counter semantics.

const (
	// Gauge is an instantaneous point-in-time value (e.g. llc_occupancy).
	Gauge ReadingKind = iota
	// Cumulative is a monotonically increasing counter (e.g. mbm_total_bytes, core_energy).
	Cumulative
)

type Registration

type Registration struct {
	// contains filtered or unexported fields
}

Registration is the handle returned by RegisterOTelInstruments. It owns the OTel batch callback registered with the caller's Meter. Callers should call Unregister (or the io.Closer-compatible Close) during shutdown to stop further observations and release the callback's reference to the Manager; leaking it keeps the Manager reachable for the lifetime of the MeterProvider.

Do not copy a Registration; use the pointer returned by RegisterOTelInstruments. Copying it by value copies the embedded sync.Once (go vet's copylocks flags this), defeating the idempotency of Unregister.

func (*Registration) Close

func (r *Registration) Close() error

Close implements io.Closer by unregistering the callback, so callers can use defer reg.Close(). It is equivalent to Unregister.

func (*Registration) Unregister

func (r *Registration) Unregister() error

Unregister removes the batch callback from the Meter, stopping further observations and releasing the callback's reference to the Manager. It is idempotent and safe to call on a nil *Registration; the first call performs the unregister and subsequent calls return the same result.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL