Documentation
¶
Overview ¶
Package monitor manages the lifecycle of per-workload resctrl monitoring groups (mon_groups). It provides runtime-agnostic operations for creating, assigning PIDs to, and removing mon_groups — independent of the config-driven pkg/rdt allocation model.
The primary use case is assigning per-pod (or per-container) RMIDs so that downstream energy-monitoring tools like Kepler can attribute hardware energy counters (e.g. Intel AET) to individual workloads.
pkg/monitor vs pkg/rdt ¶
Both packages can create and read resctrl monitoring groups, but they target different models:
Use pkg/rdt when you want config-driven RDT allocation (cache/memory bandwidth partitioning via CtrlGroups) with bundled L3 (CMT/MBM) monitoring and built-in Prometheus/OpenTelemetry export. Its monitoring surface (CreateMonGroup, GetMonData, NewCollector) is coupled to the Initialize + SetConfig lifecycle and is scoped to the L3 resource.
Use pkg/monitor when you need standalone per-workload mon_group lifecycle management decoupled from allocation: no Initialize/SetConfig required, group placement under any (optionally pkg/rdt-managed) ctrl_group, typed counter readings (gauge vs cumulative, with units) across all mon_data domains including Intel AET energy and PERF_PKG counters, key validation and canonicalization (e.g. pod UIDs), pre-fork PID assignment, and crash recovery via Reconcile.
In short: pkg/rdt monitors the classes it allocates (L3, for export); pkg/monitor manages per-workload group lifecycle and reads arbitrary counters, independent of allocation.
Usage:
monitor.SetLogger(slog.Default().WithGroup("monitor"))
mgr, err := monitor.New(monitor.Options{
KeyValidator: monitor.PodUIDValidator,
KeyCanonicalizer: monitor.CanonicalizePodUID,
})
if err != nil {
log.Fatal(err)
}
grp, err := mgr.EnsureGroup(podUID, rdtClass)
if err != nil {
log.Fatal(err)
}
if err := mgr.AssignPID(podUID, containerPID); err != nil {
log.Fatal(err)
}
// On teardown:
mgr.Remove(podUID)
Index ¶
- Variables
- func CanonicalizePodUID(key string) string
- func DefaultKeyValidator(key string) bool
- func DomainInstance(domain string) string
- func DomainPrefix(domain string) string
- func InstrumentName(domain, counter string) string
- func PodUIDValidator(key string) bool
- func SetLogger(l *slog.Logger)
- func Validate(root string) ([]string, error)
- type AttributeFunc
- type FilterFunc
- type Group
- type Manager
- func (m *Manager) AssignPID(key string, pid int) error
- func (m *Manager) EnsureGroup(key, rdtClass string) (*Group, error)
- func (m *Manager) List() []string
- func (m *Manager) ReadCounters(key string) ([]Reading, error)
- func (m *Manager) Reconcile(live []string) error
- func (m *Manager) RegisterOTelInstruments(meter metric.Meter, opts ...OTelOption) (*Registration, error)
- func (m *Manager) Remove(key string) error
- func (m *Manager) Snapshot() map[string]*Group
- type OTelOption
- type Options
- type Reading
- type ReadingKind
- type Registration
Constants ¶
This section is empty.
Variables ¶
var ( // ErrNotTracked is returned when an operation references a key that has // no active mon_group. ErrNotTracked = errors.New("monitor: key not tracked") // ErrNoRMIDs is returned when the kernel has no available RMIDs (mkdir // returns ENOSPC on the resctrl filesystem). ErrNoRMIDs = errors.New("monitor: no RMIDs available") // ErrBadKey is returned when a key fails the configured KeyValidator. ErrBadKey = errors.New("monitor: invalid key") // ErrBadClass is returned when an rdtClass name is unsafe (path // traversal, empty, or contains separators). ErrBadClass = errors.New("monitor: invalid rdt class") // ErrClassMismatch is returned by AssignPID when the target PID already // belongs to a non-root control group that differs from the mon_group's // parent ctrl_group. Writing the PID would silently overwrite its CLOSID // (its CAT/MBA allocation), so the assignment is refused instead. ErrClassMismatch = errors.New("monitor: pid belongs to a different control group") )
Typed errors returned by Manager methods.
Functions ¶
func CanonicalizePodUID ¶
CanonicalizePodUID converts a pod UID to the canonical 8-4-4-4-12 lowercase UUID form. A compact 32-character hex input has dashes inserted; a 36- character input is lowercased and returned unchanged in shape. Any other input is returned lowercased but otherwise untouched, so it is safe to use as a KeyCanonicalizer guarded by PodUIDValidator. Suitable for use as Options.KeyCanonicalizer alongside PodUIDValidator.
func DefaultKeyValidator ¶
DefaultKeyValidator accepts any non-empty key that does not contain path separators, NUL bytes, or dot-segments ("." and ".."), and rejects the specific resctrl directory names "mon_groups", "mon_data", and "info" that are kernel-managed and cannot be reconciled.
func DomainInstance ¶
DomainInstance extracts the numeric instance from a resctrl domain directory name. E.g. "mon_L3_00" → "00", "mon_PERF_PKG_01" → "01".
func DomainPrefix ¶
DomainPrefix extracts the resource prefix from a resctrl domain directory name by taking the first component after "mon_" (lowercased). E.g. "mon_L3_00" → "l3", "mon_PERF_PKG_01" → "perf".
Multi-word qualifiers (e.g. PKG in PERF_PKG) describe the instance scope (package vs core) and are available via the domain.name attribute rather than the instrument name prefix.
func InstrumentName ¶
InstrumentName derives the OTel instrument name from a resctrl domain directory name and counter file name: the domain's resource prefix followed by the counter file name with "_" replaced by ".". The mapping is mechanical (the _bytes suffix is kept) so every instrument name identifies its resctrl file.
Prometheus names are chosen by the consumer's exporter, not by this package, and depend on its translation strategy. For the counter l3.mbm.total.bytes (unit By):
UnderscoreEscapingWithSuffixes → l3_mbm_bytes_total NoUTF8EscapingWithSuffixes → l3.mbm.total.bytes_total
The underscore strategy removes every "total" word from a counter name before appending _total, so pkg/rdt's l3.mbm.total renders the same way. Consumers should select a strategy explicitly, e.g. with prometheus.WithTranslationStrategy(otlptranslator.UnderscoreEscapingWithSuffixes).
Examples:
InstrumentName("mon_L3_00", "llc_occupancy") = "l3.llc.occupancy"
InstrumentName("mon_L3_00", "mbm_local_bytes") = "l3.mbm.local.bytes"
InstrumentName("mon_L3_00", "mbm_total_bytes") = "l3.mbm.total.bytes"
InstrumentName("mon_PERF_PKG_01", "core_energy") = "perf.core.energy"
InstrumentName("mon_PERF_PKG_01", "c1_res") = "perf.c1.res"
func PodUIDValidator ¶
PodUIDValidator matches the Kubernetes pod-UID UUID shape in either of the two forms a container runtime may report:
- the standard 8-4-4-4-12 dashed form, e.g. "a1b2c3d4-e5f6-7890-abcd-ef1234567890" (containerd), and
- the compact 32-character hex form without dashes, e.g. "a1b2c3d4e5f67890abcdef1234567890" (some CRI-O versions).
Hex digits of either case are accepted. Pair this validator with CanonicalizePodUID (via Options.KeyCanonicalizer) so the on-disk mon_group directory name is always the canonical dashed form regardless of which form the runtime reported.
func SetLogger ¶
SetLogger sets the logger used by the package. Safe to call before New and concurrently with Manager operations. A nil argument resets to the default logger.
func Validate ¶
Validate checks that the resctrl filesystem at root is mounted and has monitoring support. It returns the deduplicated, sorted union of counter file names found across all domain subdirectories under root/mon_data. An error is returned if the path does not exist or has no mon_data subdirectory.
Types ¶
type AttributeFunc ¶
AttributeFunc returns additional OTel attributes for a reading. Called for each group during collection. The key is the canonicalized tracking key (e.g. dashed pod UID), path is the absolute mon_group directory.
type FilterFunc ¶
FilterFunc returns whether a Reading should be exported. Return false to suppress the reading (e.g. to gate perf counters behind a config flag).
type Group ¶
type Group struct {
// contains filtered or unexported fields
}
Group is a handle to one mon_group on the resctrl filesystem.
func (*Group) Class ¶
Class returns the rdtClass (ctrl_group) the mon_group lives under. An empty string means the mon_group is under the root resctrl group.
func (*Group) Gen ¶
Gen returns the group's generation, which increments each time the mon_group for this key is (re)created. Consumers use it to detect RMID reuse across a remove/recreate of the same key.
type Manager ¶
type Manager struct {
// contains filtered or unexported fields
}
Manager owns the lifecycle of per-workload resctrl mon_groups.
It is safe for concurrent use from multiple goroutines.
func (*Manager) AssignPID ¶
AssignPID writes pid to the group's tasks file. The kernel assigns the RMID to this PID and all future child processes. Call while the init process is created but paused (the pre-fork window) for race-free attribution.
func (*Manager) EnsureGroup ¶
EnsureGroup idempotently creates a mon_group for key under an optional pre-existing rdtClass ctrl_group ("" = root resctrl directory). It never creates the ctrl_group itself. Returns a Group handle on success.
func (*Manager) ReadCounters ¶
ReadCounters walks <group>/mon_data/<domain>/* for the tracked key and returns every readable counter file. Missing files/dirs are skipped, not errors: not every domain exposes every counter. Returns ErrNotTracked if the key is unknown. The caller key is canonicalized once here.
func (*Manager) Reconcile ¶
Reconcile removes on-disk mon_groups whose key is not present in live. Eligibility for removal requires all of:
- the directory name satisfies KeyValidator (scopes reaping to recognized name shapes — e.g. UUID-shaped names with PodUIDValidator);
- the canonicalized name is NOT in the live set;
- the directory is not the authoritative path for a tracked entry.
Kernel metadata directories (mon_groups, mon_data, info) are always skipped.
Note: Reconcile identifies groups by their directory name (the key). If the same key were to appear under multiple ctrl_groups due to misconfiguration, only the in-memory tracked instance is authoritative; duplicates under other ctrl_groups are treated as orphans and removed.
Reconcile is best-effort but observable: it continues past per-directory failures (unreadable mon_groups/ directories, failed removals) rather than aborting, and returns all such failures aggregated with errors.Join. A nil return means every eligible orphan was reaped successfully.
func (*Manager) RegisterOTelInstruments ¶
func (m *Manager) RegisterOTelInstruments(meter metric.Meter, opts ...OTelOption) (*Registration, error)
RegisterOTelInstruments registers observable OTel instruments for all counters readable by the Manager. A batch callback is registered with the provided Meter that calls Snapshot + ReadCounters on each collection cycle.
Instrument names are derived mechanically from the resctrl domain and counter file names:
mon_L3_00/llc_occupancy → l3.llc.occupancy (unit: By) mon_L3_00/mbm_local_bytes → l3.mbm.local.bytes (unit: By) mon_L3_00/mbm_total_bytes → l3.mbm.total.bytes (unit: By) mon_PERF_PKG_00/core_energy → perf.core.energy (unit: J) mon_PERF_PKG_00/activity → perf.activity (unit: farads; kernel nF × 1e-9) mon_PERF_PKG_00/c1_res → perf.c1.res mon_PERF_PKG_00/c6_res → perf.c6.res mon_PERF_PKG_00/uops_retired → perf.uops.retired mon_PERF_PKG_00/unhalted_core_cycles → perf.unhalted.core.cycles mon_PERF_PKG_00/unhalted_ref_cycles → perf.unhalted.ref.cycles mon_PERF_PKG_00/stalls_llc_miss → perf.stalls.llc.miss mon_PERF_PKG_00/stalls_llc_hit → perf.stalls.llc.hit
The L3 counters and core_energy/activity are always present when the hardware supports them. The remaining PERF_PKG counters (c1_res, c6_res, uops_retired, unhalted_*_cycles, stalls_*) require the kernel command-line parameter "rdt=perf". Counter names are platform-dependent; the library registers instruments for whatever files the kernel exposes.
The domain prefix is the first component after "mon_" (lowercased): "mon_L3_00" → "l3", "mon_PERF_PKG_00" → "perf". Multi-word qualifiers like PKG describe the instance scope; the full domain directory name is available via the "domain.name" attribute (e.g. "mon_PERF_PKG_00").
L3 instrument names are similar to pkg/rdt's RegisterOpenTelemetryInstruments but preserve the _bytes counter suffix (e.g. l3.mbm.total.bytes vs pkg/rdt's l3.mbm.total). This maintains backward compatibility with the kernel counter file names as a mechanical derivation. See InstrumentName for how Prometheus exporters render these names.
Each metric carries a "domain.id" attribute with the numeric instance (e.g. "00") and a "domain.name" attribute with the full domain directory name (e.g. "mon_PERF_PKG_00"). Additional per-group attributes can be injected via WithAttributes.
Monotonic counters (ReadingKind == Cumulative) are smoothed via an internal accumulator that suppresses brief negative deltas caused by cross-CPU aggregation races in the kernel.
The caller owns the Meter and MeterProvider lifecycle (endpoint, push interval, shutdown). Call Unregister (or Close) on the returned Registration during shutdown to remove the batch callback and release its reference to the Manager.
func (*Manager) Remove ¶
Remove deletes the mon_group for key (kernel releases the RMID) and drops all in-memory state for that key. Removing a directory that is already gone on disk is not an error, but calling Remove for a key that is not tracked returns ErrNotTracked.
The lock is held across the on-disk removal so that a concurrent EnsureGroup cannot create a new entry that is immediately clobbered, and so that a failed rmdir leaves the entry intact (retryable) without risk of state corruption. Note: AssignPID/ReadCounters may still observe the entry (fetched before Remove acquires the lock) and attempt I/O on the now-removed directory. AssignPID will return an error wrapping ENOENT; ReadCounters will return empty readings (it treats a missing mon_data directory as a non-error). Both outcomes are benign transient races — the key will be gone from the map by the time the caller retries.
type OTelOption ¶
type OTelOption func(*otelConfig)
OTelOption configures RegisterOTelInstruments.
func WithAttributes ¶
func WithAttributes(fn AttributeFunc) OTelOption
WithAttributes sets a function that provides extra attributes per group. The attributes are appended to the default domain.id and domain.name attributes on each observed metric value.
func WithFilter ¶
func WithFilter(fn FilterFunc) OTelOption
WithFilter sets a function that gates which readings are exported. If the filter returns false, the reading is silently skipped.
type Options ¶
type Options struct {
// ResctrlRoot is the resctrl mount point. Default: "/sys/fs/resctrl".
// Tests typically set this to a temp directory.
ResctrlRoot string
// KeyValidator, if set, rejects keys that do not satisfy it (e.g. the
// pod-UID UUID shape). Default: accept any non-empty key without path
// separators or NUL bytes.
//
// The validator also scopes Reconcile: only on-disk mon_group directories
// whose name satisfies KeyValidator are eligible for orphan removal.
// Directories that do not match the validator are unconditionally skipped,
// so a narrow validator (e.g. PodUIDValidator) limits the blast radius of
// reconciliation to the set of directory names it recognizes. Directories
// that DO match but are not in the live set will still be removed.
KeyValidator func(key string) bool
// KeyCanonicalizer, if set, normalizes a caller key to a canonical form
// before it is used as a mon_group directory name and as the in-memory
// tracking key. It is applied after KeyValidator. Default: identity.
//
// Pair it with a matching KeyValidator (e.g. CanonicalizePodUID alongside
// PodUIDValidator) so that keys reported in different-but-equivalent forms
// (e.g. a pod UID with or without dashes) map to a single, predictable
// on-disk directory name.
//
// It must be idempotent: canonicalizing an already-canonical key returns it
// unchanged (canon(canon(k)) == canon(k)). The bundled CanonicalizePodUID
// satisfies this; internal callers pass already-canonical keys and rely on
// it to avoid re-normalizing.
KeyCanonicalizer func(key string) string
}
Options configures a Manager.
type Reading ¶
type Reading struct {
Domain string // resctrl mon_data subdir, e.g. "mon_L3_00", "mon_PERF_PKG_01"
Name string // counter file name, e.g. "llc_occupancy", "core_energy"
Value float64 // parsed value (float to cover core_energy/activity)
Kind ReadingKind // Gauge or Cumulative
Unit string // UCUM unit of Value where known (e.g. "By", "J", "nF"), otherwise ""
}
Reading is one raw counter sample from one monitoring domain.
type ReadingKind ¶
type ReadingKind int
ReadingKind distinguishes gauge from cumulative counter semantics.
const ( // Gauge is an instantaneous point-in-time value (e.g. llc_occupancy). Gauge ReadingKind = iota // Cumulative is a monotonically increasing counter (e.g. mbm_total_bytes, core_energy). Cumulative )
type Registration ¶
type Registration struct {
// contains filtered or unexported fields
}
Registration is the handle returned by RegisterOTelInstruments. It owns the OTel batch callback registered with the caller's Meter. Callers should call Unregister (or the io.Closer-compatible Close) during shutdown to stop further observations and release the callback's reference to the Manager; leaking it keeps the Manager reachable for the lifetime of the MeterProvider.
Do not copy a Registration; use the pointer returned by RegisterOTelInstruments. Copying it by value copies the embedded sync.Once (go vet's copylocks flags this), defeating the idempotency of Unregister.
func (*Registration) Close ¶
func (r *Registration) Close() error
Close implements io.Closer by unregistering the callback, so callers can use defer reg.Close(). It is equivalent to Unregister.
func (*Registration) Unregister ¶
func (r *Registration) Unregister() error
Unregister removes the batch callback from the Meter, stopping further observations and releasing the callback's reference to the Manager. It is idempotent and safe to call on a nil *Registration; the first call performs the unregister and subsequent calls return the same result.