contention

package
v0.15.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 17, 2026 License: MIT Imports: 43 Imported by: 0

Documentation

Overview

Package contention is the committed contention observatory for the GoGraph Optimization Laboratory (rmp sprint 353, task #2678).

Why it exists

Before this package, nothing in the repository enabled Go's contention profilers. No source file called runtime.SetMutexProfileFraction or runtime.SetBlockProfileRate; the only traces were a comment in bench/audit352/gctax_test.go and prose in docs/ describing ad-hoc `go test -mutexprofile` invocations that were never committed. Contention was therefore inferred from the SHAPE of a throughput curve and never attributed to a lock. Compliance Mandate 3 requires the opposite: readiness for extreme concurrency is "proven, never presumed".

This package turns that into a repeatable instrument.

The two-window design, and why it is not one window

Mutex and block profiling perturb the very thing they measure. At full rate the runtime timestamps every lock handoff and every blocking event, which inflates wall-clock and can REORDER contention: a lock that is hot without the profiler may look cool with it, because the profiler's own bookkeeping changes the arrival pattern. A single profiled run therefore cannot supply both the effect and the cause.

So every measurement runs two windows over the identical workload:

  • the UNPROFILED window (WindowEffect) supplies the effect — throughput, p50, p99. These are the numbers that may be quoted for scaling.
  • the PROFILED window (WindowProbe) supplies the cause — mutex and block attribution per lock call site. Its throughput is recorded but must NEVER be quoted as the module's throughput; it is the probe's throughput.

Reading a scaling claim off the profiled window is the classic error this design exists to prevent.

One fresh process per WINDOW

Each window runs in its own child process. There are two independent reasons, and both were established by measurement rather than by argument.

The first is profile cumulativity. The mutex and block profiles accumulate for the lifetime of a process and the runtime exposes no way to reset them. Two workloads measured in one process contaminate each other, and the second inherits the first one's samples. An early revision ran the instrument's own controls in the shared test process, and the negative control caught it by "detecting" contention in a fixture that has no shared state at all.

The second is heap warmth, and it is why a process per MEASUREMENT was not enough. When both windows ran in one child, the unprofiled window always ran first and paid every first-touch cost — plan compilation, heap growth, page faults, pool population — while the profiled window inherited all of it warm, along with a raised GC goal. Measured on cypher-write-mem at level 16, over six fresh processes, the PROFILED window was FASTER in five of the six, by about 8% at the median. That is not a credible physical result; it is the ordering confound swamping the effect the design set out to expose, and it would have been published as a probe_slowdown below 1.0 — the instrument claiming that turning full-rate profiling on speeds the module up.

A warm-up pass narrows that gap but cannot close it, because the second window still inherits the whole of the first window's run on top of any warm-up. Alternating the order merely averages two different confounds. So the same principle the package already applies to profile state is applied to heap state: one window, one process, no inheritance.

Concurrency

Observe is NOT safe for concurrent use and must not be called from two goroutines: it mutates the process-global profiling rates. It is called once per child process, which is exactly how the sweep drives it.

Index

Constants

View Source
const (
	// TransportSim is the in-memory [sim.SimListener] pipe every committed
	// Bolt workload runs over.
	TransportSim = "sim"
	// TransportTCP is a real loopback socket on 127.0.0.1.
	TransportTCP = "tcp"

	// TransportSimReadDeadlineCleared converts the server's per-read deadline
	// to the zero Time: the mutex and Broadcast still happen, the per-wait
	// timer and goroutine do not.
	TransportSimReadDeadlineCleared = "sim-rdl-clear"
	// TransportSimReadDeadlineDropped drops the per-read deadline call
	// entirely: neither the Broadcast nor the per-wait timer happens.
	TransportSimReadDeadlineDropped = "sim-rdl-drop"
	// TransportSimWriteDeadlineDropped drops the per-record write deadline
	// call entirely, which on the pipe is one mutex acquisition and one
	// spurious Broadcast per record.
	TransportSimWriteDeadlineDropped = "sim-wdl-drop"
	// TransportSimNoDeadlines drops both.
	TransportSimNoDeadlines = "sim-no-dl"
	// TransportTCPNoDeadlines drops both on the socket. It is the CONTROL: the
	// same removal, on a transport where the deadline is cheap, must buy
	// little or nothing. Without it, a gain on the pipe arms could be the
	// removal of any per-message work rather than of the pipe's own.
	TransportTCPNoDeadlines = "tcp-no-dl"

	// TransportCommitted is not a transport at all: it is the COMMITTED
	// workload of the same shape, taken straight from [ByName], driven through
	// this campaign's replica machinery.
	//
	// It is the FIDELITY control, and it is here because without it the whole
	// campaign rests on an unchecked assumption. The sim arm is built to stand
	// in for bolt-wire-rows and bolt-wire-read, but it is not identical to
	// them: it discards the server log where they let it fall through to
	// slog.Default, it pre-dials its connections where they dial lazily inside
	// the first operation, and it reaches the client through
	// [sim.NewWireClientNetConn] rather than [sim.NewWireClient]. Whether those
	// three differences move the number is a question, not a given — and it
	// became a live one when the sim arm read 0.920 at level 8 against the
	// published bolt-wire-rows row's 1.087.
	//
	// The arm overrides the committed workload's Ops with this campaign's, so
	// the two arms do the same total work. A scaling ratio is a rate and is
	// therefore invariant to that, but the WALL CLOCK of a cell is not, and two
	// arms of one interleaved campaign must cost the same or the interleaving
	// stops being balanced.
	TransportCommitted = "committed"

	// TransportSimLazyDial dials each worker's connection inside its FIRST
	// operation through a padded [perWorker] table, exactly as boltClient
	// does, instead of pre-dialling every connection in Setup.
	TransportSimLazyDial = "sim-lazydial"
	// TransportSimDefaultLog leaves Options.Logger nil, so server.NewServer
	// falls through to slog.Default as [sim.NewSimServer] does, instead of
	// installing a discarding handler.
	TransportSimDefaultLog = "sim-defaultlog"
	// TransportSimConnClient builds the client with [sim.NewWireClient] over
	// the concrete *SimConn, as every committed Bolt workload does, instead of
	// [sim.NewWireClientNetConn] over the net.Conn interface.
	TransportSimConnClient = "sim-simconn-client"

	// TransportTCPLazyDial is [TransportSimLazyDial] on the socket, and it is
	// REQUIRED, not optional.
	//
	// The bisect above found that dialling in Setup rather than inside the
	// worker's first operation costs the pipe arm 25% of its level-8
	// throughput. If that penalty is a property of the pipe, then a pre-dialled
	// sim arm compared against a pre-dialled tcp arm understates the pipe and
	// inflates the headline. The A/B must therefore be taken with the dial site
	// held constant AND known to be neutral on both sides, which is what this
	// arm establishes.
	TransportTCPLazyDial = "tcp-lazydial"
)

Variables

This section is empty.

Functions

func Levels

func Levels() []int

Levels returns the goroutine ladder every workload walks.

It returns a fresh slice on every call, so the ladder cannot be mutated through the returned value by one caller and observed changed by another. Levels is safe for concurrent use.

func TransportArmNames

func TransportArmNames() []string

TransportArmNames returns every arm TransportWorkload accepts, sorted.

It returns a fresh slice on every call and is safe for concurrent use.

func TransportKinds

func TransportKinds() []string

TransportKinds returns the two transports the A/B compares. It deliberately excludes the attribution probes: a campaign that swept them alongside the two real transports would report probe throughput in the same table as module throughput.

It returns a fresh slice on every call and is safe for concurrent use.

func TransportQueryNames

func TransportQueryNames() []string

TransportQueryNames returns the query keys TransportWorkload accepts, sorted.

It returns a fresh slice on every call and is safe for concurrent use.

func TransportWorkloadName

func TransportWorkloadName(kind, query string) string

TransportWorkloadName is the stable name of one arm, used in filenames and in the report.

func WriteMetrics

func WriteMetrics(dir string, m *Metrics) error

WriteMetrics serialises one window's Metrics into dir, next to whatever profiles that window produced, so a sweep can be re-read without re-running it.

It is safe for concurrent use only across distinct dir values: two calls naming the same directory race on the same file.

func WriteResult

func WriteResult(dir string, r *Result) error

WriteResult serialises a paired Result into dir, so the two windows of one measurement can be read back together.

It is safe for concurrent use only across distinct dir values.

Types

type Metrics

type Metrics struct {
	Workload  string  `json:"workload"`
	Surface   string  `json:"surface"`
	Level     int     `json:"level"`
	Ops       int     `json:"ops"`
	Profiled  bool    `json:"profiled"`
	WallNanos int64   `json:"wall_nanos"`
	OpsPerSec float64 `json:"ops_per_sec"`
	P50Nanos  int64   `json:"p50_nanos"`
	P99Nanos  int64   `json:"p99_nanos"`
	MaxNanos  int64   `json:"max_nanos"`
	// LatencySampleEvery is the systematic-sampling stride used for the
	// percentiles: 1 means every operation was timed. Recorded so a reader can
	// see how the percentiles were obtained rather than assuming.
	LatencySampleEvery int   `json:"latency_sample_every"`
	LatencySamples     int   `json:"latency_samples"`
	Errors             int64 `json:"errors"`
	// GoMaxProcs and NumCPU are recorded because a scaling number is
	// meaningless without them.
	GoMaxProcs int `json:"gomaxprocs"`
	NumCPU     int `json:"numcpu"`
	// PerfCores and EffCores split NumCPU by core class, and they are recorded
	// because NumCPU ALONE IS MISLEADING on a heterogeneous machine.
	//
	// Measured: on a 4P+6E host, throughput on a read workload peaks at 3-4
	// goroutines and decays thereafter, while NumCPU and GOMAXPROCS both report
	// 10. A reader who takes 10 as the scaling region concludes that something
	// is capping throughput at ~1.5x and goes hunting a lock that does not
	// exist (rmp #2691). The scaling region ends near PerfCores.
	//
	// Zero means the split could not be determined, which is the honest value
	// for a platform that does not report it -- not an assertion of homogeneity.
	PerfCores int `json:"perf_cores"`
	EffCores  int `json:"eff_cores"`
}

Metrics is the effect half of a measurement. It is a plain value: safe to copy and to read concurrently, carrying no internal synchronisation.

func Observe

func Observe(w Workload, level int, win Window, outDir string) (Metrics, error)

Observe runs ONE window of a measurement in the current process and, for WindowProbe, writes that window's profiles into outDir.

One window per process is deliberate: see the package documentation. A caller that wants a Result runs Observe twice, in two fresh children, and pairs the halves.

Observe is NOT safe for concurrent use: for WindowProbe it sets the process-global profiling rates. On every path — success, error, and panic — it restores the mutex profile fraction to the value it found and RESETS the block profile rate to 0. The reset is not a restore: the runtime exposes no getter for the block rate, so a rate the caller had set before calling Observe is lost and must be set again afterwards.

func ReadMetrics

func ReadMetrics(dir string) (Metrics, error)

ReadMetrics reads back what WriteMetrics wrote.

ReadMetrics is safe for concurrent use.

type Op

type Op func(ctx context.Context, worker, iter int) error

Op is one unit of work, executed concurrently by every worker goroutine. It must be safe for concurrent use — that is the property under test. worker is the goroutine's index in [0,level) and iter counts that worker's calls, so an Op can carve a disjoint key space and measure the MECHANISM rather than data conflicts.

type Result

type Result struct {
	Effect Metrics `json:"effect"` // unprofiled: quotable throughput
	Probe  Metrics `json:"probe"`  // profiled: attribution only
	// ProfileDir is the PROBE child's directory, holding mutex.pb.gz,
	// block.pb.gz, cpu.pb.gz and goroutine.pb.gz. The effect child writes no
	// profiles at all.
	ProfileDir string `json:"profile_dir"`
}

Result pairs the two windows of one measurement. Each half is produced by its own child process. Like Metrics it is a plain value carrying no internal synchronisation.

type Window

type Window string

Window selects which half of a measurement to run. The two windows are deliberately run in separate processes; see the package documentation.

const (
	// WindowEffect is the unprofiled window. It writes no profiles, and its
	// throughput and latency are the quotable numbers.
	WindowEffect Window = "effect"
	// WindowProbe is the profiled window. It writes cpu, mutex, block and
	// goroutine profiles, and its throughput is the probe's, never the
	// module's.
	WindowProbe Window = "probe"
)

func ParseWindow

func ParseWindow(s string) (Window, bool)

ParseWindow converts a command-line argument into a Window.

type Workload

type Workload struct {
	// Name is the stable identifier used in filenames and reports.
	Name string
	// Surface names the module packages the workload is meant to drive, so a
	// sweep can state which surfaces it reached and which it did not.
	Surface string
	// Ops is the total number of operations across all workers. It is fixed
	// per workload rather than per level so that every level does the SAME
	// total work: a fraction of a fixed workload in a fixed window is a rate,
	// and comparing rates across levels is the whole point.
	Ops int
	// Setup builds the fixture and returns the operation plus a teardown. dir
	// is a writable directory private to this measurement.
	Setup func(dir string) (op Op, teardown func() error, err error)
}

Workload is one exercise of a module surface.

A Workload is an immutable description and is safe for concurrent use. All mutable state lives in whatever Setup builds, and Setup is called once per measurement window.

func All

func All() []Workload

All is the registry of workloads the sweep drives. Each names the module surface it is meant to reach, so a sweep can state its coverage honestly rather than implying it touched everything.

Op counts are sized so that every measured window lasts of the order of a second at the SLOWEST rung of the ladder. The first counts committed with this package were far smaller, and the sweep for rmp #2679 measured what that cost: at level 8 cypher-read-label-small closed in 18 ms, cypher-write-mem in 30 ms and lpg-neighbours-read in 13 ms. In an 18 ms window the one-shot plan compile is 30.51% of all blocked nanoseconds, so the profile ranked a start-up cost above every steady-state lock; and because a fixed cold cost falls on every rung alike, it dragged scaling_vs_1 towards 1.0 and made the module look as though it scaled better than it does.

All is safe for concurrent use: every call builds a fresh slice of fresh Workload values, so no caller can mutate the registry another caller sees.

func ByName

func ByName(name string) (Workload, bool)

ByName returns the workload with the given name.

ByName is safe for concurrent use: it allocates a fresh Workload per call and shares no state between callers.

func TransportWorkload

func TransportWorkload(kind, query string, level int) (Workload, error)

TransportWorkload builds one arm of the transport A/B.

kind is TransportSim or TransportTCP; query is one of TransportQueryNames. level is the goroutine count the arm will be driven at, and it is a parameter because the arm PRE-ESTABLISHES one connection per worker during Setup, keeping connection cost out of the measured window on both arms.

The returned Workload is an immutable description and is safe for concurrent use; everything mutable lives in what its Setup builds.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL