profile

package
v0.2.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 17, 2026 License: Apache-2.0 Imports: 12 Imported by: 0

Documentation

Overview

Package profile records what tasks actually cost this model, and uses it to recalibrate the ruler the planner sizes tasks with.

The sizing anchors began as a prior — three invented examples of a too-small, a right-sized, and a too-big task — with a note that they should be replaced by measurement once an executor existed. This is that replacement.

It is calibration, not learning, and the difference matters. Nothing here fits a model or predicts a number. It collects what real tasks cost, and when there is enough evidence that the current ruler is wrong, it asks once for three new examples drawn from tasks that actually ran. The output is still three sentences in a prompt; only their provenance changes.

Index

Constants

View Source
const (
	BucketDirect = "direct"
	BucketReflex = "reflex"
	BucketWhole  = "whole"
)

BucketDirect labels work dispatched without a planner size judgment. Keeping it separate lets the compiler learn direct-job costs without teaching the ruler that an unmeasured task was atomic.

BucketWhole is the shape that had no name at all, and its absence was expensive. An undivided goal, a single-leaf remainder and a node spliced in after planning all run as one worker over the whole job and none of them was ever sized, so each was journaled under the empty string — a bucket no reader looks up, which is why the commonest shape in the product could never reach the evidence floor and never acquired an expectation of its own. Worse, the unlabelled rows were still the population the cheapest-leaf floor was taken over, so a one-turn no-op's 7,215 tokens became the advertised price of existing as a leaf. Labelled, the shape accrues its own row and the floor is taken over shapes somebody named.

View Source
const MinSamples = 8

MinSamples is how much evidence is needed before the ruler may be rewritten. Eight leaves is not statistics and this is not pretending to be — it is only enough to tell a systematically wrong anchor from one unlucky task.

Variables

This section is empty.

Functions

func FanInOf

func FanInOf(count int) *int

FanInOf is the pointer a caller writes into a Record.

func Identity

func Identity(model string) string

Identity is the model a history belongs to.

Two layers, and only the first is always there. The free one is the doctrine's own normalisation — the leading "~" is codeaf's routing marker and not part of any slug, and case is not an identity — which is what a surface with no catalog gets. The second is whatever resolver was installed, applied over the first.

The answer is memoised per spelling for the life of the process, because the installed resolver reads a catalog that warms in the background: asked before it lands and again after, it would honestly give two answers, and a file key that moved halfway through a run would split a history inside one process rather than across two. First answer wins; the merge in Load is what makes an early, unresolved answer cost nothing but a merge later.

func UseIdentity

func UseIdentity(resolve Resolve)

UseIdentity installs the catalog-backed resolution for this process. It is the same shape as plan.UseAnchors and is installed at the same place and time: once at launch, by the surface that owns the catalog.

Installing a resolver clears the memo, so a process that installs late still gets one consistent answer from that point on; what was written under the unresolved spelling before it is not lost but adopted, by the merge in Load.

Types

type Profile

type Profile struct {
	Model      string   `json:"model"`
	Subharness string   `json:"subharness"`
	Anchors    string   `json:"anchors,omitempty"` // empty means the built-in prior
	Records    []Record `json:"records"`
	// contains filtered or unexported fields
}

Profile is the accumulated experience of one model running one kind of work.

Keyed by model and worker because capability is a property of the executor, not of the project. There is one worker, so in practice there is one file per model; the key keeps its shape so that a profile written by a build with more than one is simply a file nothing opens.

func Load

func Load(dir, model, subharness string) (*Profile, error)

Load reads the profile for a model and subharness, returning an empty one when there is nothing recorded yet.

The file is keyed on the model's IDENTITY rather than on its spelling (see identity.go). With no resolver installed that is the free normalisation alone, which slug already collapsed, so every existing profile file keeps its exact name; with one installed, two spellings of one model open one history, and the history written under the other spelling is adopted on the first miss.

The result is always a fresh value owning its own records: callers Add to a profile and Save it, so a remembered parse must never become shared mutable state.

func (*Profile) Add

func (p *Profile) Add(records ...Record) []Record

Add appends measurements and returns the records as they were journaled. Each expectation is taken before its record enters the profile, so a leaf can never make its own prediction look better.

func (*Profile) BoundaryEvidence

func (p *Profile) BoundaryEvidence(limit int) []Record

BoundaryEvidence is the newest handful of records that say something about where this worker's edge is.

It is a separate pick from Evidence rather than a fourth band inside it because the two answer different questions. Evidence samples the observed range — cheapest, middle, overran — and a note about fit is not a point on that range: a run that finished in four turns because the job was trivial for this worker sits in the same band as one that finished in four turns because the worker is fast, and only one of them is evidence that the boundary is in the wrong place. Newest first, because a boundary that has already moved once is described by what happened after it moved.

func (*Profile) Evidence

func (p *Profile) Evidence(each int) (small, middle, large []Record)

Evidence picks the tasks worth showing a recalibration call: the cheapest few, the ones nearest the middle, and the ones that ran out of budget. Real examples at each end of the observed range are what a ruler is made of.

func (*Profile) Measure

func (p *Profile) Measure() Spread

Measure summarises the recorded work.

func (*Profile) MeasureReflex

func (p *Profile) MeasureReflex() ReflexStats

MeasureReflex summarises reflex records without admitting them as planner ruler evidence. An unverified completion still counts here: this statistic asks whether the micro-leaf finished, not whether it should rate a model.

func (*Profile) NeedsRecalibration

func (p *Profile) NeedsRecalibration() (bool, string)

NeedsRecalibration decides whether the ruler is worth rewriting.

Two guards, both about not chasing noise. There must be enough samples to see a pattern, and the pattern must actually contradict the current ruler — either tasks are routinely exhausting their budget, which means the anchors are letting too much into one node, or almost nothing ever comes close, which means they are splitting work that did not need splitting.

func (*Profile) RecordsSnapshot

func (p *Profile) RecordsSnapshot() ([]Record, time.Time)

RecordsSnapshot returns a stable copy of the measurements and the profile file's last modification time. Derived views use the latter only for legacy records written before per-record timestamps existed.

func (*Profile) Save

func (p *Profile) Save() error

Save writes the profile back.

type Record

type Record struct {
	Title   string `json:"title"`
	Summary string `json:"summary"`
	// Time is when this observation landed. Older profiles omit it; readers
	// may use the profile file's modification time as a coarse fallback.
	Time time.Time `json:"time,omitempty"`
	// Sources is the touch-list: how many distinct things this leaf had to
	// visit to finish. It is a size signal and it is NOT a fan-in — a node that
	// names four datasets and waits on nobody has four sources and no inputs.
	// Reading it as one is what priced reassembly at the price of an atomic.
	Sources int `json:"sources"`
	// SourcesKnown distinguishes an observed zero from old and direct records
	// whose omitted source count decoded to zero.
	SourcesKnown bool `json:"sources_known,omitempty"`
	// FanIn is how many earlier results actually landed in this leaf: its
	// settled dependency count, measured rather than inferred. It is the number
	// a join is priced from, and it is a different question from Sources — the
	// one the fan-out prompt has never been shown, because until this field
	// existed nothing recorded it.
	//
	// Nil is "nobody counted", which every record written before this field
	// existed is, and which the join price is required to skip: counting an
	// unmeasured leaf as a zero-input one would price a join from leaves that
	// never made one.
	FanIn    *int    `json:"fan_in,omitempty"`
	Size     string  `json:"size"`  // planner prediction, or direct when none was made
	Turns    int     `json:"turns"` // what it actually took
	Tokens   int     `json:"tokens"`
	Stop     string  `json:"stop"`
	Cost     float64 `json:"cost,omitempty"`
	Promoted bool    `json:"promoted,omitempty"`

	// ExpectedTurns and ExpectedTokens are the profile's own medians for this
	// size before the record landed. Nil means the bucket had too little prior
	// evidence to make an expectation; that is distinct from a zero residual.
	ExpectedTurns  *int     `json:"expected_turns,omitempty"`
	ExpectedTokens *int     `json:"expected_tokens,omitempty"`
	Surprise       *float64 `json:"surprise,omitempty"`

	// Calibration is what the worker said about its own fit for this task —
	// exec.Outcome.Calibration, journaled. It is free text and it is only ever
	// read by a model: the recalibration call renders it beside the turns and
	// tokens, so the three anchor examples get rewritten from what the work felt
	// like as well as from what it cost. Nil on every record the generalist has
	// ever written, which is why every existing profile file and every existing
	// recalibration prompt is byte-identical.
	Calibration []string `json:"calibration,omitempty"`

	// Verdict is how the leaf actually ended. It replaced a `done` flag that was
	// the scheduler's StateDone carried across — true of a leaf that exhausted
	// its budget mid-edit as much as of one that finished — and the flag was
	// never read, because a ruler calibrated against it would have been
	// calibrated against the budget rather than against the work.
	Verdict provider.Reading `json:"verdict,omitempty"`
}

Record is one executed leaf.

func (Record) Boundary

func (r Record) Boundary() bool

Boundary reports whether this record says anything about where the edge of this worker's capacity is: a note the worker wrote about its own fit.

func (Record) FanInCount

func (r Record) FanInCount() (int, bool)

FanInCount is the measured inbound dependency count and whether anybody measured it. A record from before fan-in was recorded answers false, and a caller pricing a join has to skip it rather than read it as zero.

func (Record) HasSourceCount

func (r Record) HasSourceCount() bool

HasSourceCount keeps nonzero counts from legacy profiles usable while treating their indistinguishable zero value as unknown.

func (Record) Labelled

func (r Record) Labelled() bool

Labelled reports that somebody named this record's shape. An empty size is a record written by a surface that did not say, and it is evidence about no shape in particular — see BucketWhole for why there is now no reason to write one.

func (Record) Overran

func (r Record) Overran() bool

Overran reports a task that could not finish inside its budget — the clearest evidence that the ruler let too much into one node.

The verdict is the authority where there is one; the stop reason is read for records written before verdicts existed, so an old profile still calibrates.

type ReflexStats

type ReflexStats struct {
	Samples      int
	Successes    int
	Promotions   int
	MedianTurns  int
	MedianTokens int
	AverageCost  float64
}

ReflexStats is the measured boundary between work that finishes as one quick action and work that promotes into the compiled path.

type Resolve

type Resolve func(model string) string

Resolve turns one spelling of a model into the identity its history is kept under. It must be pure, cheap and non-blocking: it is called on the launch path, and a resolver that waited on a fetch would put that fetch in front of the first frame of a run.

type Spread

type Spread struct {
	Samples  int
	MinTurns int
	Median   int
	MaxTurns int
	Overran  int
	MedianKb int
}

Spread is what the profile knows, and it deliberately reports range rather than a single number. Identical per-city briefs came back at 9, 16 and 25 turns in one run; a median alone would encode a precision that is not there.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL