Documentation
¶
Overview ¶
Package evaluate scores csf builds on a held-out, rotated suite of past tickets before they go live (EVAL-SUITE, #416). A suite ticket is replayed from the commit its pull request started from, on the build being scored, and read at a fixed budget of tool calls by the loop's own struggle definition (ouroboros.StrugglesWithin); the merged change is the reference its files are compared with. The suite is hidden from the miners and the loop's corpus readers, rotates weekly, and never holds a ticket the loop evolves the harness on.
Index ¶
- Constants
- Variables
- func CheckClaims(ctx context.Context, body string, recompute Recompute) error
- func Citation(score Score, versus Score) (string, error)
- func Claims(body string) []string
- func Disjoint(suite Suite, evolution []int64) error
- func HiddenAssignments(corpus iofs.IFiles, hidden map[int64]bool) (map[string]bool, error)
- func Read(events []byte, budget int) (toolCalls int64, episodes int64)
- func Recall(reference []string, touched []string) float64
- func TicketOf(corpus iofs.IFiles, assignment string) (int64, error)
- func TicketOfTitle(title string) int64
- type Cited
- type Collector
- type Comparison
- type Derivation
- type Host
- type Job
- type MergedPull
- type Recompute
- type Replay
- type Replayer
- type ReplayerOption
- func WithClock(source clock.IClock, poll time.Duration) ReplayerOption
- func WithHost(host Host) ReplayerOption
- func WithLauncher(launcher proc.ILauncher) ReplayerOption
- func WithLogger(logger *slog.Logger) ReplayerOption
- func WithParallel(parallel int) ReplayerOption
- func WithState(directory string, files iofs.IFiles) ReplayerOption
- type Run
- type Score
- type Suite
- type SuiteStore
- func (store *SuiteStore) HiddenTickets(ctx context.Context) (map[int64]bool, error)
- func (store *SuiteStore) LatestSuite(ctx context.Context) (Suite, error)
- func (store *SuiteStore) Recompute(ctx context.Context, build string, version int, versus string) (string, error)
- func (store *SuiteStore) RecordReplay(ctx context.Context, replay Replay) error
- func (store *SuiteStore) RecordRun(ctx context.Context, run Run) error
- func (store *SuiteStore) RecordSuite(ctx context.Context, suite Suite) error
- func (store *SuiteStore) Replays(ctx context.Context, build string, version int) ([]Replay, error)
- func (store *SuiteStore) Runs(ctx context.Context) ([]Run, error)
- func (store *SuiteStore) ScoreOn(ctx context.Context, build string, suite Suite) (Score, error)
- func (store *SuiteStore) Scores(ctx context.Context) ([]Score, error)
- func (store *SuiteStore) Suites(ctx context.Context) ([]Suite, error)
- type Ticket
Constants ¶
const ( NodeHost = "host" NodeBurst = "burst" )
The node kinds a replay runs on: this host, or a cloud burst job.
const ( MetricScore = "csf_eval_struggle_rate_per_1k_tool_calls" MetricWallSeconds = "csf_eval_score_wall_seconds" MetricCostUSD = "csf_eval_score_cost_usd" MetricShards = "csf_eval_shards" )
The suite's series, read from the records at each scrape; the CSF dashboard's Evaluation suite row draws them.
const ( // RotationPeriod is how long one suite version stands: the struggle // rate's own reporting period, a week, so every weekly reading of the // headline is on one fixed suite. RotationPeriod = 7 * 24 * time.Hour )
Variables ¶
var ( // ErrUncited reports a pull request body that claims an improvement and // carries no citation matching a recorded score. ErrUncited = errors.New("evaluate: an improvement is claimed without a suite citation that matches a recorded score") // ErrNoRate reports a score with no tool call, which has nothing to cite. ErrNoRate = errors.New("evaluate: the score has no tool call to rate") )
var ( // ErrInvalidOption reports a nil option or one the replayer cannot use. ErrInvalidOption = errors.New("evaluate: invalid option") // ErrNoOriginal reports a suite ticket whose original run left no recipe. ErrNoOriginal = errors.New("evaluate: the ticket's original run has no recipe to replay") )
var ( // ErrUnscored reports a build with no complete score on the suite. ErrUnscored = errors.New("evaluate: the build has no complete score on the current suite") // ErrRegressed reports a build whose score is worse than the live // build's beyond the bound. ErrRegressed = errors.New("evaluate: the build regressed beyond the bound") )
var ( // ErrDatabaseRequired reports a store built without the database. ErrDatabaseRequired = errors.New("evaluate: the csfpg database is required") // ErrNoSuite reports a read of the suite before any was selected. ErrNoSuite = errors.New("evaluate: no suite has been selected; run csf eval suite") )
var ( // ErrEmptyPool reports a selection with no eligible ticket. ErrEmptyPool = errors.New("evaluate: no eligible ticket for the suite") // ErrNotDisjoint reports a suite holding a ticket the loop evolves the // harness on. ErrNotDisjoint = errors.New("evaluate: the suite and the evolution tickets overlap") )
var ReplayTools = []string{
"Read", "Write", "Edit", "Glob", "Grep", "TodoWrite",
"Bash(git:*)",
"Bash(gh issue view:*)", "Bash(gh issue list:*)", "Bash(gh pr view:*)", "Bash(gh pr list:*)", "Bash(gh pr diff:*)", "Bash(gh run list:*)", "Bash(gh run view:*)",
"Bash(tools/bazel.sh:*)", "Bash(bash tools/check-house-lint.sh:*)", "Bash(bash tools/ontology-score.sh:*)",
"Bash(bash tools/check-merge.sh:*)", "Bash(bash tools/check-generated.sh:*)", "Bash(python3 tools/check_operator_identifiers.py:*)",
"Bash(docker run:*)", "Bash(docker build:*)",
"Bash(ls:*)", "Bash(cat:*)", "Bash(head:*)", "Bash(tail:*)", "Bash(wc:*)", "Bash(grep:*)", "Bash(rg:*)", "Bash(find:*)",
"Bash(sed:*)", "Bash(awk:*)", "Bash(sort:*)", "Bash(uniq:*)", "Bash(diff:*)", "Bash(jq:*)", "Bash(stat:*)", "Bash(tree:*)",
"Bash(echo:*)", "Bash(printf:*)", "Bash(pwd)", "Bash(mkdir:*)", "Bash(cp:*)", "Bash(mv:*)", "Bash(touch:*)", "Bash(gofmt:*)",
}
ReplayTools are the tool rules a replay may use: the files, all of git (the replay repository's origin is its own local copy, so even a push stays there), read-only gh, the pinned build, the repository's own scripts and the read-only shell tools. Nothing that reaches past the replay is allowed (gh writes, csf send, merge or submit, the csf MCP tools, an arbitrary bash -c), because a replay re-runs a real brief that may say to do exactly that. A refused call is a struggle signal on every build alike, so scores stay comparable.
Functions ¶
func CheckClaims ¶
CheckClaims is the proof check (#416 Build 5): a body that claims an improvement must carry at least one citation whose line equals the one the records give. A body that claims nothing passes.
func Citation ¶
Citation is the line a pull request cites score with, measured against the score of build vs on the same suite version.
func Disjoint ¶
Disjoint checks that no suite ticket is one the loop evolves the harness on (#416 Build 7).
func HiddenAssignments ¶
HiddenAssignments lists the run directories under corpus that worked a hidden ticket: the original runs of suite tickets and every replay of them. The corpus readers skip these, so a miner run never receives a suite ticket's event log.
func Read ¶
Read is what one replay's event log contributes at budget: its tool calls and struggle episodes by the loop's definition.
func Recall ¶
Recall is the share of the reference's files the replay's change touched; a reference with no files is fully recalled by nothing, so 0.
func TicketOfTitle ¶
TicketOfTitle reads the ticket a pull request title names, 0 for none.
Types ¶
type Collector ¶
type Collector struct {
// contains filtered or unexported fields
}
Collector measures every recorded score at each scrape.
func NewSuiteCollector ¶
NewSuiteCollector measures the scores scores returns.
func (*Collector) Collect ¶
func (collector *Collector) Collect(metrics chan<- prometheus.Metric)
Collect reads the scores and sends every series; a record that cannot be read sends none.
func (*Collector) Describe ¶
func (collector *Collector) Describe(descriptions chan<- *prometheus.Desc)
Describe sends every series' description.
type Comparison ¶
type Comparison struct {
Delta float64 `json:"delta_per_1k"`
Bound float64 `json:"bound_per_1k"`
Paired int `json:"paired_tickets"`
}
Comparison is a candidate build against the live one on the same suite version: the difference of the pooled scores and the bound it is judged by, from the suite's own run-to-run variance (the spread of the per-ticket differences of the two builds' replays).
func Admit ¶
func Admit(suite Suite, candidate []Replay, live []Replay, candidateBuild string, liveBuild string) (*Comparison, error)
Admit is the score-before-live check: the candidate must have a complete score on the suite, and when the live build has one too, the candidate must not be worse beyond the bound. It returns the comparison it judged by, nil when the live build had no complete score to compare with.
func Compare ¶
func Compare(candidate []Replay, live []Replay) Comparison
Compare judges candidate against live: delta is candidate minus live (negative is better) and bound is z95 times the standard error of the mean per-ticket difference. With fewer than two paired tickets there is no variance to bound by, and the bound is infinite.
type Derivation ¶
type Derivation struct {
Pool int `json:"pool"`
Evolution int `json:"evolution_excluded"`
Budget int `json:"budget_tool_calls"`
BudgetFrom string `json:"budget_from"`
Rate float64 `json:"pool_rate_per_1k"`
Dispersion float64 `json:"dispersion"`
Effect float64 `json:"effect_log"`
Needed int `json:"replays_needed"`
Size int `json:"size"`
SizeFrom string `json:"size_from"`
Resolves float64 `json:"resolves_factor"`
}
Derivation records how a suite's size and budget follow from the pool: every number a reader needs to redo the arithmetic.
type Host ¶
type Host struct {
Admit func(ctx context.Context, request *harnessv1.CheckAgentSessionAdmissionRequest) (*harnessv1.CheckAgentSessionAdmissionResponse, error)
Submit func(ctx context.Context, request *harnessv1.SubmitAgentSessionRequest) (*harnessv1.SubmitAgentSessionResponse, error)
Get func(ctx context.Context, request *harnessv1.GetAgentSessionRequest) (*harnessv1.GetAgentSessionResponse, error)
Cancel func(ctx context.Context, request *harnessv1.CancelAgentSessionRequest) (*harnessv1.CancelAgentSessionResponse, error)
}
Host is the harness host a replay runs on, as the four operations the replayer calls: the generated client's methods, or the in-process service's.
type Job ¶
type Job struct {
Ticket Ticket `json:"ticket"`
Recipe json.RawMessage `json:"recipe"`
}
Job is one suite ticket ready to replay: the ticket and the replay recipe, encoded as protojson so a job crosses to a burst node as data.
type MergedPull ¶
type MergedPull struct {
Number int64
Title string
MergeCommit string
BaseCommit string
Files []string
}
MergedPull is one merged pull request as the pool reads it: its title, merge commit, the commit it started from (the merge commit's first parent) and the files it changed.
type Recompute ¶
Recompute is the citation line the records give for a build on a suite version measured against build vs, as Citation writes it.
type Replay ¶
type Replay struct {
Build string `json:"build"`
SuiteVersion int `json:"suite_version"`
Ticket int64 `json:"ticket"`
Node string `json:"node"`
Assignment string `json:"assignment"`
ToolCalls int64 `json:"tool_calls"`
Episodes int64 `json:"episodes"`
Recall float64 `json:"recall"`
CostUSDMicros int64 `json:"cost_usd_micros"`
Seconds int64 `json:"seconds"`
RecordedAt time.Time `json:"recorded_at"`
}
Replay is one suite ticket replayed on one build, read at the suite's budget: its tool calls and struggle episodes, the share of the reference's files it touched, what it cost and how long it ran.
type Replayer ¶
type Replayer struct {
// contains filtered or unexported fields
}
Replayer runs replays on one harness host: it submits each job, reads the replay's event log as it grows, cancels the session once the suite's budget of tool calls is spent, and records what the replay read.
func NewReplayer ¶
func NewReplayer(options ...ReplayerOption) (*Replayer, error)
NewReplayer validates the whole option set before building the replayer.
func (*Replayer) Run ¶
func (replayer *Replayer) Run(ctx context.Context, build string, suite Suite, node string, jobs []Job, record func(ctx context.Context, replay Replay) error) error
Run replays jobs on the host for build and suite at the suite's budget and hands each finished replay to record, tagged with node.
type ReplayerOption ¶
ReplayerOption configures a Replayer.
func WithClock ¶
func WithClock(source clock.IClock, poll time.Duration) ReplayerOption
WithClock grants the clock the replayer paces its reads by. Required.
func WithHost ¶
func WithHost(host Host) ReplayerOption
WithHost grants the harness host the replays run on. Required.
func WithLauncher ¶
func WithLauncher(launcher proc.ILauncher) ReplayerOption
WithLauncher grants the process capability git reads a replay's change through. Required.
func WithLogger ¶
func WithLogger(logger *slog.Logger) ReplayerOption
WithLogger grants the logger progress goes to.
func WithParallel ¶
func WithParallel(parallel int) ReplayerOption
WithParallel bounds the replays running at once on the host; the host's own admission bounds them further.
type Score ¶
type Score struct {
Build string `json:"build"`
SuiteVersion int `json:"suite_version"`
Replays int `json:"replays"`
Expected int `json:"expected"`
Complete bool `json:"complete"`
ToolCalls int64 `json:"tool_calls"`
Episodes int64 `json:"episodes"`
PerK *float64 `json:"per_1k"`
Low *float64 `json:"low"`
High *float64 `json:"high"`
Recall float64 `json:"recall"`
CostUSD float64 `json:"cost_usd"`
WallSeconds float64 `json:"wall_seconds"`
Nodes map[string]int `json:"nodes"`
}
Score is a build's result on one suite version: struggle episodes per 1,000 tool calls over its replays with the 95% Garwood interval (the headline), the mean recall of the reference, what scoring cost and how long it took, and the replays per node kind.
type Suite ¶
type Suite struct {
Version int `json:"version"`
SelectedAt time.Time `json:"selected_at"`
// Model is the model every replay of this version runs on, so two
// builds' scores differ by the build alone.
Model string `json:"model"`
// Tools are the tool rules every replay of this version may use; empty
// in a record made before they were recorded, which means ReplayTools.
Tools []string `json:"tools,omitempty"`
RotatesAt time.Time `json:"rotates_at"`
Tickets []Ticket `json:"tickets"`
EvolutionSet []int64 `json:"evolution_tickets"`
Derivation Derivation `json:"derivation"`
}
Suite is one version of the held-out suite.
func Select ¶
func Select(pool []Ticket, evolution []int64, version int, model string, now time.Time) (Suite, error)
Select draws suite version from the pool: the tickets the loop evolves the harness on are removed, the budget and size are derived from the rest, and the held-out tickets are the size first by a hash of the version and the ticket, so a version always draws the same tickets and the next draws afresh.
type SuiteStore ¶
type SuiteStore struct {
// contains filtered or unexported fields
}
SuiteStore keeps the suites, replays and scoring runs in CSF's PostgreSQL schema (the csf_eval_* tables), the one place results from every node are aggregated. It borrows the capability and never closes it.
func NewSuiteStore ¶
func NewSuiteStore(database csfpg.IDB) (*SuiteStore, error)
NewSuiteStore returns the store over a pool the binary opened through ipc/db/csfpg, or over pgmem's IDB in a spec.
func (*SuiteStore) HiddenTickets ¶
HiddenTickets is every ticket any suite version has held: once held out, a ticket stays hidden from the miners and the loop's corpus readers.
func (*SuiteStore) LatestSuite ¶
func (store *SuiteStore) LatestSuite(ctx context.Context) (Suite, error)
LatestSuite reads the current suite version.
func (*SuiteStore) Recompute ¶
func (store *SuiteStore) Recompute(ctx context.Context, build string, version int, versus string) (string, error)
Recompute gives the citation line the records hold for build on a suite version against build vs: the Recompute the proof check runs.
func (*SuiteStore) RecordReplay ¶
func (store *SuiteStore) RecordReplay(ctx context.Context, replay Replay) error
RecordReplay records one replay, replacing an earlier reading of the same build, version and ticket.
func (*SuiteStore) RecordRun ¶
func (store *SuiteStore) RecordRun(ctx context.Context, run Run) error
RecordRun records a scoring run's wall clock.
func (*SuiteStore) RecordSuite ¶
func (store *SuiteStore) RecordSuite(ctx context.Context, suite Suite) error
RecordSuite records a new suite version.
func (*SuiteStore) Runs ¶
func (store *SuiteStore) Runs(ctx context.Context) ([]Run, error)
Runs reads every scoring run, oldest first.
func (*SuiteStore) ScoreOn ¶
ScoreOn is a build's score on suite from the records, with its scoring run's wall clock when one was recorded.
type Ticket ¶
type Ticket struct {
Number int64 `json:"ticket"`
PullRequest int64 `json:"pull_request"`
BaseCommit string `json:"base_commit"`
MergeCommit string `json:"merge_commit"`
Files []string `json:"files"`
Assignment string `json:"assignment"`
// ToolCalls and Episodes are the original run's whole reading, used to
// derive the budget and the size; they are not a score.
ToolCalls int64 `json:"tool_calls"`
Episodes int64 `json:"episodes"`
}
Ticket is one past ticket in the pool or the suite: the pull request that closed it, the commit that pull request started from, its merge commit, the files the merged change touched (the reference) and the original run whose recipe a replay reuses, with how that run read at the suite budget.
func Pool ¶
func Pool(corpus iofs.IFiles, pulls []MergedPull) ([]Ticket, error)
Pool is every past ticket a suite may hold: a merged pull request that names its ticket, and a run under corpus that worked the ticket and left its recipe, the one with the most tool calls being the original. A ticket closed by several pull requests is held once, by its last.