eval

package
v1.55.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 11, 2026 License: MIT Imports: 30 Imported by: 0

Documentation

Overview

Package eval implements the local-only review evaluation toolkit.

It deliberately owns a separate registry under <NM_HOME>/eval, so opening the normal pipeline database never creates an eval table or runs an eval migration, and nothing here emits telemetry or reaches the network.

The dependency runs one way: the daemon calls AutoCapture when a run finishes (see RunManager.autoCaptureEvalCase) and RelabelRun when a source PR merges (see RunManager.relabelEvalRun). This package never calls back into the daemon, alters a gate, or influences a pipeline decision.

Index

Constants

View Source
const (
	GoldTruePositive  = "true-positive"
	GoldFalseNegative = "false-negative"
	GoldFalsePositive = "false-positive"
)

Gold kinds are the scientific labels written from the recorded gate decision (see goldFromRound). Capture writes true-positive gold from a user Fix or from an auto-fix selection on a merged run, false-negative gold from a user-added finding, and false-positive gold from a finding the human did not select for fix that shipped in a merged PR. IngestPostPRMiss writes additional false-negative gold after a green review. Adjudicated labels are never inferred from pending unmatched candidate findings.

Variables

View Source
var ErrNoCapturableReview = errors.New("run has no capturable review pass")

ErrNoCapturableReview marks the outcomes where a run simply holds nothing to freeze - no review step, no finished pass, a decision the human has not made yet, or rounds recorded before this machine started keeping replay provenance. These are ordinary states of a healthy pipeline, not failures, so automatic collection can pass over them silently while still surfacing a real capture fault. Every one of them is a deliberate refusal to invent a case: grading a candidate against a half-recorded review pass would be worse than having no case at all.

View Source
var ErrReviewDidNotPassGreen = errors.New("review did not pass green")

ErrReviewDidNotPassGreen is returned when ingest is asked to label a run whose review did not complete with a non-blocking (green) pass. A parked or blocking review is a different class: it found something, so it is not a post-PR miss.

Functions

func RenderReport

func RenderReport(reports []CandidateReport) string

RenderReport is a stable human-readable local comparison. Scores are finding-level. Unmatched candidate findings stay pending and are never called false positives. A replay with no gold is unlabeled, not a pass.

func RenderSets

func RenderSets(summaries []SetSummary) string

RenderSets is a stable human-readable preflight. It intentionally contains counts and buckets only, never source paths, URLs, diffs, or findings.

func Replay

func Replay(ctx context.Context, store *Store, opts ReplayOptions) (Session, []Evaluation, error)

Replay runs exactly the captured review pass. It does not start a daemon or use the production NM_HOME: every case is restored into a fresh temp gate and worktree. Push, PR, CI, and all fix loops are intentionally absent from the MVP subject under test.

Types

type AutoCaptureResult

type AutoCaptureResult struct {
	Captured int
	Pruned   int
	Skipped  bool
	Reason   string
}

AutoCaptureResult reports what one automatic collection pass did. Skipped is true when the run simply had nothing to freeze, which is the ordinary outcome for most runs and is not a failure.

func AutoCapture

func AutoCapture(ctx context.Context, p *paths.Paths, database *db.DB, runID string, maxCases int) (AutoCaptureResult, error)

AutoCapture freezes one finished run's review passes into the local corpus and then enforces the retention cap.

It is the single entry point for collection that nobody asked for by hand, so it deliberately keeps the same Capture the CLI uses rather than a looser variant: an automatically collected case has to be exactly as trustworthy as one a person captured, or the corpus quietly becomes a different thing from what the eval commands report on.

The caller owns the timeout and the decision to run at all. This function owns nothing of the pipeline: it opens its own registry, does its work, and closes it, so a failure here cannot reach the run that triggered it.

type BaselineMetrics

type BaselineMetrics struct {
	InputTokens      int64 `json:"input_tokens"`
	OutputTokens     int64 `json:"output_tokens"`
	CacheReadTokens  int64 `json:"cache_read_tokens"`
	FreshInputTokens int64 `json:"fresh_input_tokens"`
	DurationMS       int64 `json:"duration_ms"`
	TokensReported   bool  `json:"tokens_reported"`
}

BaselineMetrics is the recorded source-review performance baseline. A false TokensReported means the adapter did not surface trustworthy token usage.

type Candidate

type Candidate struct {
	Agent types.AgentName `json:"agent"`
	Model string          `json:"model"`
}

Candidate identifies one agent and model combination under evaluation. The canonical command-line spelling is agent+model, for example codex+gpt-5.4.

func ParseCandidate

func ParseCandidate(raw string) (Candidate, error)

ParseCandidate accepts exactly agent+model. Keeping the model explicit makes comparison records self-describing rather than silently inheriting a user's current agent default.

func (Candidate) String

func (c Candidate) String() string

type CandidateReport

type CandidateReport struct {
	Cohort        string
	Summary       EvaluationSummary
	RepeatCount   int
	Confidence    *Interval
	AverageTokens *float64
	AverageWallMS float64
	OnFrontier    bool
}

CandidateReport is one locally observed candidate slice.

func Report

func Report(store *Store) ([]CandidateReport, error)

Report loads every local evaluation result grouped by candidate. It never contacts a forge, agent provider, telemetry endpoint, or remote case store.

type Case

type Case struct {
	Manifest
	Labels   Labels          `json:"labels"`
	Decision Decision        `json:"decision"`
	Baseline BaselineMetrics `json:"baseline"`
	Dir      string          `json:"-"`
}

Case is one frozen review pass. Dir is local bookkeeping and never enters a manifest or report payload.

func Capture

func Capture(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, runID string) ([]Case, error)

Capture exports every persisted review pass from one real run. It reads the production state only; it never starts the daemon, changes a gate, fetches, or sends data anywhere.

func RelabelAll

func RelabelAll(ctx context.Context, store *Store, p *paths.Paths, database *db.DB) ([]Case, error)

func RelabelRun

func RelabelRun(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, runID string) ([]Case, error)

RelabelRun recomputes gold for every captured case of a run. It adds new auto-fix-merged / shipped-unfixed labels onto previously unlabeled findings and drops obsolete derived merge labels that the current rounds no longer support. It never overwrites adjudicated, user-fix, or ingested post-PR-miss labels. Missing cases are a no-op so a merge observed before the first capture cannot fail the pipeline.

type Decision

type Decision struct {
	Action             string   `json:"action"`
	SelectionSource    string   `json:"selection_source,omitempty"`
	SelectedFindingIDs []string `json:"selected_finding_ids,omitempty"`
	HasUserFindings    bool     `json:"has_user_findings"`
}

Decision records the human gate evidence available for the exported review pass. Approval actions were not persisted in historical rows, so Action can be "unknown". The original selections themselves are never guessed, and an unknown action supports no label at all (see hasRecordedDecision).

type Evaluation

type Evaluation struct {
	ID                string `json:"id"`
	SessionID         string `json:"session_id"`
	CaseID            string `json:"case_id"`
	Candidate         string `json:"candidate"`
	Cohort            string `json:"cohort"`
	Repeat            int    `json:"repeat"`
	StartedAt         int64  `json:"started_at"`
	CompletedAt       int64  `json:"completed_at"`
	Status            string `json:"status"`
	Error             string `json:"error,omitempty"`
	HasFindingGold    bool   `json:"has_finding_gold"`
	GoldCount         int    `json:"gold_count"`
	TruePositive      int    `json:"true_positive"`
	TruePositiveExact int    `json:"true_positive_exact,omitempty"`
	TruePositiveFuzzy int    `json:"true_positive_fuzzy,omitempty"`
	FalseNegative     int    `json:"false_negative"`
	FalsePositive     int    `json:"false_positive"`
	FalsePositiveGold int    `json:"false_positive_gold,omitempty"`
	Pending           int    `json:"pending"`
	FindingsJSON      string `json:"findings_json,omitempty"`
	FindingCount      int    `json:"finding_count"`
	InputTokens       int64  `json:"input_tokens"`
	OutputTokens      int64  `json:"output_tokens"`
	CacheReadTokens   int64  `json:"cache_read_tokens"`
	FreshInputTokens  int64  `json:"fresh_input_tokens"`
	TokensReported    bool   `json:"tokens_reported"`
	DurationMS        int64  `json:"duration_ms"`
	Model             string `json:"model,omitempty"`
}

Evaluation is one candidate replay over one case. Status is "completed" or "failed"; failures remain visible in reports and are not silently scored. Confusion-matrix fields are finding-level: unmatched candidate findings stay in Pending and are never treated as false positives.

type EvaluationSummary

type EvaluationSummary struct {
	Candidate         string
	Total             int
	Labeled           int
	TruePositive      int
	TruePositiveExact int
	TruePositiveFuzzy int
	FalseNegative     int
	FalsePositive     int
	FalsePositiveGold int
	Pending           int
	Failures          int
	InputTokens       int64
	OutputTokens      int64
	FreshInputTokens  int64
	TokensReported    int
	DurationMS        int64
}

EvaluationSummary aggregates finding-level scores. A case with no gold is unlabeled / pending, never a pass. Unmatched candidate findings stay in Pending and do not become false positives.

func SummarizeEvaluations

func SummarizeEvaluations(evaluations []Evaluation) EvaluationSummary

SummarizeEvaluations scores finding-level gold only. Unmatched candidate findings stay pending. A replay with no gold is unlabeled, not a pass.

func (EvaluationSummary) ExactRecall

func (s EvaluationSummary) ExactRecall() float64

func (EvaluationSummary) F1

func (s EvaluationSummary) F1() float64

func (EvaluationSummary) F1Conservative

func (s EvaluationSummary) F1Conservative() float64

func (EvaluationSummary) HasFalsePositiveGold

func (s EvaluationSummary) HasFalsePositiveGold() bool

func (EvaluationSummary) Precision

func (s EvaluationSummary) Precision() float64

func (EvaluationSummary) PrecisionLower

func (s EvaluationSummary) PrecisionLower() float64

func (EvaluationSummary) Recall

func (s EvaluationSummary) Recall() float64

type FindingGold

type FindingGold struct {
	ID          string `json:"id"`
	Kind        string `json:"kind"`
	Source      string `json:"source,omitempty"`
	File        string `json:"file,omitempty"`
	Line        int    `json:"line,omitempty"`
	Description string `json:"description,omitempty"`
	Severity    string `json:"severity,omitempty"`
	Action      string `json:"action,omitempty"`
}

FindingGold is one finding-level gold label for an underlying issue. Capture writes only the labels goldFromRound can support from recorded evidence; unmatched later candidate findings stay unlabeled. IngestPostPRMiss writes additional false-negative gold for a confirmed miss after a green review.

func ParsePostPRMissFinding

func ParsePostPRMissFinding(raw string) (FindingGold, error)

ParsePostPRMissFinding accepts one finding object. ID and description are required; they are the eval matcher keys. This is the typed source of truth for a confirmed post-PR miss. no-mistakes does not read firstmate ledgers or scrape GitHub review comments.

type IngestResult

type IngestResult struct {
	CaseID string
	Added  int
	Total  int
}

IngestResult reports which captured case received post-PR-miss gold.

func IngestPostPRMiss

func IngestPostPRMiss(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, runID string, misses []FindingGold) (IngestResult, error)

IngestPostPRMiss captures a run that already passed review green, then writes confirmed post-PR misses as false-negative gold on the last green review pass. Capture of an existing case is a no-op, so later ingest still attaches gold. Duplicate finding IDs are no-ops.

type Interval

type Interval struct {
	Lower float64
	Upper float64
	Cases int
}

Interval is a finite-sample recall range over cases. Repeats are averaged inside each case so a noisy provider does not inflate apparent sample size.

type Labels

type Labels struct {
	Version                 int           `json:"version"`
	Findings                []FindingGold `json:"findings,omitempty"`
	QueuedCandidateFindings int           `json:"queued_candidate_findings"`
}

Labels is a local, growing label file. Finding-level gold is the unit of truth. Queued candidate findings are kept as evidence for later adjudication and are never scored as false positives.

func (Labels) FalsePositiveCount

func (l Labels) FalsePositiveCount() int

func (Labels) HasGold

func (l Labels) HasGold() bool

func (Labels) TrueIssueCount

func (l Labels) TrueIssueCount() int

type Manifest

type Manifest struct {
	Version          int    `json:"version"`
	ID               string `json:"id"`
	SourceRunID      string `json:"source_run_id"`
	SourceRoundID    string `json:"source_round_id"`
	CapturedAt       int64  `json:"captured_at"`
	RepoFingerprint  string `json:"repo_fingerprint"`
	Branch           string `json:"branch"`
	DefaultBranch    string `json:"default_branch"`
	BaseSHA          string `json:"base_sha"`
	HeadSHA          string `json:"head_sha"`
	StartingHeadSHA  string `json:"starting_head_sha"`
	ReviewedHeadSHA  string `json:"reviewed_head_sha"`
	TrustedConfigSHA string `json:"trusted_config_sha"`
	Intent           string `json:"intent,omitempty"`
	IntentSource     string `json:"intent_source,omitempty"`
	VersionAtCapture string `json:"no_mistakes_version,omitempty"`
	BuildSHA         string `json:"no_mistakes_build_sha,omitempty"`
	ChangedFiles     int    `json:"changed_files"`
	ChangedLines     int    `json:"changed_lines"`
}

Manifest pins every input needed to recreate a review pass without storing a remote URL. The commits it names live in this repository's local object pool and the configuration it replays under sits beside it in the case directory. It carries no digest of those objects: Git object names are content hashes, so the pins are their own integrity check.

type ReplayOptions

type ReplayOptions struct {
	Set       string
	Candidate Candidate
	Repeats   int
}

ReplayOptions controls one isolated candidate comparison.

type Score

type Score struct {
	TruePositive      int
	TruePositiveExact int
	TruePositiveFuzzy int
	FalseNegative     int
	FalsePositive     int
	FalsePositiveGold int
	Pending           int
}

Score is one candidate's finding-level confusion matrix against gold. Pending is unmatched candidate findings: queued, never punished as FP.

func ScoreCandidate

func ScoreCandidate(labels Labels, findingsJSON string) Score

ScoreCandidate matches a candidate finding list against recorded gold.

  • TP: the candidate raises the same underlying issue as a true-issue gold (human-accepted Fix, auto-fix that landed in a merged PR, a human-added miss, or a confirmed post-PR miss)
  • FN: the candidate misses a true-issue gold
  • FP: only an explicit false-positive gold that the candidate still raised
  • Pending: unmatched candidate findings, never inferred as invalid

Matching is a documented cascade of strengths: exact-id, exact-text, nearby-line Jaccard, then gated containment. Assignment is one globally optimal assignment over the whole graph (see assignMatches), so neither candidate ordering nor a tier boundary can consume a candidate another gold needed. Headline recall uses the full cascade; exact vs fuzzy counts are reported separately so a threshold change is visible.

type Session

type Session struct {
	ID        string    `json:"id"`
	StartedAt time.Time `json:"started_at"`
	Set       string    `json:"set"`
	Candidate string    `json:"candidate"`
	Repeats   int       `json:"repeats"`
	CaseIDs   []string  `json:"case_ids"`
	Cohort    string    `json:"cohort"`
}

Session records the immutable local plan used for one replay batch.

type SetSummary

type SetSummary struct {
	Name           string
	Cases          int
	GoldCases      int
	TruePositive   int
	FalseNegative  int
	FalsePositive  int
	Unlabeled      int
	QueuedFindings int
	PinCount       int
	Cap            int
	Warning        string
	Composition    map[string]int
}

SetSummary lets users inspect corpus coverage before an eval consumes tokens.

func InspectSets

func InspectSets(store *Store) ([]SetSummary, error)

InspectSets summarizes all logical sets and their diversified mix.

type Store

type Store struct {
	// contains filtered or unexported fields
}

Store is the opt-in sqlite-first local registry. It is intentionally a separate database so merely opening the normal pipeline DB never creates an eval table or runs an eval migration.

func Open

func Open(root string) (*Store, error)

Open creates the local eval registry. The eval CLI, AutoCapture, and RelabelRun open it explicitly; opening the pipeline database never does.

func (*Store) Close

func (s *Store) Close() error

func (*Store) ListCases

func (s *Store) ListCases(set string) ([]Case, error)

ListCases resolves the logical sets. Diversified is gold-only, size-capped, stratified, and pinned: unlabeled cases never fill it. Pins stay until the case is pruned, loses its gold, RefreshDiversified is called, or the live cap shrinks (the read path trims oldest pins to the current cap, at most one per stratum when reconciling to 0 or a lower cap). Tune is leftover labeled cases after those pins, the set matcher thresholds may be fitted on.

func (*Store) Prune

func (s *Store) Prune(ctx context.Context, maxCases int) (int, error)

Prune bounds the corpus at maxCases by dropping the oldest cases first, and reports how many it removed. A maxCases of 0 or less keeps every case.

It never removes a case reserved by a replay session or one that already has recorded candidate replays: those evaluations are the result of tokens somebody spent, and a cohort in an eval report pins the exact case IDs it compared, so reclaiming one would silently invalidate a published comparison. When protected cases alone exceed the cap the corpus stays over it rather than deleting that evidence - the cap is a retention target, not a promise to reach a number.

func (*Store) RefreshDiversified

func (s *Store) RefreshDiversified() ([]Case, error)

RefreshDiversified rebuilds the official pin set from current gold.

func (*Store) SetDiversifiedSize

func (s *Store) SetDiversifiedSize(n int)

SetDiversifiedSize sets the official-set cap. 0 means one gold case per stratum with no Hamilton bound. Negative values restore the default cap.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL