Documentation
¶
Overview ¶
Package eval implements the local-only review evaluation toolkit.
It deliberately owns a separate registry under <NM_HOME>/eval, so opening the normal pipeline database never creates an eval table or runs an eval migration, and nothing here emits telemetry or reaches the network.
The dependency runs one way: the daemon calls AutoCapture when a run finishes (see RunManager.autoCaptureEvalCase) and RelabelRun when a source PR merges (see RunManager.relabelEvalRun). This package never calls back into the daemon, alters a gate, or influences a pipeline decision.
Index ¶
- Constants
- Variables
- func RenderReport(reports []CandidateReport) string
- func RenderSets(summaries []SetSummary) string
- func Replay(ctx context.Context, store *Store, opts ReplayOptions) (Session, []Evaluation, error)
- type AutoCaptureResult
- type BaselineMetrics
- type Candidate
- type CandidateReport
- type Case
- func Capture(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, ...) ([]Case, error)
- func RelabelAll(ctx context.Context, store *Store, p *paths.Paths, database *db.DB) ([]Case, error)
- func RelabelRun(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, ...) ([]Case, error)
- type Decision
- type Evaluation
- type EvaluationSummary
- func (s EvaluationSummary) ExactRecall() float64
- func (s EvaluationSummary) F1() float64
- func (s EvaluationSummary) F1Conservative() float64
- func (s EvaluationSummary) HasFalsePositiveGold() bool
- func (s EvaluationSummary) Precision() float64
- func (s EvaluationSummary) PrecisionLower() float64
- func (s EvaluationSummary) Recall() float64
- type FindingGold
- type IngestResult
- type Interval
- type Labels
- type Manifest
- type ReplayOptions
- type Score
- type Session
- type SetSummary
- type Store
Constants ¶
const ( GoldTruePositive = "true-positive" GoldFalseNegative = "false-negative" GoldFalsePositive = "false-positive" )
Gold kinds are the scientific labels written from the recorded gate decision (see goldFromRound). Capture writes true-positive gold from a user Fix or from an auto-fix selection on a merged run, false-negative gold from a user-added finding, and false-positive gold from a finding the human did not select for fix that shipped in a merged PR. IngestPostPRMiss writes additional false-negative gold after a green review. Adjudicated labels are never inferred from pending unmatched candidate findings.
Variables ¶
var ErrNoCapturableReview = errors.New("run has no capturable review pass")
ErrNoCapturableReview marks the outcomes where a run simply holds nothing to freeze - no review step, no finished pass, a decision the human has not made yet, or rounds recorded before this machine started keeping replay provenance. These are ordinary states of a healthy pipeline, not failures, so automatic collection can pass over them silently while still surfacing a real capture fault. Every one of them is a deliberate refusal to invent a case: grading a candidate against a half-recorded review pass would be worse than having no case at all.
var ErrReviewDidNotPassGreen = errors.New("review did not pass green")
ErrReviewDidNotPassGreen is returned when ingest is asked to label a run whose review did not complete with a non-blocking (green) pass. A parked or blocking review is a different class: it found something, so it is not a post-PR miss.
Functions ¶
func RenderReport ¶
func RenderReport(reports []CandidateReport) string
RenderReport is a stable human-readable local comparison. Scores are finding-level. Unmatched candidate findings stay pending and are never called false positives. A replay with no gold is unlabeled, not a pass.
func RenderSets ¶
func RenderSets(summaries []SetSummary) string
RenderSets is a stable human-readable preflight. It intentionally contains counts and buckets only, never source paths, URLs, diffs, or findings.
func Replay ¶
func Replay(ctx context.Context, store *Store, opts ReplayOptions) (Session, []Evaluation, error)
Replay runs exactly the captured review pass. It does not start a daemon or use the production NM_HOME: every case is restored into a fresh temp gate and worktree. Push, PR, CI, and all fix loops are intentionally absent from the MVP subject under test.
Types ¶
type AutoCaptureResult ¶
AutoCaptureResult reports what one automatic collection pass did. Skipped is true when the run simply had nothing to freeze, which is the ordinary outcome for most runs and is not a failure.
func AutoCapture ¶
func AutoCapture(ctx context.Context, p *paths.Paths, database *db.DB, runID string, maxCases int) (AutoCaptureResult, error)
AutoCapture freezes one finished run's review passes into the local corpus and then enforces the retention cap.
It is the single entry point for collection that nobody asked for by hand, so it deliberately keeps the same Capture the CLI uses rather than a looser variant: an automatically collected case has to be exactly as trustworthy as one a person captured, or the corpus quietly becomes a different thing from what the eval commands report on.
The caller owns the timeout and the decision to run at all. This function owns nothing of the pipeline: it opens its own registry, does its work, and closes it, so a failure here cannot reach the run that triggered it.
type BaselineMetrics ¶
type BaselineMetrics struct {
InputTokens int64 `json:"input_tokens"`
OutputTokens int64 `json:"output_tokens"`
CacheReadTokens int64 `json:"cache_read_tokens"`
FreshInputTokens int64 `json:"fresh_input_tokens"`
DurationMS int64 `json:"duration_ms"`
TokensReported bool `json:"tokens_reported"`
}
BaselineMetrics is the recorded source-review performance baseline. A false TokensReported means the adapter did not surface trustworthy token usage.
type Candidate ¶
Candidate identifies one agent and model combination under evaluation. The canonical command-line spelling is agent+model, for example codex+gpt-5.4.
func ParseCandidate ¶
ParseCandidate accepts exactly agent+model. Keeping the model explicit makes comparison records self-describing rather than silently inheriting a user's current agent default.
type CandidateReport ¶
type CandidateReport struct {
Cohort string
Summary EvaluationSummary
RepeatCount int
Confidence *Interval
AverageTokens *float64
AverageWallMS float64
OnFrontier bool
}
CandidateReport is one locally observed candidate slice.
func Report ¶
func Report(store *Store) ([]CandidateReport, error)
Report loads every local evaluation result grouped by candidate. It never contacts a forge, agent provider, telemetry endpoint, or remote case store.
type Case ¶
type Case struct {
Manifest
Labels Labels `json:"labels"`
Decision Decision `json:"decision"`
Baseline BaselineMetrics `json:"baseline"`
Dir string `json:"-"`
}
Case is one frozen review pass. Dir is local bookkeeping and never enters a manifest or report payload.
func Capture ¶
func Capture(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, runID string) ([]Case, error)
Capture exports every persisted review pass from one real run. It reads the production state only; it never starts the daemon, changes a gate, fetches, or sends data anywhere.
func RelabelAll ¶
func RelabelRun ¶
func RelabelRun(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, runID string) ([]Case, error)
RelabelRun recomputes gold for every captured case of a run. It adds new auto-fix-merged / shipped-unfixed labels onto previously unlabeled findings and drops obsolete derived merge labels that the current rounds no longer support. It never overwrites adjudicated, user-fix, or ingested post-PR-miss labels. Missing cases are a no-op so a merge observed before the first capture cannot fail the pipeline.
type Decision ¶
type Decision struct {
Action string `json:"action"`
SelectionSource string `json:"selection_source,omitempty"`
SelectedFindingIDs []string `json:"selected_finding_ids,omitempty"`
HasUserFindings bool `json:"has_user_findings"`
}
Decision records the human gate evidence available for the exported review pass. Approval actions were not persisted in historical rows, so Action can be "unknown". The original selections themselves are never guessed, and an unknown action supports no label at all (see hasRecordedDecision).
type Evaluation ¶
type Evaluation struct {
ID string `json:"id"`
SessionID string `json:"session_id"`
CaseID string `json:"case_id"`
Candidate string `json:"candidate"`
Cohort string `json:"cohort"`
Repeat int `json:"repeat"`
StartedAt int64 `json:"started_at"`
CompletedAt int64 `json:"completed_at"`
Status string `json:"status"`
Error string `json:"error,omitempty"`
HasFindingGold bool `json:"has_finding_gold"`
GoldCount int `json:"gold_count"`
TruePositive int `json:"true_positive"`
TruePositiveExact int `json:"true_positive_exact,omitempty"`
TruePositiveFuzzy int `json:"true_positive_fuzzy,omitempty"`
FalseNegative int `json:"false_negative"`
FalsePositive int `json:"false_positive"`
FalsePositiveGold int `json:"false_positive_gold,omitempty"`
Pending int `json:"pending"`
FindingsJSON string `json:"findings_json,omitempty"`
FindingCount int `json:"finding_count"`
InputTokens int64 `json:"input_tokens"`
OutputTokens int64 `json:"output_tokens"`
CacheReadTokens int64 `json:"cache_read_tokens"`
FreshInputTokens int64 `json:"fresh_input_tokens"`
TokensReported bool `json:"tokens_reported"`
DurationMS int64 `json:"duration_ms"`
Model string `json:"model,omitempty"`
}
Evaluation is one candidate replay over one case. Status is "completed" or "failed"; failures remain visible in reports and are not silently scored. Confusion-matrix fields are finding-level: unmatched candidate findings stay in Pending and are never treated as false positives.
type EvaluationSummary ¶
type EvaluationSummary struct {
Candidate string
Total int
Labeled int
TruePositive int
TruePositiveExact int
TruePositiveFuzzy int
FalseNegative int
FalsePositive int
FalsePositiveGold int
Pending int
Failures int
InputTokens int64
OutputTokens int64
FreshInputTokens int64
TokensReported int
DurationMS int64
}
EvaluationSummary aggregates finding-level scores. A case with no gold is unlabeled / pending, never a pass. Unmatched candidate findings stay in Pending and do not become false positives.
func SummarizeEvaluations ¶
func SummarizeEvaluations(evaluations []Evaluation) EvaluationSummary
SummarizeEvaluations scores finding-level gold only. Unmatched candidate findings stay pending. A replay with no gold is unlabeled, not a pass.
func (EvaluationSummary) ExactRecall ¶
func (s EvaluationSummary) ExactRecall() float64
func (EvaluationSummary) F1 ¶
func (s EvaluationSummary) F1() float64
func (EvaluationSummary) F1Conservative ¶
func (s EvaluationSummary) F1Conservative() float64
func (EvaluationSummary) HasFalsePositiveGold ¶
func (s EvaluationSummary) HasFalsePositiveGold() bool
func (EvaluationSummary) Precision ¶
func (s EvaluationSummary) Precision() float64
func (EvaluationSummary) PrecisionLower ¶
func (s EvaluationSummary) PrecisionLower() float64
func (EvaluationSummary) Recall ¶
func (s EvaluationSummary) Recall() float64
type FindingGold ¶
type FindingGold struct {
ID string `json:"id"`
Kind string `json:"kind"`
Source string `json:"source,omitempty"`
File string `json:"file,omitempty"`
Line int `json:"line,omitempty"`
Description string `json:"description,omitempty"`
Severity string `json:"severity,omitempty"`
Action string `json:"action,omitempty"`
}
FindingGold is one finding-level gold label for an underlying issue. Capture writes only the labels goldFromRound can support from recorded evidence; unmatched later candidate findings stay unlabeled. IngestPostPRMiss writes additional false-negative gold for a confirmed miss after a green review.
func ParsePostPRMissFinding ¶
func ParsePostPRMissFinding(raw string) (FindingGold, error)
ParsePostPRMissFinding accepts one finding object. ID and description are required; they are the eval matcher keys. This is the typed source of truth for a confirmed post-PR miss. no-mistakes does not read firstmate ledgers or scrape GitHub review comments.
type IngestResult ¶
IngestResult reports which captured case received post-PR-miss gold.
func IngestPostPRMiss ¶
func IngestPostPRMiss(ctx context.Context, store *Store, p *paths.Paths, database *db.DB, runID string, misses []FindingGold) (IngestResult, error)
IngestPostPRMiss captures a run that already passed review green, then writes confirmed post-PR misses as false-negative gold on the last green review pass. Capture of an existing case is a no-op, so later ingest still attaches gold. Duplicate finding IDs are no-ops.
type Interval ¶
Interval is a finite-sample recall range over cases. Repeats are averaged inside each case so a noisy provider does not inflate apparent sample size.
type Labels ¶
type Labels struct {
Version int `json:"version"`
Findings []FindingGold `json:"findings,omitempty"`
QueuedCandidateFindings int `json:"queued_candidate_findings"`
}
Labels is a local, growing label file. Finding-level gold is the unit of truth. Queued candidate findings are kept as evidence for later adjudication and are never scored as false positives.
func (Labels) FalsePositiveCount ¶
func (Labels) TrueIssueCount ¶
type Manifest ¶
type Manifest struct {
Version int `json:"version"`
ID string `json:"id"`
SourceRunID string `json:"source_run_id"`
SourceRoundID string `json:"source_round_id"`
CapturedAt int64 `json:"captured_at"`
RepoFingerprint string `json:"repo_fingerprint"`
Branch string `json:"branch"`
DefaultBranch string `json:"default_branch"`
BaseSHA string `json:"base_sha"`
HeadSHA string `json:"head_sha"`
StartingHeadSHA string `json:"starting_head_sha"`
ReviewedHeadSHA string `json:"reviewed_head_sha"`
TrustedConfigSHA string `json:"trusted_config_sha"`
Intent string `json:"intent,omitempty"`
IntentSource string `json:"intent_source,omitempty"`
VersionAtCapture string `json:"no_mistakes_version,omitempty"`
BuildSHA string `json:"no_mistakes_build_sha,omitempty"`
ChangedFiles int `json:"changed_files"`
ChangedLines int `json:"changed_lines"`
}
Manifest pins every input needed to recreate a review pass without storing a remote URL. The commits it names live in this repository's local object pool and the configuration it replays under sits beside it in the case directory. It carries no digest of those objects: Git object names are content hashes, so the pins are their own integrity check.
type ReplayOptions ¶
ReplayOptions controls one isolated candidate comparison.
type Score ¶
type Score struct {
TruePositive int
TruePositiveExact int
TruePositiveFuzzy int
FalseNegative int
FalsePositive int
FalsePositiveGold int
Pending int
}
Score is one candidate's finding-level confusion matrix against gold. Pending is unmatched candidate findings: queued, never punished as FP.
func ScoreCandidate ¶
ScoreCandidate matches a candidate finding list against recorded gold.
- TP: the candidate raises the same underlying issue as a true-issue gold (human-accepted Fix, auto-fix that landed in a merged PR, a human-added miss, or a confirmed post-PR miss)
- FN: the candidate misses a true-issue gold
- FP: only an explicit false-positive gold that the candidate still raised
- Pending: unmatched candidate findings, never inferred as invalid
Matching is a documented cascade of strengths: exact-id, exact-text, nearby-line Jaccard, then gated containment. Assignment is one globally optimal assignment over the whole graph (see assignMatches), so neither candidate ordering nor a tier boundary can consume a candidate another gold needed. Headline recall uses the full cascade; exact vs fuzzy counts are reported separately so a threshold change is visible.
type Session ¶
type Session struct {
ID string `json:"id"`
StartedAt time.Time `json:"started_at"`
Set string `json:"set"`
Candidate string `json:"candidate"`
Repeats int `json:"repeats"`
CaseIDs []string `json:"case_ids"`
Cohort string `json:"cohort"`
}
Session records the immutable local plan used for one replay batch.
type SetSummary ¶
type SetSummary struct {
Name string
Cases int
GoldCases int
TruePositive int
FalseNegative int
FalsePositive int
Unlabeled int
QueuedFindings int
PinCount int
Cap int
Warning string
Composition map[string]int
}
SetSummary lets users inspect corpus coverage before an eval consumes tokens.
func InspectSets ¶
func InspectSets(store *Store) ([]SetSummary, error)
InspectSets summarizes all logical sets and their diversified mix.
type Store ¶
type Store struct {
// contains filtered or unexported fields
}
Store is the opt-in sqlite-first local registry. It is intentionally a separate database so merely opening the normal pipeline DB never creates an eval table or runs an eval migration.
func Open ¶
Open creates the local eval registry. The eval CLI, AutoCapture, and RelabelRun open it explicitly; opening the pipeline database never does.
func (*Store) ListCases ¶
ListCases resolves the logical sets. Diversified is gold-only, size-capped, stratified, and pinned: unlabeled cases never fill it. Pins stay until the case is pruned, loses its gold, RefreshDiversified is called, or the live cap shrinks (the read path trims oldest pins to the current cap, at most one per stratum when reconciling to 0 or a lower cap). Tune is leftover labeled cases after those pins, the set matcher thresholds may be fitted on.
func (*Store) Prune ¶
Prune bounds the corpus at maxCases by dropping the oldest cases first, and reports how many it removed. A maxCases of 0 or less keeps every case.
It never removes a case reserved by a replay session or one that already has recorded candidate replays: those evaluations are the result of tokens somebody spent, and a cohort in an eval report pins the exact case IDs it compared, so reclaiming one would silently invalidate a published comparison. When protected cases alone exceed the cap the corpus stays over it rather than deleting that evidence - the cap is a retention target, not a promise to reach a number.
func (*Store) RefreshDiversified ¶
RefreshDiversified rebuilds the official pin set from current gold.
func (*Store) SetDiversifiedSize ¶
SetDiversifiedSize sets the official-set cap. 0 means one gold case per stratum with no Hamilton bound. Negative values restore the default cap.