eval

package
v0.58.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 26, 2026 License: MIT Imports: 10 Imported by: 0

Documentation

Overview

Package eval provides dependency-free golden-query evaluation, aggregate search-quality metrics, versioned reports, baseline comparison, and score-domain-safe threshold sweeps.

Index

Constants

View Source
const ReportSchemaVersion = 1

ReportSchemaVersion is the current serialized report contract.

View Source
const ScoreDomainLabel = "score_domain"

ScoreDomainLabel is the required case label used to isolate score domains.

Variables

This section is empty.

Functions

func CandidateFloors

func CandidateFloors(outcomes []Outcome, scoreDomain string) ([]float32, error)

CandidateFloors returns exact observed keep/drop boundaries for one score domain. The caller remains responsible for choosing a production floor.

func HashSuite

func HashSuite(suite Suite) (string, error)

HashSuite returns a stable identity for the exact suite contents.

func SummarizeByLabel

func SummarizeByLabel(outcomes []Outcome, label string) map[string]Metrics

SummarizeByLabel groups outcomes by one case label value.

func ValidateCase

func ValidateCase(c GoldenCase) error

ValidateCase validates one evaluation case without modifying it.

Types

type CaseRunner

type CaseRunner interface {
	Run(ctx context.Context, c GoldenCase) (results []Result, errCategory string, err error)
}

CaseRunner executes one golden case against a retrieval backend and returns its ordered results. The eval package stays dependency-free: callers wire their own client (which owns query text, mode, and filtering) behind this interface and map their hit type to Result.

A non-nil error reports an execution failure for that single case; the returned errCategory is a stable, sanitized identifier (e.g. "timeout", "search") recorded on the failed outcome. Returning an error does not abort the suite — the case is recorded as failed and the run continues.

type Comparison

type Comparison struct {
	Compatible  bool         `json:"compatible"`
	Mismatches  []string     `json:"mismatches,omitempty"`
	Regressions []Regression `json:"regressions,omitempty"`
}

Comparison describes report compatibility and measured regressions.

func Compare

func Compare(baseline Report, current Report, tolerances Tolerances) (Comparison, error)

Compare rejects incompatible reports and checks every common metric scope.

func (Comparison) Regressed

func (c Comparison) Regressed() bool

Regressed reports whether any compared metric exceeded tolerance.

type FloorEvaluation

type FloorEvaluation struct {
	Floor           float32 `json:"floor"`
	RetainedResults int     `json:"retained_results"`
	Metrics         Metrics `json:"metrics"`
}

FloorEvaluation contains aggregate quality after one inclusive score floor.

func SweepResultFloors

func SweepResultFloors(ctx context.Context, outcomes []Outcome, scoreDomain string, floors []float32) ([]FloorEvaluation, error)

SweepResultFloors re-evaluates outcomes after applying each inclusive floor. Callers should pass a bounded candidate set; exhaustive observed boundaries can be expensive for large reports.

type GoldenCase

type GoldenCase struct {
	ID           string            `json:"id"`
	Query        string            `json:"query"`
	Language     string            `json:"language,omitempty"`
	ContentKinds []string          `json:"content_kinds,omitempty"`
	K            int               `json:"k"`
	Expected     []GoldenKey       `json:"expected,omitempty"`
	Judgments    []Judgment        `json:"judgments,omitempty"`
	ExpectEmpty  bool              `json:"expect_empty,omitempty"`
	Labels       map[string]string `json:"labels,omitempty"`
}

GoldenCase is a validated query and its quality expectation. Case remains the original minimal compatibility type used by RecallAtK and MRR.

type GoldenKey

type GoldenKey struct {
	ContentKind string `json:"content_kind"`
	ContentID   string `json:"content_id"`
}

GoldenKey is the stable JSON content identity used by golden fixtures and reports. Key remains the original compatibility type.

type Judgment

type Judgment struct {
	Key       GoldenKey `json:"key"`
	Relevance int       `json:"relevance"`
}

Judgment assigns a relevance grade to a content item for a query. Grades range from 0 (not relevant) to 3 (highly relevant).

type Metrics

type Metrics struct {
	Cases           int `json:"cases"`
	SuccessfulCases int `json:"successful_cases"`
	FailedCases     int `json:"failed_cases"`
	JudgedCases     int `json:"judged_cases"`
	EmptyCases      int `json:"empty_cases"`

	RecallAtK      float64 `json:"recall_at_k"`
	SuccessAtK     float64 `json:"success_at_k"`
	MRRAtK         float64 `json:"mrr_at_k"`
	NDCGAtK        float64 `json:"ndcg_at_k"`
	ExactEmptyRate float64 `json:"exact_empty_rate"`

	MinResults    int     `json:"min_results"`
	MaxResults    int     `json:"max_results"`
	MeanResults   float64 `json:"mean_results"`
	MedianResults float64 `json:"median_results"`
}

Metrics contains macro-averaged quality and result-count measurements.

func Summarize

func Summarize(outcomes []Outcome) Metrics

Summarize computes macro-averaged quality metrics over successful outcomes. Execution failures are counted but excluded from quality denominators.

type Outcome

type Outcome struct {
	Status        OutcomeStatus `json:"status"`
	QualityStatus QualityStatus `json:"quality_status,omitempty"`
	Case          GoldenCase    `json:"case"`
	Results       []Result      `json:"results,omitempty"`
	ResultCount   int           `json:"result_count"`
	ErrorCategory string        `json:"error_category,omitempty"`

	Judged         bool    `json:"judged"`
	RecallAtK      float64 `json:"recall_at_k,omitempty"`
	SuccessAtK     float64 `json:"success_at_k,omitempty"`
	ReciprocalRank float64 `json:"reciprocal_rank,omitempty"`
	NDCGAtK        float64 `json:"ndcg_at_k,omitempty"`

	EmptyExpected bool `json:"empty_expected"`
	ExactEmpty    bool `json:"exact_empty"`
}

Outcome is the evaluation of one case. Status and QualityStatus keep execution failure distinct from successful empty, hit, and miss outcomes.

func Evaluate

func Evaluate(c GoldenCase, results []Result) (Outcome, error)

Evaluate computes metrics for one successful retrieval while preserving the caller's result order. Duplicate returned keys contribute only at their first raw rank, but ResultCount retains the actual number returned.

func Failed

func Failed(c GoldenCase, category string) Outcome

Failed creates an execution-failure outcome. Category should be a stable, sanitized identifier such as "timeout" or "keyword_search".

type OutcomeStatus

type OutcomeStatus string

OutcomeStatus records retrieval execution success or failure.

const (
	OutcomeStatusSuccess OutcomeStatus = "success"
	OutcomeStatusFailed  OutcomeStatus = "failed"
)

type QualityStatus

type QualityStatus string

QualityStatus records the relevance result independently of execution status.

const (
	QualityStatusUnjudged          QualityStatus = "unjudged"
	QualityStatusHit               QualityStatus = "hit"
	QualityStatusMiss              QualityStatus = "miss"
	QualityStatusExactEmpty        QualityStatus = "exact_empty"
	QualityStatusUnexpectedResults QualityStatus = "unexpected_results"
)

type Regression

type Regression struct {
	Scope    string  `json:"scope"`
	Metric   string  `json:"metric"`
	Baseline float64 `json:"baseline"`
	Current  float64 `json:"current"`
	Delta    float64 `json:"delta"`
}

Regression identifies one metric outside its configured tolerance.

type Report

type Report struct {
	SchemaVersion int                           `json:"schema_version"`
	Identity      ReportIdentity                `json:"identity"`
	ContentID     string                        `json:"content_id"`
	GroupLabels   []string                      `json:"group_labels,omitempty"`
	Metrics       Metrics                       `json:"metrics"`
	Breakdowns    map[string]map[string]Metrics `json:"breakdowns,omitempty"`
	Outcomes      []Outcome                     `json:"outcomes"`
}

Report is a validated deterministic evaluation artifact.

func BuildReport

func BuildReport(identity ReportIdentity, outcomes []Outcome, groupLabels ...string) (Report, error)

BuildReport validates, normalizes, and deterministically orders outcomes.

func RunSuite

func RunSuite(ctx context.Context, s Suite, runner CaseRunner, identity ReportIdentity, groupLabels ...string) (Report, error)

RunSuite executes every case in the suite through the runner, evaluates each against its expectation, and builds one deterministic report. Per-case execution failures are captured as failed outcomes rather than aborting the whole run; a nil runner, an invalid suite, or a report-build error aborts.

type ReportIdentity

type ReportIdentity struct {
	DatasetID   string `json:"dataset_id"`
	SuiteID     string `json:"suite_id"`
	CandidateID string `json:"candidate_id"`
}

ReportIdentity binds a report to corpus, suite, and candidate configuration.

type Result

type Result struct {
	Key   GoldenKey `json:"key"`
	Score float32   `json:"score"`
}

Result is one ordered retrieval result and its score. The score domain is supplied by the caller and must not be mixed with other domains in a sweep.

type Suite

type Suite struct {
	ID    string       `json:"id"`
	Cases []GoldenCase `json:"cases"`
}

Suite is a portable golden-query fixture. ID is a human-managed version; reports independently hash the validated contents.

func ParseSuite

func ParseSuite(r io.Reader) (Suite, error)

ParseSuite decodes and validates one golden-query suite.

type Tolerances

type Tolerances struct {
	RecallAtKDrop      float64 `json:"recall_at_k_drop"`
	SuccessAtKDrop     float64 `json:"success_at_k_drop"`
	MRRAtKDrop         float64 `json:"mrr_at_k_drop"`
	NDCGAtKDrop        float64 `json:"ndcg_at_k_drop"`
	ExactEmptyRateDrop float64 `json:"exact_empty_rate_drop"`
	FailedCaseIncrease int     `json:"failed_case_increase"`
}

Tolerances defines maximum allowed absolute metric drops and failure growth.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL