Documentation
¶
Overview ¶
Package eval provides dependency-free golden-query evaluation, aggregate search-quality metrics, versioned reports, baseline comparison, and score-domain-safe threshold sweeps.
Index ¶
- Constants
- func CandidateFloors(outcomes []Outcome, scoreDomain string) ([]float32, error)
- func HashSuite(suite Suite) (string, error)
- func SummarizeByLabel(outcomes []Outcome, label string) map[string]Metrics
- func ValidateCase(c GoldenCase) error
- type CaseRunner
- type Comparison
- type FloorEvaluation
- type GoldenCase
- type GoldenKey
- type Judgment
- type Metrics
- type Outcome
- type OutcomeStatus
- type QualityStatus
- type Regression
- type Report
- type ReportIdentity
- type Result
- type Suite
- type Tolerances
Constants ¶
const ReportSchemaVersion = 1
ReportSchemaVersion is the current serialized report contract.
const ScoreDomainLabel = "score_domain"
ScoreDomainLabel is the required case label used to isolate score domains.
Variables ¶
This section is empty.
Functions ¶
func CandidateFloors ¶
CandidateFloors returns exact observed keep/drop boundaries for one score domain. The caller remains responsible for choosing a production floor.
func SummarizeByLabel ¶
SummarizeByLabel groups outcomes by one case label value.
func ValidateCase ¶
func ValidateCase(c GoldenCase) error
ValidateCase validates one evaluation case without modifying it.
Types ¶
type CaseRunner ¶
type CaseRunner interface {
Run(ctx context.Context, c GoldenCase) (results []Result, errCategory string, err error)
}
CaseRunner executes one golden case against a retrieval backend and returns its ordered results. The eval package stays dependency-free: callers wire their own client (which owns query text, mode, and filtering) behind this interface and map their hit type to Result.
A non-nil error reports an execution failure for that single case; the returned errCategory is a stable, sanitized identifier (e.g. "timeout", "search") recorded on the failed outcome. Returning an error does not abort the suite — the case is recorded as failed and the run continues.
type Comparison ¶
type Comparison struct {
Compatible bool `json:"compatible"`
Mismatches []string `json:"mismatches,omitempty"`
Regressions []Regression `json:"regressions,omitempty"`
}
Comparison describes report compatibility and measured regressions.
func Compare ¶
func Compare(baseline Report, current Report, tolerances Tolerances) (Comparison, error)
Compare rejects incompatible reports and checks every common metric scope.
func (Comparison) Regressed ¶
func (c Comparison) Regressed() bool
Regressed reports whether any compared metric exceeded tolerance.
type FloorEvaluation ¶
type FloorEvaluation struct {
Floor float32 `json:"floor"`
RetainedResults int `json:"retained_results"`
Metrics Metrics `json:"metrics"`
}
FloorEvaluation contains aggregate quality after one inclusive score floor.
func SweepResultFloors ¶
func SweepResultFloors(ctx context.Context, outcomes []Outcome, scoreDomain string, floors []float32) ([]FloorEvaluation, error)
SweepResultFloors re-evaluates outcomes after applying each inclusive floor. Callers should pass a bounded candidate set; exhaustive observed boundaries can be expensive for large reports.
type GoldenCase ¶
type GoldenCase struct {
ID string `json:"id"`
Query string `json:"query"`
Language string `json:"language,omitempty"`
ContentKinds []string `json:"content_kinds,omitempty"`
K int `json:"k"`
Expected []GoldenKey `json:"expected,omitempty"`
Judgments []Judgment `json:"judgments,omitempty"`
ExpectEmpty bool `json:"expect_empty,omitempty"`
Labels map[string]string `json:"labels,omitempty"`
}
GoldenCase is a validated query and its quality expectation. Case remains the original minimal compatibility type used by RecallAtK and MRR.
type GoldenKey ¶
type GoldenKey struct {
ContentKind string `json:"content_kind"`
ContentID string `json:"content_id"`
}
GoldenKey is the stable JSON content identity used by golden fixtures and reports. Key remains the original compatibility type.
type Judgment ¶
Judgment assigns a relevance grade to a content item for a query. Grades range from 0 (not relevant) to 3 (highly relevant).
type Metrics ¶
type Metrics struct {
Cases int `json:"cases"`
SuccessfulCases int `json:"successful_cases"`
FailedCases int `json:"failed_cases"`
JudgedCases int `json:"judged_cases"`
EmptyCases int `json:"empty_cases"`
RecallAtK float64 `json:"recall_at_k"`
SuccessAtK float64 `json:"success_at_k"`
MRRAtK float64 `json:"mrr_at_k"`
NDCGAtK float64 `json:"ndcg_at_k"`
ExactEmptyRate float64 `json:"exact_empty_rate"`
MinResults int `json:"min_results"`
MaxResults int `json:"max_results"`
MeanResults float64 `json:"mean_results"`
MedianResults float64 `json:"median_results"`
}
Metrics contains macro-averaged quality and result-count measurements.
type Outcome ¶
type Outcome struct {
Status OutcomeStatus `json:"status"`
QualityStatus QualityStatus `json:"quality_status,omitempty"`
Case GoldenCase `json:"case"`
Results []Result `json:"results,omitempty"`
ResultCount int `json:"result_count"`
ErrorCategory string `json:"error_category,omitempty"`
Judged bool `json:"judged"`
RecallAtK float64 `json:"recall_at_k,omitempty"`
SuccessAtK float64 `json:"success_at_k,omitempty"`
ReciprocalRank float64 `json:"reciprocal_rank,omitempty"`
NDCGAtK float64 `json:"ndcg_at_k,omitempty"`
EmptyExpected bool `json:"empty_expected"`
ExactEmpty bool `json:"exact_empty"`
}
Outcome is the evaluation of one case. Status and QualityStatus keep execution failure distinct from successful empty, hit, and miss outcomes.
func Evaluate ¶
func Evaluate(c GoldenCase, results []Result) (Outcome, error)
Evaluate computes metrics for one successful retrieval while preserving the caller's result order. Duplicate returned keys contribute only at their first raw rank, but ResultCount retains the actual number returned.
func Failed ¶
func Failed(c GoldenCase, category string) Outcome
Failed creates an execution-failure outcome. Category should be a stable, sanitized identifier such as "timeout" or "keyword_search".
type OutcomeStatus ¶
type OutcomeStatus string
OutcomeStatus records retrieval execution success or failure.
const ( OutcomeStatusSuccess OutcomeStatus = "success" OutcomeStatusFailed OutcomeStatus = "failed" )
type QualityStatus ¶
type QualityStatus string
QualityStatus records the relevance result independently of execution status.
const ( QualityStatusUnjudged QualityStatus = "unjudged" QualityStatusHit QualityStatus = "hit" QualityStatusMiss QualityStatus = "miss" QualityStatusExactEmpty QualityStatus = "exact_empty" QualityStatusUnexpectedResults QualityStatus = "unexpected_results" )
type Regression ¶
type Regression struct {
Scope string `json:"scope"`
Metric string `json:"metric"`
Baseline float64 `json:"baseline"`
Current float64 `json:"current"`
Delta float64 `json:"delta"`
}
Regression identifies one metric outside its configured tolerance.
type Report ¶
type Report struct {
SchemaVersion int `json:"schema_version"`
Identity ReportIdentity `json:"identity"`
ContentID string `json:"content_id"`
GroupLabels []string `json:"group_labels,omitempty"`
Metrics Metrics `json:"metrics"`
Breakdowns map[string]map[string]Metrics `json:"breakdowns,omitempty"`
Outcomes []Outcome `json:"outcomes"`
}
Report is a validated deterministic evaluation artifact.
func BuildReport ¶
func BuildReport(identity ReportIdentity, outcomes []Outcome, groupLabels ...string) (Report, error)
BuildReport validates, normalizes, and deterministically orders outcomes.
func RunSuite ¶
func RunSuite(ctx context.Context, s Suite, runner CaseRunner, identity ReportIdentity, groupLabels ...string) (Report, error)
RunSuite executes every case in the suite through the runner, evaluates each against its expectation, and builds one deterministic report. Per-case execution failures are captured as failed outcomes rather than aborting the whole run; a nil runner, an invalid suite, or a report-build error aborts.
type ReportIdentity ¶
type ReportIdentity struct {
DatasetID string `json:"dataset_id"`
SuiteID string `json:"suite_id"`
CandidateID string `json:"candidate_id"`
}
ReportIdentity binds a report to corpus, suite, and candidate configuration.
type Result ¶
Result is one ordered retrieval result and its score. The score domain is supplied by the caller and must not be mixed with other domains in a sweep.
type Suite ¶
type Suite struct {
ID string `json:"id"`
Cases []GoldenCase `json:"cases"`
}
Suite is a portable golden-query fixture. ID is a human-managed version; reports independently hash the validated contents.
type Tolerances ¶
type Tolerances struct {
RecallAtKDrop float64 `json:"recall_at_k_drop"`
SuccessAtKDrop float64 `json:"success_at_k_drop"`
MRRAtKDrop float64 `json:"mrr_at_k_drop"`
NDCGAtKDrop float64 `json:"ndcg_at_k_drop"`
ExactEmptyRateDrop float64 `json:"exact_empty_rate_drop"`
FailedCaseIncrease int `json:"failed_case_increase"`
}
Tolerances defines maximum allowed absolute metric drops and failure growth.