Documentation
¶
Overview ¶
Package stats computes derived, read-only statistics (pass rates, durations, token and tool-call distributions, and error diagnostics) over MindTrial result sets, filtered and grouped by task/run metadata. It is analytical output derived from canonical results, not a canonical result format itself, and is not persisted back into result artifacts.
Index ¶
Constants ¶
This section is empty.
Variables ¶
var ( // ErrInvalidStatus is returned when a --status filter value is not recognized. ErrInvalidStatus = errors.New("invalid status filter value") // ErrInvalidTagMode is returned when a --tag-mode value is not recognized. ErrInvalidTagMode = errors.New("invalid tag-mode value") )
var ErrInvalidDimension = errors.New("invalid group-by dimension")
ErrInvalidDimension is returned when a group-by dimension name is not recognized.
var ErrInvalidOutputFormat = errors.New("invalid stats output format")
ErrInvalidOutputFormat is returned for an unrecognized stats output format.
Functions ¶
Types ¶
type Dimension ¶
type Dimension string
Dimension identifies a result/task attribute that stats records can be grouped or filtered by.
const ( // DimensionProvider groups/filters by the AI provider name. DimensionProvider Dimension = "provider" // DimensionRun groups/filters by the provider's run configuration name. DimensionRun Dimension = "run" // DimensionModel groups/filters by the resolved model identifier. DimensionModel Dimension = "model" // DimensionSuite groups/filters by the task's suite label. DimensionSuite Dimension = "suite" // DimensionCategory groups/filters by the task's category label. DimensionCategory Dimension = "category" // DimensionDifficulty groups/filters by the task's difficulty label. DimensionDifficulty Dimension = "difficulty" // DimensionTag groups/filters by task tag. Unlike the other dimensions, a single result // can carry multiple tags, so grouping by tag causes that result to be counted in every // tag group it belongs to (groups overlap and are not additive). DimensionTag Dimension = "tag" )
func ParseDimensions ¶
ParseDimensions parses a comma-separated list of dimension names into validated, lower-cased Dimension values. Blank entries are ignored. Each dimension may appear at most once: a repeated dimension has no sensible use case and would otherwise produce ambiguous combinations downstream (see dimensionCombinations/buildRecord).
type Filters ¶
type Filters struct {
Providers []string
Runs []string
Models []string
Suites []string
Categories []string
Difficulties []string
Statuses []string
Tags []string
TagMode TagMode
}
Filters selects the subset of results that contribute to computed stats. Each field is matched case-insensitively; multiple values within one field are combined with a logical OR (any match is sufficient), except Tags, whose combination is controlled by TagMode. A nil/empty field imposes no restriction on that attribute.
type OutputFormat ¶
type OutputFormat string
OutputFormat selects how Write renders computed stats records.
const ( // OutputFormatText renders a human-readable, column-aligned table. OutputFormatText OutputFormat = "text" // OutputFormatCSV renders comma-separated values with a header row. OutputFormatCSV OutputFormat = "csv" // OutputFormatJSON renders a single indented JSON array of records. OutputFormatJSON OutputFormat = "json" // OutputFormatJSONL renders one JSON object per line, one line per record. OutputFormatJSONL OutputFormat = "jsonl" )
func ParseOutputFormat ¶
func ParseOutputFormat(value string) (OutputFormat, error)
ParseOutputFormat validates and normalizes a --stats-format value. A blank value defaults to OutputFormatText.
type Record ¶
type Record struct {
// Dimensions maps each requested group-by dimension name to this record's value for it
// (e.g. {"provider": "openai", "run": "gpt-4"}).
Dimensions map[string]string
// Count is the number of non-skipped results (Passed + Failed + Error).
Count int
Passed int
Failed int
Error int
Skipped int
// PassRate, Accuracy and ErrorRate are percentages (0-100) rounded to 2 decimals. See
// formatters.PassRate/AccuracyRate/ErrorRate for the exact formulas; a zero denominator
// yields 0, not an undefined value.
PassRate float64
Accuracy float64
ErrorRate float64
// Duration metrics are computed over non-skipped results only.
TotalDuration *time.Duration
MedianDuration *time.Duration
StddevDuration *time.Duration
// Token/tool-call metrics reflect only the candidate answer and any subsequent error
// (never judge/validation usage), matching the HTML report's dynamic summary. A metric
// is nil when no contributing result reported it.
//
// Input metrics resolve cache tokens according to InputTokenAccounting (see
// TokenUsage.EffectiveInputTokens), so they are comparable across providers regardless
// of whether cache reads/writes are already part of InputTokens.
TotalInputTokens *int64
MedianInputTokens *float64
StddevInputTokens *float64
// Output metrics resolve reasoning tokens according to OutputTokenAccounting (see
// TokenUsage.GeneratedTokens), so they are comparable across providers regardless of
// whether reasoning is already part of OutputTokens.
TotalOutputTokens *int64
MedianOutputTokens *float64
StddevOutputTokens *float64
// Reasoning/cache metrics sum only the counts providers actually reported, so they stay
// nil when none did rather than implying zero.
TotalReasoningTokens *int64
MedianReasoningTokens *float64
StddevReasoningTokens *float64
TotalCacheReadTokens *int64
MedianCacheReadTokens *float64
StddevCacheReadTokens *float64
TotalCacheWriteTokens *int64
MedianCacheWriteTokens *float64
StddevCacheWriteTokens *float64
TotalToolCalls *int64
MedianToolCalls *float64
StddevToolCalls *float64
// EstimatedCandidateCost derives from the static prices the candidate run was configured
// with and is never a billed amount. It is nil when any contributing result reported
// usage that could not be priced, so a partial sum is never mistaken for a complete total.
EstimatedCandidateCost *float64
// CandidateCostCurrency is the ISO 4217 code EstimatedCandidateCost is expressed in,
// empty when unknown.
CandidateCostCurrency string
// TransientErrors and ResponseParsingErrors count Error-kind results whose
// ErrorDetails.Transient/ResponseParsing flag is explicitly true. The two are not
// mutually exclusive: a result can set both.
TransientErrors int
ResponseParsingErrors int
}
Record holds aggregated metrics for one group of results, keyed by the requested group-by dimensions.
func ComputeStats ¶
ComputeStats filters and groups results according to groupBy and filters, returning one Record per distinct combination of group-by dimension values. Records are returned in a deterministic order, sorted by dimension values in groupBy order.
type TagMode ¶
type TagMode string
TagMode determines how multiple Filters.Tags values are combined.
const ( // TagModeAll requires a result to carry every filtered tag (logical AND). TagModeAll TagMode = "all" // TagModeAny requires a result to carry at least one filtered tag (logical OR). TagModeAny TagMode = "any" // TagModeDefault is the mode used when Filters.TagMode/a --tag-mode value is blank. TagModeDefault = TagModeAll )
func ParseTagMode ¶
ParseTagMode validates and normalizes a --tag-mode value. A blank value defaults to TagModeDefault.