stats

package
v0.25.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 2, 2026 License: MPL-2.0 Imports: 15 Imported by: 0

Documentation

Overview

Package stats computes derived, read-only statistics (pass rates, durations, token and tool-call distributions, and error diagnostics) over MindTrial result sets, filtered and grouped by task/run metadata. It is analytical output derived from canonical results, not a canonical result format itself, and is not persisted back into result artifacts.

Index

Constants

This section is empty.

Variables

View Source
var (
	// ErrInvalidStatus is returned when a --status filter value is not recognized.
	ErrInvalidStatus = errors.New("invalid status filter value")
	// ErrInvalidTagMode is returned when a --tag-mode value is not recognized.
	ErrInvalidTagMode = errors.New("invalid tag-mode value")
)
View Source
var ErrInvalidDimension = errors.New("invalid group-by dimension")

ErrInvalidDimension is returned when a group-by dimension name is not recognized.

View Source
var ErrInvalidOutputFormat = errors.New("invalid stats output format")

ErrInvalidOutputFormat is returned for an unrecognized stats output format.

Functions

func Write

func Write(format OutputFormat, groupBy []Dimension, records []Record, out io.Writer) error

Write renders records in the given format to out. groupBy determines the dimension column order for the text/CSV formats; it is ignored for json/jsonl, which serialize Record.Dimensions as a map instead.

Types

type Dimension

type Dimension string

Dimension identifies a result/task attribute that stats records can be grouped or filtered by.

const (
	// DimensionProvider groups/filters by the AI provider name.
	DimensionProvider Dimension = "provider"
	// DimensionRun groups/filters by the provider's run configuration name.
	DimensionRun Dimension = "run"
	// DimensionModel groups/filters by the resolved model identifier.
	DimensionModel Dimension = "model"
	// DimensionSuite groups/filters by the task's suite label.
	DimensionSuite Dimension = "suite"
	// DimensionCategory groups/filters by the task's category label.
	DimensionCategory Dimension = "category"
	// DimensionDifficulty groups/filters by the task's difficulty label.
	DimensionDifficulty Dimension = "difficulty"
	// DimensionTag groups/filters by task tag. Unlike the other dimensions, a single result
	// can carry multiple tags, so grouping by tag causes that result to be counted in every
	// tag group it belongs to (groups overlap and are not additive).
	DimensionTag Dimension = "tag"
)

func ParseDimensions

func ParseDimensions(commaSeparated string) ([]Dimension, error)

ParseDimensions parses a comma-separated list of dimension names into validated, lower-cased Dimension values. Blank entries are ignored. Each dimension may appear at most once: a repeated dimension has no sensible use case and would otherwise produce ambiguous combinations downstream (see dimensionCombinations/buildRecord).

type Filters

type Filters struct {
	Providers    []string
	Runs         []string
	Models       []string
	Suites       []string
	Categories   []string
	Difficulties []string
	Statuses     []string
	Tags         []string
	TagMode      TagMode
}

Filters selects the subset of results that contribute to computed stats. Each field is matched case-insensitively; multiple values within one field are combined with a logical OR (any match is sufficient), except Tags, whose combination is controlled by TagMode. A nil/empty field imposes no restriction on that attribute.

func (Filters) Validate

func (f Filters) Validate() error

Validate checks that every filter value is well-formed (currently only Statuses needs parsing), without requiring any results to be present. Useful for failing fast on a bad CLI filter before loading input files.

type OutputFormat

type OutputFormat string

OutputFormat selects how Write renders computed stats records.

const (
	// OutputFormatText renders a human-readable, column-aligned table.
	OutputFormatText OutputFormat = "text"
	// OutputFormatCSV renders comma-separated values with a header row.
	OutputFormatCSV OutputFormat = "csv"
	// OutputFormatJSON renders a single indented JSON array of records.
	OutputFormatJSON OutputFormat = "json"
	// OutputFormatJSONL renders one JSON object per line, one line per record.
	OutputFormatJSONL OutputFormat = "jsonl"
)

func ParseOutputFormat

func ParseOutputFormat(value string) (OutputFormat, error)

ParseOutputFormat validates and normalizes a --stats-format value. A blank value defaults to OutputFormatText.

type Record

type Record struct {
	// Dimensions maps each requested group-by dimension name to this record's value for it
	// (e.g. {"provider": "openai", "run": "gpt-4"}).
	Dimensions map[string]string

	// Count is the number of non-skipped results (Passed + Failed + Error).
	Count   int
	Passed  int
	Failed  int
	Error   int
	Skipped int

	// PassRate, Accuracy and ErrorRate are percentages (0-100) rounded to 2 decimals. See
	// formatters.PassRate/AccuracyRate/ErrorRate for the exact formulas; a zero denominator
	// yields 0, not an undefined value.
	PassRate  float64
	Accuracy  float64
	ErrorRate float64

	// Duration metrics are computed over non-skipped results only.
	TotalDuration  *time.Duration
	MedianDuration *time.Duration
	StddevDuration *time.Duration

	// Token/tool-call metrics reflect only the candidate answer and any subsequent error
	// (never judge/validation usage), matching the HTML report's dynamic summary. A metric
	// is nil when no contributing result reported it.
	//
	// Input metrics resolve cache tokens according to InputTokenAccounting (see
	// TokenUsage.EffectiveInputTokens), so they are comparable across providers regardless
	// of whether cache reads/writes are already part of InputTokens.
	TotalInputTokens  *int64
	MedianInputTokens *float64
	StddevInputTokens *float64

	// Output metrics resolve reasoning tokens according to OutputTokenAccounting (see
	// TokenUsage.GeneratedTokens), so they are comparable across providers regardless of
	// whether reasoning is already part of OutputTokens.
	TotalOutputTokens  *int64
	MedianOutputTokens *float64
	StddevOutputTokens *float64

	// Reasoning/cache metrics sum only the counts providers actually reported, so they stay
	// nil when none did rather than implying zero.
	TotalReasoningTokens  *int64
	MedianReasoningTokens *float64
	StddevReasoningTokens *float64

	TotalCacheReadTokens  *int64
	MedianCacheReadTokens *float64
	StddevCacheReadTokens *float64

	TotalCacheWriteTokens  *int64
	MedianCacheWriteTokens *float64
	StddevCacheWriteTokens *float64

	TotalToolCalls  *int64
	MedianToolCalls *float64
	StddevToolCalls *float64

	// EstimatedCandidateCost derives from the static prices the candidate run was configured
	// with and is never a billed amount. It is nil when any contributing result reported
	// usage that could not be priced, so a partial sum is never mistaken for a complete total.
	EstimatedCandidateCost *float64

	// CandidateCostCurrency is the ISO 4217 code EstimatedCandidateCost is expressed in,
	// empty when unknown.
	CandidateCostCurrency string

	// TransientErrors and ResponseParsingErrors count Error-kind results whose
	// ErrorDetails.Transient/ResponseParsing flag is explicitly true. The two are not
	// mutually exclusive: a result can set both.
	TransientErrors       int
	ResponseParsingErrors int
}

Record holds aggregated metrics for one group of results, keyed by the requested group-by dimensions.

func ComputeStats

func ComputeStats(results runners.Results, groupBy []Dimension, filters Filters) ([]Record, error)

ComputeStats filters and groups results according to groupBy and filters, returning one Record per distinct combination of group-by dimension values. Records are returned in a deterministic order, sorted by dimension values in groupBy order.

type TagMode

type TagMode string

TagMode determines how multiple Filters.Tags values are combined.

const (
	// TagModeAll requires a result to carry every filtered tag (logical AND).
	TagModeAll TagMode = "all"
	// TagModeAny requires a result to carry at least one filtered tag (logical OR).
	TagModeAny TagMode = "any"
	// TagModeDefault is the mode used when Filters.TagMode/a --tag-mode value is blank.
	TagModeDefault = TagModeAll
)

func ParseTagMode

func ParseTagMode(value string) (TagMode, error)

ParseTagMode validates and normalizes a --tag-mode value. A blank value defaults to TagModeDefault.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL