Documentation
¶
Overview ¶
Package runners provides interfaces and implementations for executing MindTrial tasks and collecting their results.
Index ¶
- Variables
- func MergeResults(resultSets ...Results) (Results, MergeStats)
- func NewEmittingLogger(logger zerolog.Logger, emitter eventEmitter) logging.Logger
- type AnswerDetails
- type AsyncResultSet
- type Details
- type EmittingLogger
- type ErrorDetails
- type EvaluationMetadata
- type InputTokenAccounting
- type MergeStats
- type OutputTokenAccounting
- type Pricing
- type ResultKind
- type ResultSet
- type Results
- type RetryPolicy
- type RunConfigSnapshot
- type RunMergeStats
- type RunResult
- type Runner
- type SemanticValidationDetails
- type TaskMetadata
- type TaskRuntimeSettings
- type TokenUsage
- type ToolCallOutput
- type ToolCallSummary
- type ToolUsage
- type ValidationDetails
- type ValidationMethod
Constants ¶
This section is empty.
Variables ¶
var ( // ErrToolNotFound is returned when a required tool is not found in the available tools. ErrToolNotFound = errors.New("required tool not found") // ErrServiceNotFound is returned when a selected task references an unavailable service. ErrServiceNotFound = errors.New("required service not found") // ErrEvaluationSeed is returned when an evaluation randomness seed cannot be created. ErrEvaluationSeed = errors.New("failed to create evaluation seed") // ErrInvalidTaskRuntimeConfig is returned when task service configuration is inconsistent. ErrInvalidTaskRuntimeConfig = errors.New("invalid task runtime configuration") )
Functions ¶
func MergeResults ¶ added in v0.19.0
func MergeResults(resultSets ...Results) (Results, MergeStats)
MergeResults combines multiple result sets into one. When the same (Provider, Run, Task) tuple exists in multiple inputs, the last occurrence wins. New tasks for an existing run are grouped with that run's entries rather than appended to the end. Runs within a provider keep first-seen order. Tasks within each run keep insertion order, with replaced entries keeping their position.
Types ¶
type AnswerDetails ¶ added in v0.6.0
type AnswerDetails struct {
// Title is a descriptive header for the response produced by the target AI model.
Title string
// Explanation of the answer produced by the target AI model.
Explanation []string
// ActualAnswer is the raw answer from the target AI model split into lines.
ActualAnswer []string
// ExpectedAnswer is a set of all acceptable correct answers, each being an array of lines.
ExpectedAnswer [][]string
// Usage contains token usage statistics for generating the answer.
Usage TokenUsage
// ToolUsage contains aggregated execution statistics for any tools invoked while
// producing the answer.
ToolUsage map[string]ToolUsage `json:"ToolUsage,omitempty"`
// ToolCalls contains a log of every individual invocation attempt made while producing
// the answer, including attempts that never actually ran. Tracked separately from
// ToolUsage, which only reflects invocations that actually ran.
ToolCalls []ToolCallSummary `json:"ToolCalls,omitempty"`
}
AnswerDetails defines structured information about the AI model's response to a task.
type AsyncResultSet ¶
type AsyncResultSet interface {
// GetResults returns the task results for each provider.
// The call will block until the run is finished.
GetResults() Results
// ProgressEvents returns a channel that emits run progress as a value between 0 and 1.
// The channel will be closed when the run is finished.
ProgressEvents() <-chan float32
// MessageEvents returns a channel that emits run log messages.
// The channel will be closed when the run is finished.
MessageEvents() <-chan string
// Cancel stops the ongoing run execution.
Cancel()
}
AsyncResultSet extends the basic ResultSet interface to provide asynchronous operation capabilities. It offers channels for monitoring progress and receiving messages during execution, as well as the ability to cancel the ongoing run.
type Details ¶ added in v0.6.0
type Details struct {
// Answer contains details about the AI model's response and reasoning process.
Answer AnswerDetails
// Validation contains details about the answer verification and assessment.
Validation ValidationDetails
// Error contains details about any errors that occurred during task execution.
Error ErrorDetails
}
Details encapsulates comprehensive information about task execution and validation.
type EmittingLogger ¶ added in v0.5.0
type EmittingLogger struct {
// contains filtered or unexported fields
}
EmittingLogger implements the logging.Logger interface and additionally emits log messages as events through the provided event emitter. This allows log messages to be broadcasted to UI components or other consumers.
func (*EmittingLogger) Error ¶ added in v0.5.0
func (l *EmittingLogger) Error(ctx context.Context, level slog.Level, err error, msg string, args ...any)
Error logs an error at the specified level with optional format arguments. The error and message are logged by the logger and emitted as an event.
func (*EmittingLogger) Message ¶ added in v0.5.0
Message logs a message at the specified level with optional format arguments. The message is logged by the logger and emitted as an event.
func (*EmittingLogger) WithContext ¶ added in v0.5.0
func (l *EmittingLogger) WithContext(context string) logging.Logger
WithContext returns a new Logger that appends the specified context to the existing prefix.
type ErrorDetails ¶ added in v0.6.0
type ErrorDetails struct {
// Title provides a summary description of the error.
Title string
// Message contains the primary error message.
Message string
// Details contains any additional error information in a generic structure.
Details map[string][]string
// Usage contains token usage statistics if available even in error scenarios.
// This is typically populated if the error occurs when parsing the generated response.
Usage TokenUsage
// ToolUsage contains aggregated execution statistics for any tools invoked prior to
// the error.
ToolUsage map[string]ToolUsage `json:"ToolUsage,omitempty"`
// ToolCalls contains a log of every individual invocation attempt made prior to the
// error, including attempts that never actually ran. Tracked separately from ToolUsage,
// which only reflects invocations that actually ran.
ToolCalls []ToolCallSummary `json:"ToolCalls,omitempty"`
// ResponseParsing is true when a model response was received but MindTrial could not
// parse or unmarshal it into the expected result structure; nil otherwise. When
// Transient is also nil, this kind of failure should generally be treated as
// non-transient (retrying with the same input is unlikely to help).
ResponseParsing *bool `json:"ResponseParsing,omitempty"`
// Transient indicates whether the error appears temporary/external (true), appears
// permanent/hard (false), or is unknown (nil). This is a best-effort classification,
// not a complete error taxonomy.
Transient *bool `json:"Transient,omitempty"`
// FromValidation is true when Usage/ToolUsage/ToolCalls above belong to the judge's
// validation attempt rather than the candidate response (i.e. validation itself
// failed, after a candidate answer was already produced), so aggregate consumers
// (e.g. the stats package, the HTML report's dynamic summary) must not attribute them
// to the candidate. False (the default) covers every other error, whose usage, if any,
// is the candidate's own.
FromValidation bool `json:"FromValidation,omitempty"`
}
ErrorDetails defines structured information about errors that occurred during execution.
type EvaluationMetadata ¶
type EvaluationMetadata struct {
// Seed is shared by every task attempt of the evaluation; passing it to
// --evaluation-seed reproduces the seed-derived inputs of that evaluation.
Seed string
}
EvaluationMetadata describes the evaluation invocation that produced a result.
type InputTokenAccounting ¶ added in v0.20.2
type InputTokenAccounting string
InputTokenAccounting describes how cached input token counts relate to InputTokens.
const ( // InputTokenAccountingCacheTokensSeparate indicates that cache read and write // tokens are separate from InputTokens and must be added to obtain total input usage. InputTokenAccountingCacheTokensSeparate InputTokenAccounting = "cache_tokens_separate" // InputTokenAccountingCacheTokensIncluded indicates that cache read and write // tokens are informational subsets already included in InputTokens. InputTokenAccountingCacheTokensIncluded InputTokenAccounting = "cache_tokens_included" )
type MergeStats ¶ added in v0.19.0
type MergeStats struct {
// Runs maps provider name → run name → merge statistics for that run.
Runs map[string]map[string]RunMergeStats
}
MergeStats holds statistics collected during a MergeResults operation.
type OutputTokenAccounting ¶ added in v0.22.0
type OutputTokenAccounting string
OutputTokenAccounting describes how reasoning token counts relate to OutputTokens.
const ( // OutputTokenAccountingReasoningTokensIncluded indicates that reasoning is already // represented by OutputTokens, so ReasoningTokens is an informational subset. OutputTokenAccountingReasoningTokensIncluded OutputTokenAccounting = "reasoning_tokens_included" // OutputTokenAccountingReasoningTokensSeparate indicates that reasoning is excluded from // OutputTokens and must be added to obtain total generated usage. OutputTokenAccountingReasoningTokensSeparate OutputTokenAccounting = "reasoning_tokens_separate" )
type Pricing ¶ added in v0.22.0
type Pricing struct {
Currency string
InputPerMillion *float64
OutputPerMillion *float64
CacheReadPerMillion *float64
CacheWritePerMillion *float64
ReasoningPerMillion *float64
}
Pricing mirrors config.Pricing for use in RunConfigSnapshot, preserving the price assumptions that cost estimates for this result were based on. Rates are per million tokens; a nil rate is unknown, while zero is a valid free rate.
type ResultKind ¶
type ResultKind int
ResultKind represents the task execution result status.
const ( Success ResultKind = iota Failure Error NotSupported )
Success indicates that task finished successfully with correct result. Failure indicates that task finished successfully but with incorrect result. Error indicates that task failed to produce a result. NotSupported indicates that task could not finish because the provider does not support the required features.
type ResultSet ¶
type ResultSet interface {
// GetResults returns the task results for each provider.
GetResults() Results
}
ResultSet represents the outcome of executing a set of tasks.
type Results ¶
Results stores task results for each provider.
func (Results) ProviderResultsByRunAndKind ¶
func (r Results) ProviderResultsByRunAndKind(provider string) map[string]map[ResultKind][]RunResult
ProviderResultsByRunAndKind organizes results by run configuration and result kind.
type RetryPolicy ¶ added in v0.22.0
RetryPolicy mirrors config.RetryPolicy for use in RunConfigSnapshot, keeping this package's result types free of a dependency on the config package's own types.
type RunConfigSnapshot ¶ added in v0.22.0
type RunConfigSnapshot struct {
Name string
Model string
MaxRequestsPerMinute int
TextOnly bool
DisableStructuredOutput bool
ModelParameters map[string]interface{}
RetryPolicy RetryPolicy
Pricing *Pricing
}
RunConfigSnapshot is an artifact-safe representation of the effective run configuration, suitable for persisting in results without leaking API keys or other secrets.
type RunMergeStats ¶ added in v0.19.0
type RunMergeStats struct {
// Total is the number of results for this run after merging.
Total int
// Updated is the number of unique tasks whose results were replaced by a later input.
Updated int
}
RunMergeStats holds per-run statistics collected during a merge operation.
type RunResult ¶
type RunResult struct {
// TraceID is a globally unique identifier for this specific task result, used for tracing and correlation.
TraceID string
// Kind indicates the result status.
Kind ResultKind
// Task is the name of the executed task.
Task string
// Provider is the name of the AI provider that executed the task.
Provider string
// Run is the name of the provider's run configuration used.
Run string
// RunConfig contains the effective run configuration used to produce this result,
// with any API keys or other secrets omitted.
RunConfig RunConfigSnapshot
// Evaluation describes the evaluation invocation that produced this result.
Evaluation EvaluationMetadata
// Got is the actual answer received from the AI model.
// For plain text response format, this should be a string that follows the format instruction precisely.
// For structured schema-based response format, this will be any object that conforms to the schema.
Got interface{}
// Want are the accepted valid answer(s) for the task.
// For plain text response format: contains string values that should follow the format instruction precisely.
// For structured schema-based response format: contains object values that conform to the schema.
Want utils.ValueSet
// TaskMetadata carries optional descriptive labels copied from the originating task.
TaskMetadata TaskMetadata
// Details contains comprehensive information about the generated response and validation assessment.
Details Details
// Duration is the cumulative time the AI model itself spent generating a response,
// summed across every conversation turn's model request (network + inference).
// It excludes local tool execution time (see ToolCalls/ToolUsage) and any subsequent
// validation time, so it is not the total wall-clock time spent processing the task.
Duration time.Duration
}
RunResult represents the outcome of executing a single task.
type Runner ¶
type Runner interface {
// Run executes all given tasks against all run configurations and returns when done.
Run(ctx context.Context, tasks []config.Task) (ResultSet, error)
// Start executes all given tasks against all run configurations asynchronously.
// It returns immediately and the execution continues in the background,
// offering progress updates and messages through the returned result set.
Start(ctx context.Context, tasks []config.Task) (AsyncResultSet, error)
// Close releases resources when the runner is no longer needed.
Close(ctx context.Context)
}
Runner executes tasks on configured AI providers.
func NewDefaultRunner ¶
func NewDefaultRunner(ctx context.Context, cfg []config.ProviderConfig, judges []config.JudgeConfig, tools []config.ToolConfig, logger zerolog.Logger) (Runner, error)
NewDefaultRunner creates a new Runner that executes tasks on all configured providers in parallel. The individual runs on a single provider are executed sequentially by default, or in parallel when the provider's MaxParallelRequestsPerMinute is set to a value greater than 0. It returns an error if any provider initialization fails. The runner has no task-scoped services.
func NewDefaultRunnerWithRuntime ¶
func NewDefaultRunnerWithRuntime(ctx context.Context, cfg []config.ProviderConfig, judges []config.JudgeConfig, tools []config.ToolConfig, runtimeSettings TaskRuntimeSettings, logger zerolog.Logger) (Runner, error)
NewDefaultRunnerWithRuntime creates a runner like NewDefaultRunner that additionally provides the task-scoped execution infrastructure described by runtimeSettings.
type SemanticValidationDetails ¶ added in v0.22.0
type SemanticValidationDetails struct {
// Verdict contains the raw verdict produced by the judge.
Verdict interface{}
// JudgeName identifies the judge configuration used.
JudgeName string
// Provider is the name of the AI provider that executed the judge task.
Provider string
// Variant is the name of the judge's configuration variant used.
Variant string
// VariantConfig is the effective variant configuration used by the judge, with any
// API keys or other secrets omitted.
VariantConfig RunConfigSnapshot
}
SemanticValidationDetails identifies the judge variant that evaluated a response and the raw verdict it produced.
type TaskMetadata ¶ added in v0.20.0
type TaskMetadata struct {
// Suite is an optional grouping label for organizing related tasks (e.g. a benchmark suite name).
Suite string
// Category is an optional classification label for the task (e.g. "math", "coding").
Category string
// Difficulty is an optional free-form difficulty label for the task (e.g. "easy", "hard").
Difficulty string
// Tags is an optional set of free-form labels for filtering and grouping tasks.
Tags []string
}
TaskMetadata carries optional descriptive labels from the originating task into the result.
type TaskRuntimeSettings ¶
type TaskRuntimeSettings struct {
// Services lists the task-scoped services available to tools.
Services []config.ServiceConfig
// CustomValidators lists the Docker-backed validators available to tasks.
CustomValidators []config.ValidatorConfig
// EvaluationSeed fixes the seed shared by all task attempts of every evaluation.
// If empty, each evaluation generates its own random seed.
EvaluationSeed string
}
TaskRuntimeSettings configures optional task-scoped execution infrastructure.
type TokenUsage ¶ added in v0.6.1
type TokenUsage struct {
// InputTokens is the input token count reported by the provider.
InputTokens *int64 `` /* 199-byte string literal not displayed */
// OutputTokens is the number of generated output tokens.
OutputTokens *int64 `` /* 200-byte string literal not displayed */
// ReasoningTokens is the number of reasoning (thinking) tokens reported by the provider.
ReasoningTokens *int64 `` /* 265-byte string literal not displayed */
// InputCacheWriteTokens is the number of input tokens written
// into a provider prompt cache.
InputCacheWriteTokens *int64 `` /* 172-byte string literal not displayed */
// InputCacheReadTokens is the number of input tokens read from a
// provider prompt cache.
InputCacheReadTokens *int64 `` /* 167-byte string literal not displayed */
// InputTokenAccounting defines how the cache token counts relate to InputTokens.
InputTokenAccounting InputTokenAccounting `` /* 567-byte string literal not displayed */
// OutputTokenAccounting defines how ReasoningTokens relates to OutputTokens.
OutputTokenAccounting OutputTokenAccounting `` /* 559-byte string literal not displayed */
}
TokenUsage represents token usage consumed by an LLM request. Values are optional and may be nil if not available.
func (TokenUsage) EffectiveInputTokens ¶ added in v0.22.0
func (u TokenUsage) EffectiveInputTokens() *int64
EffectiveInputTokens returns the total input tokens consumed, resolving the cache counters according to InputTokenAccounting. Returns nil when no relevant counter was reported or the accounting mode is not recognized.
func (TokenUsage) GeneratedTokens ¶ added in v0.22.0
func (u TokenUsage) GeneratedTokens() *int64
GeneratedTokens returns the total tokens generated, resolving reasoning tokens according to OutputTokenAccounting. Returns nil when neither counter was reported or the accounting mode is not recognized.
type ToolCallOutput ¶ added in v0.20.0
type ToolCallOutput struct {
// Bytes is the total size of the output stream, regardless of Truncated.
Bytes int64
// Preview is a truncated prefix of the output stream, or nil if not captured or empty.
Preview *string
// Truncated indicates whether Preview was cut short of the full output.
Truncated bool
}
ToolCallOutput holds a size-limited preview of a tool call's output stream.
type ToolCallSummary ¶ added in v0.20.0
type ToolCallSummary struct {
// Tool is the name of the tool this call invoked.
Tool string
// CallID identifies this call, letting a specific invocation be correlated between this
// summary, the corresponding tool-call log lines (which include the same ID in their
// prefix), and - when the calling provider's API assigns its own tool-call ID and that
// ID was reused here - the provider's own API error messages. Falls back to a
// generated ID when the calling provider does not supply one, so this is never empty,
// but its shape/format is not guaranteed to be consistent across providers.
CallID string
// ConversationTurn is the 1-based conversation turn this call was made during, or 0 if
// unknown.
ConversationTurn int
// StartedAt is when this call began (start of setup, before the underlying process runs).
StartedAt time.Time
// CompletedAt is when this call finished, successfully or not.
CompletedAt time.Time
// Duration is the wall-clock duration of the underlying process's runtime, not including
// setup/teardown overhead. It is nil when no process ever ran (e.g. an infrastructure_error).
Duration *time.Duration
// WallTime is the wall-clock duration of the entire call attempt, from setup through
// output retrieval - i.e. Duration plus setup/teardown overhead. Unlike Duration, this is
// always set, even for calls whose underlying process never ran.
WallTime time.Duration
// ExitCode is the underlying process's exit code, or nil if no exit code is known.
ExitCode *int64
// TimedOut indicates the call was aborted due to exceeding its configured timeout.
TimedOut bool
// Status is one of: "success", "nonzero_exit", "empty_output", "timeout",
// "invalid_arguments", "infrastructure_error".
Status string
// Stdout is a size-limited capture of the call's standard output, or nil if no output was
// ever captured.
Stdout *ToolCallOutput
// Stderr is a size-limited capture of the call's standard error, or nil if no output was
// ever captured.
Stderr *ToolCallOutput
// ErrorMessage is a short explanation of the failure when Status is not "success".
ErrorMessage string
}
ToolCallSummary records the outcome of a single tool invocation.
type ToolUsage ¶ added in v0.11.0
type ToolUsage struct {
// CallCount is the number of times the tool's underlying process actually ran.
CallCount *int64 `json:"CallCount,omitempty"`
// TotalDuration is the cumulative execution time for the tool's underlying process.
TotalDuration *time.Duration `json:"TotalDuration,omitempty"`
}
ToolUsage represents aggregated execution statistics captured for a tool during execution. It only reflects invocations whose underlying process actually ran (regardless of exit code); invocation attempts that failed before that point (e.g. invalid arguments, or an infrastructure error during setup) do not affect these aggregates. See ToolCallSummary/ToolCalls for a complete per-invocation log that does include such attempts.
type ValidationDetails ¶ added in v0.6.0
type ValidationDetails struct {
// Title identifies the type of validation assessment performed.
Title string
// Explanation contains detailed analysis of why the validation succeeded or failed.
Explanation []string
// Usage contains token usage statistics for the response validation step.
// This is typically populated when using an LLM judge validator.
Usage TokenUsage
// ToolUsage contains aggregated execution statistics for any tools invoked during
// validation.
ToolUsage map[string]ToolUsage `json:"ToolUsage,omitempty"`
// ToolCalls contains a log of every individual invocation attempt made during
// validation, including attempts that never actually ran. Tracked separately from
// ToolUsage, which only reflects invocations that actually ran.
ToolCalls []ToolCallSummary `json:"ToolCalls,omitempty"`
// Semantic contains the judge's raw verdict and variant provenance, populated
// when validation was performed by an LLM judge; nil otherwise.
Semantic *SemanticValidationDetails `json:"Semantic,omitempty"`
// Method identifies which validation strategy was used.
Method ValidationMethod `json:"Method,omitempty"`
}
ValidationDetails defines structured information about answer verification and correctness assessment.
type ValidationMethod ¶ added in v0.23.0
type ValidationMethod string
ValidationMethod identifies which validation strategy was used for a task.
const ( // ValidationMethodExact indicates exact/canonical value matching. ValidationMethodExact ValidationMethod = "exact" // ValidationMethodSchema indicates JSON Schema validation: the raw candidate // answer is validated directly against the single expected-result JSON Schema // without canonicalization or normalization. The schema's $schema is optional // and defaults to Draft 2020-12; normalization flags (case-sensitive, // ignore-whitespace, trim-lines) are ignored and the mode is mutually // exclusive with judge validation. ValidationMethodSchema ValidationMethod = "schema" // ValidationMethodSemantic indicates LLM judge semantic validation. ValidationMethodSemantic ValidationMethod = "semantic" // ValidationMethodCustom indicates trusted user-defined validation. ValidationMethodCustom ValidationMethod = "custom" )