README
¶
MindTrial
MindTrial lets you test a single AI language model (LLM) or evaluate multiple models side-by-side. It supports providers like OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, and OpenRouter. You can create your own custom tasks with text prompts, plain text or structured JSON response formats, optional file attachments, and tool use for enhanced capabilities; validate responses through exact value matching, an LLM judge for semantic evaluation, or trusted Docker-backed custom validators; and get results in easy-to-read HTML, CSV, and JSON formats.
Quick Start Guide
-
Install the tool:
go install github.com/petmal/mindtrial/cmd/mindtrial@latest -
Run with default settings:
mindtrial run
Prerequisites
- Go 1.26
- Docker (for tool execution; Docker Engine 25.0+ for task services)
- API keys from your chosen AI providers
Key Features
- Compare multiple AI models at once
- Create custom evaluation tasks using simple YAML files
- Attach files or images to prompts for visual tasks
- Enable tool use for tasks with secure sandboxed execution
- Use LLM judges for semantic validation of complex and creative tasks
- Evaluate stateful tool use with task-scoped Docker services and trusted custom validators
- Get results in HTML, CSV, and JSON formats
- Merge and compare results from multiple runs
- Easy to extend with new AI models
- Smart rate limiting to prevent API overload
- Interactive mode with terminal-based UI
Basic Usage
-
Display available commands and options:
mindtrial help -
Run with custom configuration and output options:
mindtrial --config="custom-config.yaml" --tasks="custom-tasks.yaml" --output-dir="./results" --output-basename="custom-tasks-results" run -
Run with specific output formats (CSV only, no HTML):
mindtrial --csv=true --html=false run -
Run in interactive mode to select models and tasks before starting:
mindtrial --interactive run -
Merge results from multiple runs into a single output:
mindtrial --input="results-1.json" --input="results-2.json" --html=true --csv=true --output-basename="merged" merge-results -
Compute derived statistics (pass rate, durations, token usage, and more) grouped by provider and run:
mindtrial --input="results.json" stats
Merging Results
The merge-results command combines results from multiple trial runs into a single output. Input files are specified with the --input flag (can be repeated). Currently, only JSON is supported as the input format. Use the --json=true flag during trial runs to generate JSON output files that can later be merged. The merged output can be generated in any of the supported formats (HTML, CSV, JSON) using the corresponding flags.
[!TIP] You can also use
merge-resultswith a single input file to convert between formats. For example, if you store results in JSON, you can convert them to HTML or CSV at any time:mindtrial --input="results.json" --html=true --csv=true merge-results
[!TIP] If some results failed due to transient errors (e.g., network timeouts), you can re-run only the failed tasks and merge the new results into the original set. Because
merge-resultsuses a last-in-wins strategy for duplicate entries (same provider, run, and task), the corrected results will replace the failed ones.
Computing Statistics
The stats command computes derived statistics (pass rate, accuracy, error rate, duration, token usage, tool calls, error diagnostics, and estimated candidate cost) over one or more result files, filtered and grouped by task/run metadata. Like merge-results, input files are specified with the --input flag (can be repeated); when multiple files are given, they are merged first using the same last-in-wins strategy before stats are computed. This is derived analytical output, not a canonical result format, so it is only written to stdout and is not persisted back into result artifacts. Progress messages are written to stderr, so stdout stays safe to redirect straight into a file or parser for every --stats-format.
mindtrial --input="results.json" --group-by="provider,run,model" --stats-format="csv" stats
- --group-by: Comma-separated grouping dimensions:
provider,run,model,suite,category,difficulty,tag(default:provider,run); each dimension may appear at most once. Grouping bytagis exploded: a task tagged with multiple tags contributes to each of those tag groups, so tag groups overlap and are not additive; duplicate tags on the same task count once. Results missing a value for a grouping dimension are grouped under(unspecified)rather than dropped. - --stats-format: Output format:
text,csv,json, orjsonl(default:text). - --provider, --run, --model, --suite, --category, --difficulty, --status: Restrict stats to matching results; each can be specified multiple times (combined with OR).
--statusacceptspassed,failed,error, orskipped. Pass(unspecified)to match results missing a value for that field. - --tag: Restrict stats to results carrying the given tag; can be specified multiple times. Combined according to
--tag-mode. Pass(unspecified)to match untagged results. - --tag-mode: How multiple
--tagfilters combine:all(default, every tag must be present) orany(at least one tag must be present).
[!NOTE] Token, tool-call, and duration metrics reflect only the candidate answer and any subsequent error (not judge/validation usage), matching the HTML report's dynamic summary.
Median*/Stddev*metrics require at least two contributing samples; otherwise they are omitted.
TotalInputTokensis normalized: providers that count cache reads/writes separately from input (currently Anthropic) have them added, and providers that already include them do not.TotalOutputTokensis normalized the same way: providers that count reasoning separately from output (currently Google and xAI) have it added, and providers that already include it do not.TotalReasoningTokens,TotalCacheReadTokens, andTotalCacheWriteTokenssum only the counts providers actually reported.EstimatedCandidateCost/CandidateCostCurrencyare priced with the candidate run's own prices and never include judge/validation usage.
Example: given a result file with these three tasks:
| Provider | Run | Status | Tags |
|---|---|---|---|
| openai | gpt-4 | Passed | visual, spatial |
| openai | gpt-4 | Failed | visual |
| anthropic | claude | Passed | text |
Grouping by the default provider,run treats each provider/run combination as one group:
$ mindtrial --input="results.json" --stats-format="csv" stats
provider,run,Count,Passed,Failed,...
anthropic,claude,1,1,0,...
openai,gpt-4,2,1,1,...
Grouping by tag instead explodes each task into every tag it carries, so tag groups overlap rather than partition the input (the first task counts toward both visual and spatial):
$ mindtrial --input="results.json" --group-by="tag" --stats-format="csv" stats
tag,Count,Passed,Failed,...
spatial,1,1,0,...
text,1,1,0,...
visual,2,1,1,...
Configuration Guide
MindTrial uses two simple YAML files to control everything:
1. config.yaml - Application Settings
Controls how MindTrial operates, including:
- Where to save results
- Which AI models to use
- API settings and rate limits
2. tasks.yaml - Task Definitions
Defines what you want to evaluate, including:
- Questions/prompts for the AI
- Expected answers
- Response format rules
[!TIP] New to MindTrial? Start with the example files provided and modify them for your needs.
[!TIP] Use interactive mode with the
--interactiveflag to select model configurations and tasks before running, without having to edit configuration files.
config.yaml
This file defines the tool's settings and target model configurations evaluated during the trial run. The main sections include:
- output-dir: Path to the directory where results will be saved.
- task-source: Path to the file with definitions of tasks to run.
- providers: List of providers (i.e. target LLM configurations) to execute tasks during the trial run.
- name: Name of the LLM provider (e.g. openai).
- client-config: Configuration for this provider's client (e.g. API key).
- max-parallel-requests-per-minute: Enables parallel execution of runs within this provider and limits the aggregate number of API requests per minute across all runs. Set to
0or omit for sequential execution (default). - runs: List of runs (i.e. model configurations) for this provider. Unless disabled, all configurations will be trialed.
- name: A unique display-friendly name to be shown in the results.
- model: Model name must be exactly as defined by the backend service's API (e.g. gpt-4o-mini).
- disable-structured-output: Disable structured JSON responses for this run and force plain-text answers.
When enabled, MindTrial:
- treats the model's entire response as the final answer (title and explanation are filled with placeholders)
- forces the model to use plain-text response mode
- skips tasks that require schema-based (
response-result-format) JSON outputs
- text-only: Skip tasks that require native file input to the model API.
When enabled, only tasks without native file input will be executed.
Tasks with files that are only available to local tools (
access: [local]) are still executed. This is useful for text-only models that cannot process images or other files natively.
[!TIP] Use
text-onlyfor models that do not support vision capabilities, such as text-only language models hosted on platforms like OpenRouter.
[!TIP] If the model can output JSON as plain text but cannot follow a provider-enforced schema, prefer
text-response-format. Usedisable-structured-outputonly when the model cannot reliably output JSON at all.
[!IMPORTANT] For models that accept an explicit
response-formatparameter (e.g. OpenRouter), ensure the response format is unset or set to plain text whendisable-structured-outputis enabled; otherwise the run will fail.
[!IMPORTANT] The
disable-structured-outputflag cannot be used in judge configurations, as judges require structured responses for evaluation.
[!NOTE] The OpenAI provider routes GPT-5 and newer model families through the Responses API and currently relies on stored response state (
previous_response_id) for multi-turn and tool-calling flows. As a result, OpenAI Zero Data Retention (ZDR) is not currently supported for those models. Legacy OpenAI models that still use the Chat Completions API are unaffected.
[!IMPORTANT] All provider names must match exactly:
- openai: OpenAI GPT models
- google: Google Gemini models
- anthropic: Anthropic Claude models
- deepseek: DeepSeek open-source models
- mistralai: Mistral AI models
- xai: xAI (Grok) models
- alibaba: Alibaba (Qwen) models
- moonshotai: Moonshot AI (Kimi) models
- openrouter: OpenRouter-hosted models
[!TIP] Instead of a literal value,
client-config.api-keycan reference an environment variable using the{{.Env.NAME}}placeholder (e.g."{{.Env.OPENAI_API_KEY}}"), so secrets don't need to be committed to the config file. Ifapi-keyis omitted entirely (or left blank), each provider falls back to its own default environment variable:
- openai:
OPENAI_API_KEY- google:
GOOGLE_API_KEY- anthropic:
ANTHROPIC_API_KEY- deepseek:
DEEPSEEK_API_KEY- mistralai:
MISTRAL_API_KEY- xai:
XAI_API_KEY- alibaba:
DASHSCOPE_API_KEY- moonshotai:
MOONSHOT_API_KEY- openrouter:
OPENROUTER_API_KEYThis fallback also applies to judge provider configurations under
judges[].provider.client-config.
[!NOTE] Anthropic and DeepSeek providers support configurable request timeout in the
client-configsection:
- request-timeout: Sets the timeout duration for API requests (i.e. thinking).
Alibaba and Moonshot AI providers support endpoint configuration in the
client-configsection:
- endpoint: Specifies the network endpoint URL for the API. If not specified, defaults are:
- Alibaba: Singapore endpoint (
https://dashscope-intl.aliyuncs.com/compatible-mode/v1) for better international access. For China mainland, usehttps://dashscope.aliyuncs.com/compatible-mode/v1.- Moonshot AI: Public API endpoint (
https://api.moonshot.ai/v1).
[!NOTE] Some models support additional model-specific runtime configuration parameters. These can be provided in the
model-parameterssection of the run configuration.Currently supported parameters for OpenAI models include:
- text-response-format: If
true, use plain-text response format (less reliable) for compatibility with models that do not supportJSON.- reasoning-effort: Controls effort on reasoning for reasoning models. (values:
none,minimal,low,medium,high,xhigh,max). Themaxlevel is supported on GPT-5.6 and later reasoning models. Legacy models may not support all values.- reasoning-context: Controls which prior reasoning items are reused across conversation turns (persisted reasoning). (values:
auto,current_turn,all_turns). Supported on GPT-5.6 and later reasoning models.autouses the model's default;current_turnkeeps reasoning from the active turn only, without rendering earlier turns' reasoning into the next call;all_turnsrenders compatible reasoning items from earlier turns into the next call as well, which only has an effect when prior response items are available (e.g. viaprevious_response_id, which MindTrial already relies on for its multi-turn tool-calling loop).- reasoning-mode: Selects an alternate reasoning execution mode. (values:
pro). Pro mode performs additional model work before returning a single final answer, increasing latency and token usage; use selectively for demanding tasks. Supported on GPT-5.6 and later models.- verbosity: Controls how many output tokens are generated. (values:
low,medium,high). May not be supported by legacy models.- temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
- presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
- frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
- max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response.
Currently supported parameters for OpenRouter models include:
- response-format: Controls response format (values:
json-schema,json-object,text, default:json-schema). MindTrial adjusts its parsing logic based on this value, so always use this typed parameter rather than passingresponse_formatdirectly (see note below about extra parameters).- temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
- top-k: Limits token selection to top K candidates (range: 0 or above, default: 0). Value 0 disables this setting.
- min-p: Filters tokens below minimum probability relative to most likely token (range: 0.0 to 1.0, default: 0.0).
- top-a: Considers tokens with sufficiently high probability relative to most likely token (range: 0.0 to 1.0, default: 0.0).
- presence-penalty: Penalizes token repetition based on presence in input (range: -2.0 to 2.0, default: 0.0).
- frequency-penalty: Penalizes token repetition based on frequency in input (range: -2.0 to 2.0, default: 0.0).
- repetition-penalty: Reduces repetition of tokens from input (range: 0.0 to 2.0, default: 1.0).
- max-tokens: Sets upper limit on generated tokens (range: 1 or above, limited by model context).
- max-completion-tokens: Sets the modern upper limit on generated tokens. It cannot be combined with
max-tokens.- reasoning-effort: Controls reasoning depth (values:
none,minimal,low,medium,high,xhigh,max). It maps to OpenRouter's native top-levelreasoning_effortfield.- seed: Enables deterministic sampling when supported.
- parallel-tool-calls: Enables parallel function calling during tool use (default: true).
- verbosity: Controls response verbosity (values:
low,medium,high, default:medium).- server-tools: Injects provider-managed server-side tools into every request for this run. Each entry requires a
type(the tool identifier, e.g.openrouter:fusion) and an optionalparametersmap whose fields are tool-specific. Server tools are appended to the request's tools array alongside any local Docker-based tools. Use an extra parameter (e.g.tool_choice: required) to control invocation behavior.Any additional parameters not listed above can be specified directly in
model-parametersand will be forwarded to the OpenRouter API. Use for provider-specific or OpenRouter-specific parameters. Prefer typed parameters where they exist. Note: if both a typed parameter and an equivalent extra parameter are specified (e.g.,max-tokens: 100andmax_tokens: 500), the extra parameter takes precedence and the API will receive the extra parameter's value.Currently supported parameters for Anthropic models include:
- max-tokens: Controls the maximum number of tokens available to the model for generating a response.
- thinking-budget-tokens: Enables extended thinking with a fixed token budget, giving the model more reasoning capacity on complex tasks. Must be at least 1024 and less than
max-tokens. Ignored wheneffortis also set. Without either setting, thinking follows the model default. Deprecated: Claude Opus 4.7+ removed fixed thinking budgets; setting this returns a 400 error. Useeffortinstead.- effort: Enables adaptive extended thinking and guides how deeply the model reasons before responding, from quick answers (
low) to thorough multi-step reasoning (max) (values:low,medium,high,xhigh,max). Withouteffortorthinking-budget-tokens, thinking follows the model default. When set,thinking-budget-tokensis ignored. Usemax-tokensto cap total output (thinking + response text). Thexhighlevel is recommended for coding and agentic use cases on Claude Opus 4.7+.- thinking: Explicitly overrides the thinking mode (values:
disabled,between_tools).disabledexplicitly disables thinking and cannot be combined withthinking-budget-tokensoreffortabovehigh.between_tools, available on Claude Sonnet 5.5, disables up-front thinking while allowing progress thinking between tool calls; it supportslow,medium, andhigheffort and cannot be combined withthinking-budget-tokens. When omitted, the model's default thinking behaviour is used unlesseffortorthinking-budget-tokensrequests an explicit mode.- temperature: Controls randomness/creativity of responses (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused and deterministic outputs. Deprecated: Claude Opus 4.7+ rejects non-default values with a 400 error.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs. Deprecated: Claude Opus 4.7+ rejects non-default values with a 400 error.
- top-k: Limits tokens considered for each position to top K options. Higher values allow more diverse outputs. Deprecated: Claude Opus 4.7+ rejects any value with a 400 error.
- stream: If
true, enables streaming mode for the API response. Streaming is recommended for requests with largemax-tokensvalues, especially when extended thinking is enabled, to prevent HTTP timeouts on long-running requests. Responses are streamed incrementally and buffered internally before processing.- legacy-structured-output: If
true, uses tool-based structured output instead of native JSON schema output. This is a workaround for models that have difficulty producing valid responses with nativeoutput_config.formatconstrained decoding when extended thinking is enabled. When set, the provider registers asubmit_responsetool and instructs the model to use it to submit its response.- prompt-cache-ttl: Enables Anthropic prompt caching and selects the cache lifetime. Supported values are
5mand1h. When set, MindTrial enables top-level automatic caching and places one explicit cache breakpoint on the most reusable request prefix: the final configured local tool, the final system block, or the final cacheable initial user content block. When omitted, no prompt cache controls are added.Currently supported parameters for Google models include:
- text-response-format: If
true, use plain-text response format (less reliable) for compatibility with models that do not supportJSON. This setting applies to all tasks, including those with and without tools enabled.- text-response-format-with-tools: If
true, forces plain-text response format when tools are enabled (required for pre-Gemini 3 models). Iffalseor unset, uses JSON schema mode with tools (Gemini 3+ default behavior). This setting only applies to tasks with tools enabled.- thinking-level: Controls the maximum depth of the model's internal reasoning process (values:
minimal,low,medium,high). The default is model-dependent — for example, Gemini 3 Pro defaults tohigh, while Gemini 3.5 Flash defaults tomedium.minimalminimizes reasoning for lowest latency (does not guarantee thinking is disabled),lowminimizes latency and cost for simple tasks,mediumbalances reasoning depth and latency, whilehighmaximizes reasoning depth for complex tasks (the model may take longer but output is more carefully reasoned).- media-resolution: Controls the maximum number of tokens allocated per input image (values:
low,medium,high). Higher resolutions improve fine text reading and small detail identification but increase token usage and latency.lowuses 280 tokens;mediumuses 560 tokens;highuses 1120 tokens. If unspecified, the model uses optimal defaults.- temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs. For Gemini 3, it's recommended to keep temperature at default 1.0 for optimal reasoning performance.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
- top-k: Limits tokens considered for each position to top K options. Higher values allow more diverse outputs.
- presence-penalty: Penalizes new tokens based on whether they appear in the text so far. Positive values discourage reuse of tokens, increasing vocabulary. Negative values encourage token reuse.
- frequency-penalty: Penalizes new tokens based on their frequency in the text so far. Positive values discourage frequent tokens proportionally. Negative values encourage token repetition.
- seed: Seed used for deterministic generation. When set, the model attempts to provide consistent responses for identical inputs.
Currently supported parameters for DeepSeek models include:
- temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
- presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
- frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
- thinking: Toggles reasoning (thinking) mode for V4 and newer thinking-capable models. When enabled, the model produces chain-of-thought reasoning before the final answer (values:
enabled,disabled; default:enabledfor V4 models). Note:temperature,top-p,presence-penalty, andfrequency-penaltyare silently ignored in thinking mode.- reasoning-effort: Controls how deeply the model reasons in thinking mode (values:
low,medium,high,xhigh,max). Default ishigh. Currently onlyhighandmaxare distinct effective values —lowandmediumare mapped tohigh, andxhighis mapped tomax.- max-tokens: Controls the maximum number of tokens generated by the model.
Currently supported parameters for Mistral AI models include:
- temperature: Controls randomness/creativity of responses (range: 0.0 to 1.5). Lower values produce more focused and deterministic outputs.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
- max-tokens: Controls the maximum number of tokens available to the model for generating a response.
- presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
- frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
- random-seed: Provides the seed to use for random sampling. If set, requests will generate deterministic results.
- prompt-mode: When set to "reasoning", instructs the model to reason if supported.
- reasoning-effort: Controls reasoning depth for current models (values:
none,minimal,low,medium,high,xhigh). This is independent of the legacyprompt-mode;highcan increase output-token cost.- safe-prompt: Enables content filtering to ensure outputs comply with usage policies.
Currently supported parameters for xAI models include:
- temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
- max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response.
- presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
- frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
- reasoning-effort: Controls effort on reasoning for supported reasoning-capable models (values:
low,medium,high,xhigh). Not all xAI reasoning models (i.e. Grok 4) accept this parameter; Grok 4.5 defaults tohighwhen unset. Thexhighlevel is supported on Grok 4.6 and later.- seed: Integer seed to request deterministic sampling when possible. Determinism is best-effort. xAI makes a best-effort to return repeatable outputs for identical inputs when
seedand other parameters are the same.Currently supported parameters for Alibaba models include:
- response-format: Selects
json-schema,json-object, ortext. MindTrial applies its legacy schema-instruction behavior when this is omitted and structured output is not disabled. Cannot be combined with the deprecatedtext-response-formatordisable-legacy-json-modeproperties.- stream: If
true, enables streaming mode for the API response. Some models (e.g. QwQ, QVQ, and Qwen-Omni) require streaming to be enabled. Responses are streamed incrementally and buffered internally before processing.- enable-thinking: Enables hybrid thinking on supported Qwen models.
- preserve-thinking: Preserves
reasoning_contentacross tool-call turns. Preserved reasoning is included in later input-token counts and billing.- thinking-budget: Optional positive token budget for thinking. This is distinct from
max-tokens/max-completion-tokens, which limit the complete generated response. Mutually exclusive withreasoning-effort.- reasoning-effort: Controls reasoning depth for supported Qwen models (e.g. Qwen 3.8 Max) (values:
none,minimal,low,medium,high,xhigh,max). Mutually exclusive withthinking-budget.- temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
- max-tokens: Controls the maximum number of tokens available to the model for generating a response. Deprecated: use
max-completion-tokensinstead. Mutually exclusive withmax-completion-tokens.- max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response, including reasoning tokens for thinking models. Mutually exclusive with
max-tokens.- presence-penalty: Penalizes new tokens based on whether they appear in the text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage introducing new topics.
- frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
- seed: Makes text generation more deterministic by using the same seed value. When using the same seed and keeping other parameters unchanged, the model makes best-effort to return consistent outputs for identical inputs.
- text-response-format: If
true, use plain-text response format (less reliable) for compatibility with models that do not supportJSON(for example, when thinking is enabled on certain Qwen models). Deprecated: useresponse-format: textinstead. Cannot be combined withresponse-format.- disable-legacy-json-mode: Compatibility toggle that controls legacy prompt injection for JSON formatting. Default:
false(legacy mode on), which adds an explicit JSON formatting instruction to the prompt for improved compatibility with most Qwen models. Setting this totruedisables the legacy prompt injection. For best compatibility and reliable JSON responses, keep this set tofalseunless you are certain the target model works correctly without legacy prompt injection. Deprecated: useresponse-format: json-schemainstead. Cannot be combined withresponse-format.
Currently supported parameters for Moonshot AI models include:
- temperature: Controls randomness/creativity of responses (range: 0.0 to 1.0, default: 0.0). Higher values make output more random, while lower values make it more focused and deterministic. Moonshot AI recommends 0.6 for
kimi-k2models and 1.0 forkimi-k2-thinkingmodels.- top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs. Generally, change either this or temperature, but not both at the same time.
- max-tokens: Controls the maximum number of tokens to generate for the chat completion.
- max-completion-tokens: Controls the modern maximum number of generated tokens. It cannot be combined with deprecated
max-tokens.- reasoning-effort: Controls Kimi K3 reasoning depth (values:
low,high,max).- response-format: Selects
json-schema,json-object, ortext. MindTrial appliesjson-objectwhen this is omitted (unless structured output is disabled). Kimi K3 can use nativejson-schema.- stream: Enables streaming and usage accumulation for long-running responses.
- presence-penalty: Penalizes new tokens based on whether they appear in the text (range: -2.0 to 2.0, default: 0.0). Positive values increase the likelihood of the model discussing new topics.
- frequency-penalty: Penalizes new tokens based on their existing frequency in the text (range: -2.0 to 2.0, default: 0.0). Positive values reduce the likelihood of the model repeating the same phrases verbatim.
- thinking: Toggles the reasoning (thinking) capability for thinking-capable models such as
kimi-k2.6. Accepted values areenabled(default forkimi-k2.6) anddisabled. Older Kimi models that do not support this parameter should omit it.- preserve-thinking: Enables Moonshot's Preserved Thinking feature for
kimi-k2.6, which preserves the model's chain-of-thought across model calls that share the same conversation context (e.g. successive calls in a tool-using task), so the model can build on its earlier reasoning. Accepted value:all; when omitted, prior reasoning is dropped between calls — reducing token cost at the expense of chain-of-thought continuity. Older Kimi models do not support this parameter and should omit it.For
kimi-k2.5andkimi-k2.6, Moonshot AI fixestemperature,top-p,presence-penalty, andfrequency-penaltyto model-specific defaults — supplying any of these parameters will cause the API to reject the request.
[!NOTE] The results will be saved to
<output-dir>/<output-basename>.<format>. If the result output file already exists, it will be replaced. If the log file already exists, it will be appended to.
[!TIP] The following placeholders are available for output paths and names:
- {{.Year}}: Current year
- {{.Month}}: Current month
- {{.Day}}: Current day
- {{.Hour}}: Current hour
- {{.Minute}}: Current minute
- {{.Second}}: Current second
[!TIP] If
log-fileand/oroutput-basenameis blank, the log and/or output will be written to thestdout.
[!NOTE] MindTrial processes tasks across different AI providers simultaneously (in parallel). However, when running multiple configurations from the same provider (e.g. different OpenAI models), these are processed one after another (sequentially) by default.
[!TIP] To run multiple configurations from the same provider in parallel, set
max-parallel-requests-per-minuteon the provider. This enables parallel execution of all runs within that provider, while limiting the aggregate number of API requests per minute across all runs to the specified value. When set to0(or omitted), runs execute sequentially (the default behavior).
[!TIP] Models can use the
max-requests-per-minuteproperty in their run configurations to limit the number of requests made per minute.
[!TIP] To automatically retry failed requests due to rate limiting or other transient errors, set
retry-policyat the provider level to apply to all runs. An individual run configuration can override this by setting its ownretry-policy:
- max-retry-attempts: Maximum number of retry attempts (default: 0 means no retry).
- initial-delay-seconds: Initial delay before the first retry in seconds.
Retries use exponential backoff starting with the initial delay.
[!TIP] To estimate what a trial run costs, set
pricingat the application, provider, or run level. Rates are per million tokens:
- currency: ISO 4217 code the rates are expressed in (default:
USD).- input-per-million: Price per million uncached input tokens.
- output-per-million: Price per million generated output tokens.
- cache-read-per-million: Price per million input tokens read from a prompt cache (default: the input price).
- cache-write-per-million: Price per million input tokens written to a prompt cache (default: the input price).
- reasoning-per-million: Price per million reasoning tokens (default: the output price).
pricingis inherited as a whole, not field-by-field: the nearest level (run, then provider, then application) that sets any field is used in its entirety, replacing rather than merging with whatever a less specific level configured. To override a single rate, restate the full price list at that level:config: pricing: currency: USD input-per-million: 1.25 output-per-million: 10.00 providers: - name: openai runs: - name: "GPT-5.2" model: "gpt-5.2" # Overriding output-per-million requires restating the rest of the list. pricing: currency: USD input-per-million: 1.25 output-per-million: 12.00Estimated costs are reported by the
statscommand and are always estimates derived from these configured rates, never billed amounts. A rate left unset is treated as unknown rather than free, so any estimate depending on it is omitted instead of being understated. The effective prices are also recorded in JSON results, so historical estimates stay reproducible.
[!TIP] To disable all run configurations for a given provider, set
disabled: trueon that provider. An individual run configuration can override this by settingdisabled: false(e.g. to enable just that one configuration).
Example snippet from config.yaml:
# config.yaml
config:
log-file: ""
output-dir: "./results/{{.Year}}-{{.Month}}-{{.Day}}/"
output-basename: "{{.Hour}}-{{.Minute}}-{{.Second}}"
task-source: "./tasks.yaml"
providers:
- name: openai
disabled: true
client-config:
# Resolved from the OPENAI_API_KEY environment variable at load time.
api-key: "{{.Env.OPENAI_API_KEY}}"
retry-policy:
max-retry-attempts: 5
initial-delay-seconds: 30
runs:
- name: "4o-mini - latest"
disabled: false
model: "gpt-4o-mini"
max-requests-per-minute: 3
- name: "o1-mini - latest"
model: "o1-mini"
max-requests-per-minute: 3
model-parameters:
text-response-format: true
- name: "o3-mini - latest (high reasoning)"
model: "o3-mini"
max-requests-per-minute: 3
model-parameters:
reasoning-effort: "high"
- name: openrouter
retry-policy:
max-retry-attempts: 5
initial-delay-seconds: 30
max-parallel-requests-per-minute: 30
client-config:
api-key: "<your-api-key>"
runs:
- name: "OpenAI GPT-5.2 (xhigh reasoning)"
model: "openai/gpt-5.2"
max-requests-per-minute: 20
model-parameters:
verbosity: "medium"
# Pass-through parameters use OpenAI API naming (underscores).
reasoning_effort: "xhigh"
- name: "GPT via Fusion (xhigh panel + judge, high outer)"
# Outer model.
model: "~openai/gpt-latest"
max-requests-per-minute: 2
model-parameters:
# Outer model reasoning effort.
reasoning:
effort: "high"
server-tools:
- type: openrouter:fusion
parameters:
# Inner panel models.
analysis_models:
- "~anthropic/claude-opus-latest"
- "~openai/gpt-latest"
- "~google/gemini-pro-latest"
# Judge model.
model: "~anthropic/claude-opus-latest"
# Inner panel + judge parameters.
reasoning:
effort: "xhigh"
max_completion_tokens: 65536
max_tool_calls: 16
- name: "Google Gemma 3 27B IT (free)"
model: "google/gemma-3-27b-it:free"
max-requests-per-minute: 3
model-parameters:
response-format: "text"
- name: google
client-config:
api-key: "<your-api-key>"
runs:
- name: "Gemini 2.5 Pro - latest"
model: "gemini-2.5-pro"
max-requests-per-minute: 3
model-parameters:
text-response-format-with-tools: true
- name: "Gemini 3 Pro - latest"
model: "gemini-3-pro-preview"
max-requests-per-minute: 3
model-parameters:
thinking-level: "high"
media-resolution: "high"
- name: anthropic
client-config:
api-key: "<your-api-key>"
runs:
- name: "Claude 3.7 Sonnet - latest"
model: "claude-3-7-sonnet-latest"
max-requests-per-minute: 5
model-parameters:
max-tokens: 4096
- name: "Claude 3.7 Sonnet - latest (extended thinking)"
model: "claude-3-7-sonnet-latest"
max-requests-per-minute: 5
model-parameters:
max-tokens: 8192
thinking-budget-tokens: 2048
stream: true
- name: "Claude 4.6 Opus - latest (max adaptive thinking)"
model: "claude-opus-4-6"
max-requests-per-minute: 5
model-parameters:
max-tokens: 65536
effort: max
stream: true
- name: "Claude Opus 4.7 (xhigh adaptive thinking)"
model: "claude-opus-4-7"
max-requests-per-minute: 5
model-parameters:
max-tokens: 65536
effort: xhigh
stream: true
- name: "Claude Opus 5 (xhigh adaptive thinking with prompt caching)"
model: "claude-opus-5"
max-requests-per-minute: 5
model-parameters:
max-tokens: 65536
effort: xhigh
stream: true
prompt-cache-ttl: 5m
- name: deepseek
client-config:
api-key: "<your-api-key>"
request-timeout: 10m
runs:
- name: "DeepSeek-V3.1 - latest (thinking mode)"
model: "deepseek-reasoner"
max-requests-per-minute: 15
- name: mistralai
client-config:
api-key: "<your-api-key>"
runs:
- name: "Mistral Large - latest"
model: "mistral-large-latest"
max-requests-per-minute: 5
retry-policy:
max-retry-attempts: 5
initial-delay-seconds: 30
- name: "Mistral Medium 3.5 - latest (high reasoning)"
model: "mistral-medium-3-5"
max-requests-per-minute: 5
model-parameters:
max-tokens: 65536
reasoning-effort: "high"
- name: alibaba
client-config:
api-key: "<your-api-key>"
endpoint: "https://dashscope-intl.aliyuncs.com/compatible-mode/v1" # Singapore region
retry-policy:
max-retry-attempts: 5
initial-delay-seconds: 30
runs:
- name: "Qwen3-Max-Preview"
model: "qwen3-max-preview"
max-requests-per-minute: 30
- name: "Qwen3-Max-Preview - unstructured"
model: "qwen3-max-preview"
disable-structured-output: true
max-requests-per-minute: 30
- name: "Qwen-VL-Max-Latest"
model: "qwen-vl-max-latest"
max-requests-per-minute: 30
model-parameters:
disable-legacy-json-mode: true
- name: "Qwen3-Next-80B-A3B-Thinking"
model: "qwen3-next-80b-a3b-thinking"
max-requests-per-minute: 30
model-parameters:
text-response-format: true
- name: "QVQ-Max (vision reasoning)"
model: "qvq-max"
max-requests-per-minute: 30
model-parameters:
stream: true # Required for QvQ models
- name: "Qwen3.7 Plus - latest (thinking)"
model: "qwen3.7-plus"
max-requests-per-minute: 30
model-parameters:
enable-thinking: true
preserve-thinking: true
max-tokens: 65536
stream: true
- name: "Qwen3.8 Max - latest (xhigh reasoning)"
model: "qwen3.8-max"
max-requests-per-minute: 30
model-parameters:
reasoning-effort: "xhigh"
preserve-thinking: true
max-completion-tokens: 65536
response-format: "json-object"
stream: true
- name: moonshotai
client-config:
api-key: "<your-api-key>"
runs:
- name: "Kimi K2 - latest (thinking)"
model: "kimi-k2-thinking"
text-only: true # Skip tasks that require native file input
max-requests-per-minute: 3
model-parameters:
temperature: 1.0
max-tokens: 16000
- name: "Kimi K2.6 (thinking)"
model: "kimi-k2.6"
max-requests-per-minute: 3
model-parameters:
max-tokens: 32000
thinking: enabled # default for kimi-k2.6; "disabled" turns off reasoning
preserve-thinking: all # preserve chain-of-thought across model calls (Preserved Thinking)
- name: "Kimi K3 - latest (max reasoning)"
model: "kimi-k3"
max-requests-per-minute: 3
model-parameters:
reasoning-effort: "max"
max-completion-tokens: 65536
response-format: "json-schema"
stream: true
tasks.yaml
This file defines the tasks to be executed on all enabled run configurations. Each task defines the following properties; name, prompt, and response-result-format are always required, and expected-result is required unless the task uses a custom validator:
- name: A unique display-friendly name to be shown in the results.
- prompt: The prompt (i.e. task) that will be sent to the AI model.
- response-result-format: Defines how the AI should format the final answer to the prompt. This can be either:
- Plain text format: A string instruction describing the expected answer format (e.g., "single number", "list of words separated by commas").
- Structured schema format: A JSON schema object defining the structure of the expected response for complex data (e.g., objects with specific fields and types).
- expected-result: Defines the accepted valid answer(s) to the prompt. The format depends on the
response-result-formattype:- For plain text format: A string value or list of string values that follow the format instruction precisely.
- For structured schema format: An object value or list of object values that conform to the JSON schema definition.
Only one expected result needs to match for the response to be considered correct.
With a custom validator,
expected-resultis optional trusted reference data that is passed to the validator unchanged and does not need to matchresponse-result-format.
Optionally, a task can include a list of files to be sent along with the prompt:
-
files: A list of files to attach to the prompt. Each file entry defines the following properties:
- name: A unique name for the file. Every attached file is announced to the model in the prompt as
[file: <name>], regardless of itsaccesssetting, and local tools mount the file under this exact name. - uri: The path or URI to the file. Local file paths and remote HTTP/HTTPS URLs are supported. The content is loaded on demand: it is sent with the request for files with
nativeaccess, and copied to the tool auxiliary directory for files withlocalaccess. - type: The MIME type of the file (e.g.,
image/png,image/jpeg). If omitted, the tool will attempt to infer the type based on the file extension or content. - options: Optional per-file processing options that override the
file-optionsdefaults from thetask-configsection.- image-detail: Controls the fidelity level at which the model processes input images (values:
auto,low,medium,high,original). If the provider does not natively support the requested level, the next higher level or the highest available level is selected. If not set or unknown, the provider uses its own default behavior. Currently only the OpenAI provider honors this setting. It also controls how PDF pages are rendered as images, but only on the OpenAI Responses API; the Chat Completions API does not accept a detail level for file inputs. Theoriginallevel has no PDF equivalent and is treated ashigh. - access: Controls how the file is exposed. Supported values are
native(provider native file/multimodal input) andlocal(local Docker tools). When omitted, the stable default is[native, local]. A per-fileaccessreplaces the inheritedfile-optionsvalue entirely (no union). An explicit empty list[]is invalid. The setting applies to every file type including images, so an image withaccess: [local]is never sent to the model and can only be inspected through a tool. Tasks with files that are only available to local tools (access: [local]) do not require provider file support and are not skipped bytext-only.
- image-detail: Controls the fidelity level at which the model processes input images (values:
- name: A unique name for the file. Every attached file is announced to the model in the prompt as
-
file-options: Default file processing options for all task files in
task-config. Individual files can override viafiles[].options.- image-detail: Default image fidelity (same values as above).
- access: Default access list (same semantics as above). Example:
file-options: { access: [native, local] }.
[!NOTE] If a task requires native file input (any file with
nativeaccess), it will be skipped for provider configurations that do not support file uploads or the specific file type. Tasks with onlylocalaccess never require native file support.
[!IMPORTANT] A file with
localaccess is only readable if the task also enables at least one tool that definesauxiliary-dir. Otherwise the model sees the[file: <name>]reference but has no way to read the contents.
[!NOTE] Currently supported image types include:
image/jpeg,image/jpg,image/png,image/gif,image/webp. Support may vary by provider. Native non-image (document) input is additionally supported by:
- OpenAI: the complete accepted file types list — PDF; Word, Excel, PowerPoint, Pages, Keynote, Google Docs/Sheets/Slides, RTF and OpenDocument text; CSV, TSV and IIF; and a broad set of text and code formats including plain text, Markdown, HTML, XML, CSS, JSON, YAML, TOML, calendar, vCard, subtitles, email and most programming languages
- Google:
application/pdf,application/json,text/plain,text/html,text/css,text/xml,text/csv,text/rtfandtext/javascript, per the supported content types. Only PDF is read with document vision; the other types are extracted as plain text, so charts and formatting are lost.- Anthropic:
application/pdfandtext/plain, per the citations documentation. A plain text document is sent verbatim rather than base64-encoded; other text formats such as CSV or Markdown must declaretype: "text/plain"explicitly to use this path.- OpenRouter, Mistral AI:
application/pdf- Alibaba, Moonshot AI, xAI: images only
- DeepSeek: images only on vision-capable models
A file with
nativeaccess whose type is not supported by the selected provider makes the task unsupported; it is never silently downgraded tolocalaccess.
Example file access configuration:
task-config:
file-options:
access: [native, local] # default for all files
tasks:
- name: "vision task"
prompt: "Describe the image."
response-result-format: "single sentence"
expected-result: "A cat."
files:
- name: "picture"
uri: "./taskdata/cat.png"
type: "image/png"
# inherits access: [native, local]
- name: "local-only task"
prompt: "Use the data file to answer."
response-result-format: "single number"
expected-result: "42"
files:
- name: "data"
uri: "./taskdata/data.csv"
type: "text/csv"
options:
access: [local] # mounted to tool auxiliary-dir only, never sent natively
[!TIP] To disable all tasks by default, set
disabled: truein thetask-configsection. An individual task can override this by settingdisabled: false(e.g. to enable just that one task).
Optionally, a task can also carry descriptive metadata used for filtering and grouping:
- suite: A grouping label for organizing related tasks (e.g. a benchmark suite name).
- category: A classification label for the task (e.g.
"math","coding"). - difficulty: A free-form difficulty label for the task (e.g.
"easy","hard"). - tags: A list of free-form labels for filtering and grouping tasks.
These fields are optional and have no effect on task execution or validation. When present, they are included in the JSON results (TaskMetadata), can be filtered on in the HTML report alongside the existing status/task filters, and are used for the suite/category/difficulty cycling hotkeys (s/c/d) and tag search (/) in the interactive task picker's checklist.
- name: "math problem"
suite: "arithmetic-basics"
category: "math"
difficulty: "easy"
tags: ["smoke", "regression"]
prompt: "What is 2 + 2?"
response-result-format: "single number"
expected-result: "4"
Structured Response Formats
MindTrial supports two types of response formats for tasks:
For tasks where the final answer can be represented as a text value:
- name: "math problem"
prompt: "What is 2 + 2?"
response-result-format: "single number"
expected-result: "4"
For tasks requiring complex structured answers, you can define a JSON schema that describes the expected response format:
- name: "perfect square check"
prompt: "For each number in [4, 9, 10, 16], determine if it's a perfect square and if so, provide the square root."
response-result-format:
type: array
items:
type: object
additionalProperties: false
properties:
number:
type: integer
is_perfect_square:
type: boolean
square_root:
type: integer
required: ["number", "is_perfect_square"]
expected-result:
- - number: 4
is_perfect_square: true
square_root: 2
- number: 9
is_perfect_square: true
square_root: 3
- number: 10
is_perfect_square: false
- number: 16
is_perfect_square: true
square_root: 4
[!IMPORTANT] Structured schema format caveats:
- Semantic validation (LLM judges) cannot be used with structured schema-based response formats.
- All expected results must be objects that conform to the same schema. For array schemas, the entire expected array must be wrapped in a single list item under
expected-resultto avoid treating each array element as a separate expected answer.- Models must support structured JSON response generation for reliable results.
- The OpenAI provider requires JSON schemas to have
additionalProperties: falseand all fields must be required (no optional fields allowed). Other providers may be more flexible.
System Prompt
The system prompt controls how the response format instruction is presented to the AI model.
You can customize this template globally for all tasks in the task-config section, and override it for individual tasks if needed. The template uses Go's template syntax and can reference {{.ResponseResultFormat}} to include the task's response-result-format.
Default system prompt for all tasks is:
Provide the final answer in exactly this format: {{.ResponseResultFormat}}
- system-prompt: A configuration section for the system prompt.
- template: The template string for the system prompt instruction. If not specified, uses the default.
- enable-for: Controls when system prompt should be sent to AI models. Options:
"all": Send system prompt for all tasks (both plain text and structured schema formats)."text": Send system prompt only for tasks with plain text response format (default)."none": Do not send system prompt.
[!NOTE] For structured schema response formats, the JSON schema is automatically passed to the AI model through the provider's structured response mechanism, making explicit format instructions in the system prompt optional.
Validation Rules
These rules control how the validator compares the model's answer to the expected results. By default, comparisons are case-insensitive and only trim leading and trailing whitespace.
You can set validation rules globally for all tasks in the task-config section, and override them for individual tasks if needed; any option not specified at the task level will inherit the global setting from task-config:
- validation-rules: Controls how model responses are validated against expected results.
- case-sensitive: If
true, comparison is case-sensitive. Iffalse(default), comparison ignores case. - ignore-whitespace: If
true, all whitespace (spaces, tabs, newlines) is removed before comparison. Iffalse(default), only leading/trailing whitespace is trimmed, and internal whitespace is preserved. - trim-lines: If
true, trims leading and trailing whitespace from each line before comparison while preserving internal spaces within lines. CRLF line endings are normalized to LF. This option is ignored whenignore-whitespaceis enabled. Iffalse(default), lines are not individually trimmed. - schema-validation: If
true, validates the raw candidate answer against the singleexpected-resultJSON Schema. See Schema-Based Validation for details. Mutually exclusive withjudgeandcustom-validator. - judge: Optional LLM-based semantic validation. See Judge-Based Validation for details. Mutually exclusive with
schema-validationandcustom-validator.- enabled: If
true, uses an LLM judge to evaluate semantic equivalence. Iffalse(default), uses exact value matching. - name: The name of the judge configuration defined in the
config.yamlfile. - variant: The specific run variant from the judge's provider to use.
- enabled: If
- custom-validator: Name of a trusted Docker-backed validator defined in the
config.yamlfile that decides whether the answer is correct. See Custom Validators for details. Mutually exclusive withschema-validationandjudge. Set it to""in a task to opt out of an inherited custom validator.
- case-sensitive: If
Judge-Based Validation
For complex or open-ended tasks where exact value matching is insufficient, you can configure LLM judges to evaluate responses semantically. This is particularly useful for creative writing, reasoning tasks, or when multiple valid answer formats exist.
How it works: Instead of comparing text exactly, an LLM judge evaluates whether the model's response semantically matches the expected result, considering meaning and intent rather than exact wording.
To use judge validation:
-
Define and configure judge models in
config.yaml:config: # ... existing configuration ... judges: - name: "mistral-judge" # A unique name for the judge configuration. provider: name: "mistralai" client-config: api-key: "<your-api-key>" runs: - name: "fast" model: "mistral-medium-latest" max-requests-per-minute: 30 model-parameters: temperature: 0.20 random-seed: 847629 - name: "reasoning" model: "magistral-medium-latest" max-requests-per-minute: 30 model-parameters: prompt-mode: "reasoning" temperature: 0.20 random-seed: 847629 - name: "deepseek-judge" provider: name: "deepseek" client-config: api-key: "<your-api-key>" runs: - name: "fast" model: "deepseek-chat" max-requests-per-minute: 30 model-parameters: temperature: 0.20 - name: "reasoning" model: "deepseek-reasoner" max-requests-per-minute: 30 -
Enable judge validation in
tasks.yaml:# Enable globally for all tasks. task-config: validation-rules: judge: enabled: true name: "mistral-judge" variant: "fast" tasks: # ... tasks will use judge validation by default ... # Override per-task (inherit global settings and override specific options). task-config: validation-rules: judge: enabled: false # Default: use exact value matching. name: "mistral-judge" variant: "fast" tasks: - name: "exact matching task" prompt: "What is 2+2?" response-result-format: "single number" expected-result: "4" # Inherits global validation-rules (exact value matching). - name: "creative writing task" prompt: "Write a short story about..." response-result-format: "short story narrative" expected-result: "A creative and engaging short story" validation-rules: judge: enabled: true # Override: enable judge validation for this task. # Inherits name: "mistral-judge" and variant: "fast" from global config. - name: "complex reasoning task" prompt: "Analyze this philosophical argument..." response-result-format: "structured analysis with reasoning" expected-result: "A thoughtful analysis with logical reasoning" validation-rules: judge: enabled: true variant: "reasoning" # Override: use reasoning run variant instead of fast. # Inherits name: "mistral-judge" from global config.
Schema-Based Validation
For tasks where exact value matching is too strict but an LLM judge is unnecessary, you can validate responses against a JSON Schema. This is particularly useful for numeric ranges, tolerances, bounding boxes, or any structured output where the set of acceptable answers is better expressed as constraints.
How it works: Instead of comparing the candidate answer to literal expected values, the raw candidate answer is validated directly against the single expected-result JSON Schema without canonicalization, normalization, or type coercion. The schema's $schema is optional and defaults to Draft 2020-12; normalization flags (case-sensitive, ignore-whitespace, trim-lines) are ignored and the mode is mutually exclusive with judge.
[!IMPORTANT] Schema validation performs no type coercion. A plain-text
response-result-formatalways produces a string, so itsexpected-resultschema must accept strings (e.g.,type: stringwithpattern). A schema withtype: numberwill not match the string"10"— use a structuredresponse-result-formatsuch astype: numberwhen you need numeric validation.
To use schema validation, enable it in tasks.yaml:
# Simple range — any number between 9.9 and 10.1 is accepted.
task-config:
tasks:
- name: "numeric range"
prompt: "Pick a number between 9.9 and 10.1"
response-result-format:
type: number
validation-rules:
schema-validation: true
expected-result:
$schema: "https://json-schema.org/draft/2020-12/schema"
type: number
minimum: 9.9
maximum: 10.1
Plain-text responses can still use schema validation — the schema just needs to describe a string:
- name: "hex colour"
prompt: "Return a six-digit hexadecimal colour starting with #"
response-result-format: "a six-digit hexadecimal colour starting with #"
validation-rules:
schema-validation: true
expected-result:
$schema: "https://json-schema.org/draft/2020-12/schema"
type: string
pattern: "^#[0-9A-Fa-f]{6}$"
When the model is asked to produce structured JSON, response-result-format and the validator schema are distinct — the former is sent to the model, the latter is evaluator-only:
- name: locate-object
prompt: >
Locate the red object in the image and return its normalized
bounding box.
response-result-format:
type: object
properties:
x:
type: number
minimum: 0
maximum: 1
y:
type: number
minimum: 0
maximum: 1
width:
type: number
minimum: 0
maximum: 1
height:
type: number
minimum: 0
maximum: 1
required: [x, y, width, height]
additionalProperties: false
expected-result:
$schema: "https://json-schema.org/draft/2020-12/schema"
type: object
properties:
x:
type: number
minimum: 0.31
maximum: 0.35
y:
type: number
minimum: 0.42
maximum: 0.46
width:
type: number
minimum: 0.19
maximum: 0.23
height:
type: number
minimum: 0.27
maximum: 0.31
required: [x, y, width, height]
additionalProperties: false
validation-rules:
schema-validation: true
[!TIP] Use
anyOf/oneOf/allOf/enuminside the single schema to express alternatives instead of multiple outerexpected-resultvalues. A literal object containing$schemawithoutschema-validation: trueis still compared via standard value matching — the canonical values are compared exactly (e.g., when the model was asked to generate a schema).
Judge Prompt Customization
MindTrial automatically applies a built-in semantic evaluation template that compares candidate responses against expected answers. For advanced use cases, you can customize the judge prompt template, response format, and acceptance criteria.
Judge prompts can be customized in the validation-rules.judge.prompt section of your tasks.yaml file, either globally in task-config or individually per task.
Customization Fields:
- template: Custom prompt template for the judge (supports template variables listed below).
- verdict-format: Expected response format from the judge (plain text instruction or JSON schema).
- passing-verdicts: Set of verdict values that indicate a passing evaluation, or a single explicit JSON Schema object (identified by a
$schemafield) for threshold-style criteria (e.g. a minimum score).
[!IMPORTANT]
- template is independently optional: you can customize it without also overriding
verdict-format/passing-verdicts, and vice versa.verdict-formatandpassing-verdictsmust either both be specified together, or both left unset to fall back to the built-in defaults; specifying only one leaves the other ambiguous and is rejected.- When
passing-verdictsis a set of literal value(s) (the common case), each value must conform to theverdict-formatstructure.- When
passing-verdictsis an explicit JSON Schema, it is validated as a standalone schema and matched directly against the judge's raw verdict, without any case/whitespace normalization.
[!TIP] The following template variables are available for judge prompts:
- {{.OriginalTask.Prompt}}: The original task prompt
- {{.OriginalTask.ResponseResultFormat}}: Format instruction from the task
- {{.OriginalTask.ExpectedResults}}: Array of expected answers
- {{.Candidate.Response}}: The model's response being evaluated
- {{.Rules.CaseSensitive}}: Boolean case-sensitive validation flag
- {{.Rules.IgnoreWhitespace}}: Boolean ignore whitespace flag
- {{.Rules.TrimLines}}: Boolean trim lines flag
- {{.Verdict.Format}}: The resolved verdict format, rendered as plain text or pretty-printed JSON schema
A sample task from tasks.yaml:
# tasks.yaml
task-config:
disabled: true
file-options:
image-detail: high
system-prompt:
enable-for: "text"
template: |
Provide the final answer in exactly this format: {{.ResponseResultFormat}}
Treat every substring enclosed in `<` and `>` as a variable placeholder.
Substitute only the raw value in place of `<variable name>`, removing the `<` and `>` characters.
Do not add any extra words, punctuation, quotes, or whitespace beyond what the format string shows.
validation-rules:
case-sensitive: false
ignore-whitespace: false
tasks:
- name: "riddle - split words - v1"
disabled: false
prompt: |-
There are four 8-letter words (animals) that have been split into 2-letter pieces.
Find these four words by putting appropriate pieces back together:
RR TE KA DG EH AN SQ EL UI OO HE LO AR PE NG OG
response-result-format: |-
list of words in alphabetical order separated by ", "
system-prompt:
template: "Provide the final answer in exactly this format: {{.ResponseResultFormat}}"
expected-result: |-
ANTELOPE, HEDGEHOG, KANGAROO, SQUIRREL
- name: "visual - shapes - v1"
prompt: |-
The attached picture contains various shapes marked by letters.
It also contains a set of same shapes that have been rotated marked by numbers.
Your task is to find all matching pairs.
response-result-format: |-
<shape number>: <shape letter> pairs separated by ", " and ordered by shape number
expected-result: |-
1: G, 2: F, 3: B, 4: A, 5: C, 6: D, 7: E
validation-rules:
ignore-whitespace: true
files:
- name: "picture"
uri: "./taskdata/visual-shapes-v1.png"
type: "image/png"
options:
image-detail: original
- name: "riddle - anagram - v3"
prompt: |-
Two words (each individual word is a fruit) have been combined and their letters arranged in alphabetical order forming a single group.
Find the original words for each of these 2 groups:
1. AACEEGHPPR
2. ACEILMNOOPRT
response-result-format: |-
1. <word>, <word>
2. <word>, <word>
(words in each group must be alphabetically ordered)
expected-result:
- |
1. GRAPE, PEACH
2. APRICOT, MELON
- |
1. GRAPE, PEACH
2. APRICOT, LEMON
- name: "chemistry - observable phenomena - v1"
disabled: true
prompt: |-
What are the primary observable results of mixing household vinegar (an aqueous solution of acetic acid, $CH_3COOH$) with baking soda (sodium bicarbonate, $NaHCO_3$)?
response-result-format: |-
Provide a bulleted list of the main, directly observable phenomena. Focus on what one would see and hear. Do not include the chemical equation.
expected-result: |-
The response must correctly identify the two main observable results of the chemical reaction.
Crucially, it must mention the production of a gas, described as fizzing, bubbling, or effervescence.
It should also note that the solid baking soda dissolves or disappears as it reacts with the vinegar.
validation-rules:
judge:
enabled: true
name: "mistral-judge"
variant: "reasoning"
# Uses default judge prompt configuration for semantic evaluation.
- name: "code quality - custom judge"
prompt: |-
Write a Python function that finds the maximum value in a list.
response-result-format: |-
complete Python function with proper naming and structure
expected-result: |-
A well-written Python function that correctly finds the maximum value with good practices
validation-rules:
judge:
enabled: true
name: "mistral-judge"
variant: "reasoning"
prompt:
template: |-
Evaluate this Python code for both correctness and quality:
{{.Candidate.Response}}
Criteria: 1) Correctly finds max value, 2) Proper function name/structure, 3) Handles edge cases, 4) Good Python style
verdict-format:
type: object
properties:
quality_score:
type: string
enum: ["excellent", "good", "poor"]
required: ["quality_score"]
additionalProperties: false
passing-verdicts:
- quality_score: "excellent"
- quality_score: "good"
- name: "essay quality - score threshold"
prompt: |-
Write a short essay explaining photosynthesis for a middle school audience.
response-result-format: |-
a clear, well-organized short essay
expected-result: |-
A clear and accurate explanation of photosynthesis appropriate for the target audience
validation-rules:
judge:
enabled: true
name: "mistral-judge"
variant: "reasoning"
prompt:
template: |-
Score this essay from 0 to 100 for clarity, accuracy, and audience appropriateness:
{{.Candidate.Response}}
verdict-format:
type: object
properties:
score:
type: integer
minimum: 0
maximum: 100
required: ["score"]
additionalProperties: false
passing-verdicts:
# An explicit JSON Schema (identified by "$schema") is matched directly against
# the judge's raw verdict, enabling threshold-style criteria.
$schema: "https://json-schema.org/draft/2020-12/schema"
type: object
properties:
score:
type: integer
exclusiveMinimum: 80
required: ["score"]
- name: "structured response - log parsing"
prompt: |-
Parse the following log lines and extract the timestamp, log level, and message for each. If a user ID is present, extract that as well.
Log lines:
[2025-09-14 10:30:00] INFO: User 'admin' logged in successfully.
[2025-09-14 10:31:15] WARN: System memory usage is high.
response-result-format:
type: array
items:
type: object
additionalProperties: false
properties:
timestamp:
type: string
format: "date-time"
level:
type: string
enum: ["INFO", "WARN", "ERROR"]
message:
type: string
user_id:
type: string
required: ["timestamp", "level", "message", "user_id"]
expected-result:
- - timestamp: "2025-09-14T10:30:00Z"
level: "INFO"
message: "User 'admin' logged in successfully."
user_id: "admin"
- timestamp: "2025-09-14T10:31:15Z"
level: "WARN"
message: "System memory usage is high."
user_id: ""
Tools
MindTrial supports tool use for tasks, allowing AI models to execute external tools during task solving. Tools are executed in sandboxed Docker containers with resource limits and network isolation.
Tools must be defined in config.yaml under the tools section. Each tool defines how to execute a specific capability:
- name: A unique name for the tool.
- image: Docker image to use for the tool execution.
- description: A detailed description of what the tool does and how to use it. This description is provided to the LLM to help it understand when and how to use the tool. Be specific and avoid ambiguity to help the LLM choose the correct tool and provide appropriate parameters.
- parameters: JSON schema defining the tool's input parameters. The LLM will generate the actual parameter values based on this schema. Provide comprehensive descriptions that explain parameter purpose and format.
- parameter-files: Mapping of parameter names to container file paths where argument values should be written. Argument values are converted to strings, non-string values are marshaled to JSON. The tool's command should read these files as needed.
- auxiliary-dir: Directory path inside the container where task files with
localaccess will be automatically mounted. If specified, files withlocalaccess attached to the task will be mounted to this directory using each file's unique referencenameexactly as provided. Files in this directory are reset between tool calls. - shared-dir: Directory path inside the container that persists across all tool calls within a single task. If specified, files created in this directory will be available for any subsequent tool calls but will be removed when the task completes.
- command: Command to run inside the container. The standard output of the command execution is captured and passed back to the LLM as is.
- env: Environment variables to set in the container.
- dependencies: Task services the tool can access, each with a
servicename and optionalenvtemplates.
[!IMPORTANT] Tool use requires Docker to be installed and running on the system. Tools are executed in isolated containers with no network access, unless they depend on task services.
Example tool definition in config.yaml:
config:
tools:
- name: python-code-executor
image: python:latest
description: |
Executes Python 3 code in a secure, sandboxed environment to perform calculations, data manipulation, or algorithmic tasks.
IMPORTANT:
- Only the Python standard library is available. No third-party packages (like pandas or numpy) can be imported.
- The environment has no network access.
- Task files made available to this tool are mounted under /app/data/ using their [file: filename] names.
- Use standard file operations like open('/app/data/filename', 'r') to read attached files, where 'filename' matches the name shown in [file: filename] references.
- A persistent shared directory is available at /app/shared/ that persists across ALL tool calls within the same task (regardless of which tool is being called). Files created in this directory will be available in any subsequent tool call.
- Any files or changes outside of /app/shared/ are ephemeral and will be reset between tool calls.
- The code must print its final result to standard output to be returned.
parameters:
type: object
properties:
code:
type: string
description: "A string containing a self-contained Python 3 script. The script must use the `print()` function to return a final result. Example: `print(sum([i for i in range(101) if i % 2 == 0]))`. To read attached files, use open('/app/data/filename', 'r') where 'filename' matches what appears in [file: filename] references."
required:
- code
additionalProperties: false
parameter-files:
code: /app/main.py
auxiliary-dir: /app/data
shared-dir: /app/shared
command:
- python
- /app/main.py
env:
PYTHONIOENCODING: "UTF-8"
PYTHONUNBUFFERED: "1"
PYTHONHASHSEED: "847629"
You can configure tool selection globally for all tasks in the task-config section, and override it for individual tasks if needed. Tools must be defined in config.yaml first.
- tool-selector: Configuration for tool availability during task execution.
- disabled: If
true, no tools are available for tasks (default:false). - tools: List of tools to make available, with per-tool limits.
- name: Name of the tool as defined in
config.yaml. - disabled: If
true, this tool is not available (default:false). - max-calls: Maximum number of times this tool can be called per task (optional).
- timeout: Maximum execution time per tool call (e.g.,
60s, optional). - max-memory-mb: Maximum memory usage in MB per tool call (optional).
- cpu-percent: Maximum CPU usage as percentage per tool call (optional).
- name: Name of the tool as defined in
- service-inputs: Inputs for the task services that the task uses, keyed by service name and then by input name (optional). Inputs set in
task-configapply to every task, and a task can override single inputs. Inputs only take effect for tasks that use the service.
- disabled: If
Example tool configuration in tasks.yaml:
task-config:
tool-selector:
disabled: false
tools:
- name: python-code-executor
disabled: false
max-calls: 10
timeout: 60s
max-memory-mb: 512
cpu-percent: 25
tasks:
- name: "math calculation"
prompt: "Calculate the sum of even numbers from 1 to 100."
response-result-format: "single number"
expected-result: "2550"
# Inherits global tool-selector configuration.
- name: "simple math"
prompt: "What is 2 + 2?"
response-result-format: "single number"
expected-result: "4"
tool-selector:
tools:
- name: python-code-executor
disabled: true # Selectively disable tool for this simple task.
You can set a maximum number of conversation turns per task to act as a safety net against infinite conversation loops (e.g., when a model repeatedly requests exhausted tools). The limit can be configured globally in the task-config section, and overridden for individual tasks if needed. A value of 0 means unlimited.
- max-turns: Maximum number of conversation turns allowed per task (default:
0, unlimited).
Example configuration in tasks.yaml:
task-config:
max-turns: 100 # Default limit for all tasks.
tasks:
- name: "trivia - geography - Asia"
prompt: "What is the capital of Japan?"
response-result-format: "city name"
expected-result: "Tokyo"
max-turns: 200 # Override: allow more turns for this task.
- name: "trivia - geography"
prompt: "What is the capital of Australia?"
response-result-format: "city name"
expected-result: "Canberra"
# Inherits the global limit of 100 turns.
Task Services
A task service is a Docker container that keeps state while the model works on a task, such as a simulated environment, a database, or a web shop. Services are used together with:
- tools, which let the model read and change the service state, and
- custom validators, which can check the final service state after the model answers.
Services and custom validators are independent of each other. Tools that use services work with any validation method (exact match, schema, judge, or custom validator), and a custom validator does not need any services.
[!IMPORTANT] Task services require Docker Engine 25.0 or newer (Engine API 1.44+). Before an evaluation that uses services starts, MindTrial checks the Docker daemon and stops with an error if it is too old.
Services are defined in config.yaml under the services section:
- name: A unique name for the service, used in
dependenciesandservice-inputs. - image: Docker image used to run the service.
- command: Command overriding the image's default command (optional).
- env: Environment variables set for every instance of the service (optional).
- endpoint: Where tools and validators connect to the service.
- port: Container port on which the service accepts connections.
- scheme: URL scheme,
http(default) orhttps.
- input-env: Inputs that tasks can set, mapping each input name to the environment variable that receives its value in the service container (optional). An input must not set a different value for a variable that
envalready defines. - healthcheck: Command that checks whether the service is ready, run inside the service container without a shell (optional). If set, the service is ready once the command succeeds. If not set, the service is ready as soon as its container is running.
- startup-timeout: Maximum time a service with a
healthcheckmay take to become ready (optional, default15s). The check is repeated until it succeeds or this time runs out, and a single check may also take up to this long. - max-memory-mb: Maximum memory available to the service container in MB (optional).
- cpu-percent: Maximum CPU usage as a percentage of total host CPU (optional).
A tool or custom validator gets access to a service by listing it in dependencies:
- dependencies: List of services the tool or validator can access.
- service: Name of the service.
- env: Environment variables set in the tool or validator container (optional). Each value is a template filled in from the running service, for example
CART_URL: "{{ .Endpoint }}".
A task sets service inputs with service-inputs in its tool selector:
tool-selector:
service-inputs:
cart-service: # service name
customer_id: "42" # input name declared in the service's input-env
- Input values must be strings, numbers, or booleans.
- String values are templates, so a task can derive an input from the evaluation seed, e.g.
"{{ hash .Evaluation.Seed .Task.Name }}". MindTrial generates a new evaluation seed for each evaluation, writes it to the log, and records it in every result (Evaluation.Seedin the JSON output); pass the same seed with--evaluation-seedto get the same inputs again. service-inputsset intask-configapply to every task, and a task can override single inputs. Inputs only take effect for tasks that use the service.- When an input is not set, the service uses its own default.
How services run:
- Which services start: The services that the task's enabled tools and its custom validator depend on. Tasks that need no services start none.
- When services start: Before the model receives the prompt. Every service must be ready within its
startup-timeout(15 seconds by default); otherwise the attempt fails and its services are removed. - One set of services per attempt: An attempt is one try of one task by one run configuration. Each attempt gets its own service instances, so runs never share state. A retry is a new attempt with fresh services started from the same inputs. If a service does not derive its initial state from its inputs (for example, from a seed input), a retry may start from a different state.
- Validation: After a successful attempt, a custom validator that depends on a service connects to the same instance that the model's tools used. The services are removed after validation.
- Network isolation: Each service has its own internal Docker network without internet access. A tool or validator joins only the networks of the services in its
dependencies, and one without dependencies has no network access. Services cannot reach each other. - Service failure: A service that stops during an attempt is not restarted, because that would silently reset its state. The next tool call or validation that needs it fails, and the task result is an error.
Custom Validators
A custom validator is a Docker container that you provide to decide whether the model's answer is correct. Use it when an answer can only be checked by running code, for example to run generated code against hidden tests or to check the final state of a simulated environment.
A custom validator does not need task services. Without dependencies, it runs in an isolated container with no network access. With dependencies, it can inspect the same service instances that the model's tools used.
How it works: After a successful model attempt, MindTrial runs the validator container once. It passes the task and the model's answer to the validator through templated command arguments, environment variables, and files. The validator prints its verdict as JSON on standard output.
[!IMPORTANT] Custom validators are trusted: they are part of your evaluation setup, not tools for the model, and they decide task outcomes. Only use validator images you control.
Validators are defined in config.yaml under the validators section:
- name: A unique name for the validator, used by
custom-validatorin task validation rules. - image: Docker image used to run the validator.
- command: Command overriding the image's default command (optional). Each argument is a template and is passed directly to Docker, without a shell.
- env: Environment variables set in the validator container (optional). Each value is a template.
- template-files: Files created for each validation and mounted read-only into the validator container (optional).
- path: Absolute path of the file inside the container. Each path must be unique.
- template: Template that produces the file content.
- dependencies: Task services the validator can access (optional).
- timeout: Maximum duration of one validator run (e.g.,
60s). If not set, there is no timeout. - max-memory-mb: Maximum memory available to the validator container in MB (optional).
- cpu-percent: Maximum CPU usage as a percentage of total host CPU (optional).
[!TIP] Pass the model's answer through
template-files. Usecommandandenvonly for short values you control, such as names, seeds, and flags. The model's answer can be of any size and may contain characters that command arguments and environment variables cannot carry; if the validator cannot start, the task result is an error rather than a failed answer.
A task selects a validator with the custom-validator validation rule:
response-result-formatis still required: it tells the model how to write its final answer.expected-resultis optional. When set, it is passed to the validator unchanged as reference data and does not need to matchresponse-result-format.case-sensitive,ignore-whitespace, andtrim-linesare passed to the validator, which decides whether to use them. MindTrial does not change the answer.
Example of a validator that runs hidden tests without any services:
# config.yaml
config:
validators:
- name: hidden-tests
image: example/code-checker:latest
command: [check, --solution, /input/solution.py]
template-files:
- path: /input/solution.py
template: "{{ .Candidate.Response }}"
timeout: 60s
# tasks.yaml
task-config:
tasks:
- name: fizzbuzz
prompt: "Write a Python function fizzbuzz(n) that returns the FizzBuzz sequence from 1 to n."
response-result-format: "complete Python source code"
validation-rules:
custom-validator: hidden-tests
The validator must exit with code 0 and print exactly one JSON object to standard output, for example:
{"correct": false, "title": "2 of 5 tests failed", "explanation": "fizzbuzz(15) returned '15' instead of 'FizzBuzz'."}
The object must match this schema:
{
"type": "object",
"properties": {
"correct": {"type": "boolean"},
"title": {"type": "string", "pattern": "\\S"},
"explanation": {"type": "string", "pattern": "\\S"}
},
"required": ["correct", "title", "explanation"],
"additionalProperties": false
}
correct: truemarks the answer as correct, andcorrect: falsemarks it as failed.titleandexplanationmust contain non-whitespace text.- Anything else makes the task result an error, not a failed answer: a non-zero exit code, a timeout, a Docker error, missing or unknown fields, or any output after the JSON object.
Results checked by a custom validator record custom as their validation method. Failed answers are shown exactly as the model returned them, without a comparison against expected-result. Validator runs are not counted as model tool calls.
Templates
Service inputs, dependency environment variables, and custom validator settings use Go template syntax. MindTrial checks template syntax before any task runs, and using a field that does not exist is an error.
Each kind of template has its own fields:
| Template | Available fields |
|---|---|
service-inputs string values |
.Evaluation.Seed, .Task.Name, .Provider.Name, .Run.Name |
dependencies[].env values |
.Name, .Host, .Port, .Endpoint |
Validator command, env, and template-files |
.OriginalTask.*, .Candidate.Response, .Rules.*, .Evaluation.Seed, .Task.Name, .Provider.Name, .Run.Name |
Fields:
- {{ .Evaluation.Seed }}: Evaluation seed shared by all providers, runs, tasks, and attempts of one evaluation.
- {{ .Task.Name }}, {{ .Provider.Name }}, {{ .Run.Name }}: Names of the task, the provider, and the run configuration.
- {{ .Name }}: Name of the service.
- {{ .Host }}: Network host name of the service. MindTrial generates host names, so use
.Hostor.Endpointinstead of the service name. - {{ .Port }}: Endpoint port of the service.
- {{ .Endpoint }}: Endpoint URL of the service, e.g.
http://<host>:8080. - {{ .OriginalTask.Prompt }}: The task prompt.
- {{ .OriginalTask.ResponseResultFormat }}: The task's
response-result-format. - {{ .OriginalTask.ExpectedResults }}: List of the task's expected results. It is an empty list (not
null) when the task has noexpected-result. - {{ .Candidate.Response }}: The model's final answer. For structured response formats, this is a structured value.
- {{ .Rules.CaseSensitive }}, {{ .Rules.IgnoreWhitespace }}, {{ .Rules.TrimLines }}: The task's validation flags.
The validator fields use the same names as judge prompt templates.
Helper functions:
- hash: Returns a stable unsigned 64-bit number derived from all of its arguments, e.g.
{{ hash .Evaluation.Seed .Task.Name }}. - json: Encodes its argument as JSON, e.g.
{{ json .OriginalTask.ExpectedResults }}.
[!IMPORTANT] Go prints structured values in its own format, which is not JSON. Use
jsonto pass structured data, e.g.{{ json .Candidate.Response }}.
Example: Stateful Shopping Cart
In this example, the model uses the cart tool to change a shopping cart held by the cart-service service. After the model answers, the cart-state validator checks the same cart:
# config.yaml
config:
services:
- name: cart-service
image: example/cart-service:latest
command: [cart-server]
env:
CART_CURRENCY: CAD
endpoint:
port: 8080
input-env:
customer_id: CART_CUSTOMER_ID
healthcheck: [cart-server, --check]
tools:
- name: cart
image: example/cart-client:latest
command: [cart-client, --request, /input/request.json]
description: >
Modify the current customer's shopping cart by adding or removing items.
parameters:
type: object
properties:
request:
type: object
properties:
action:
type: string
enum: [ADD, REMOVE]
description: Operation to perform on the cart.
item:
type: string
enum: [apple, orange, bread]
description: Item to add or remove.
quantity:
type: integer
minimum: 1
description: Number of items to add or remove.
required: [action, item, quantity]
additionalProperties: false
required: [request]
additionalProperties: false
parameter-files:
request: /input/request.json
dependencies:
- service: cart-service
env:
CART_URL: "{{ .Endpoint }}"
validators:
- name: cart-state
image: example/cart-validator:latest
command: [cart-validator, --expected, /input/expected.json]
template-files:
- path: /input/expected.json
template: "{{ json .OriginalTask.ExpectedResults }}"
timeout: 60s
dependencies:
- service: cart-service
env:
CART_URL: "{{ .Endpoint }}"
CART_CURRENCY is a fixed service setting. customer_id is an input that each task can set; the service receives it as CART_CUSTOMER_ID. The task below derives the customer ID from the evaluation seed, so all providers and runs in one evaluation start with the same customer. Passing the same --evaluation-seed again gives the same customer:
# tasks.yaml
task-config:
tasks:
- name: update-shopping-cart
prompt: >
Use the cart tool to add two apples and one loaf of bread.
response-result-format: "short summary of the final cart contents"
expected-result:
items:
apple: 2
bread: 1
tool-selector:
tools:
- name: cart
service-inputs:
cart-service:
customer_id: "{{ hash .Evaluation.Seed .Task.Name }}"
validation-rules:
custom-validator: cart-state
The expected-result object is reference data for the validator; it does not have to match the plain-text response-result-format. The validator reads it from /input/expected.json as [{"items":{"apple":2,"bread":1}}] and compares it with the cart state it gets from CART_URL.
Command Reference
mindtrial [options] [command]
Commands:
run Start the trials
merge-results Merge results from multiple runs
stats Compute derived statistics from result files
help Show help
version Show version
Options:
--config string Configuration file path (default: config.yaml)
--tasks string Task definitions file path
--output-dir string Results output directory
--output-basename string Base filename for results; replace if exists; blank = stdout
--html Generate HTML output (default: true)
--csv Generate CSV output (default: false)
--json Generate JSON output (default: false)
--input string Input result file path for merge-results/stats; can be specified multiple times
--log string Log file path; append if exists; blank = stdout
--verbose Enable detailed logging
--debug Enable low-level debug logging (implies --verbose)
--interactive Enable interactive interface for run configuration, and real-time progress monitoring (default: false)
--evaluation-seed string Seed for reproducible evaluation behavior (e.g., derived task service inputs); generated for each evaluation when omitted
--group-by string Comma-separated stats grouping dimensions: provider, run, model, suite, category, difficulty, tag (default: provider,run)
--stats-format string Stats output format: text, csv, json, or jsonl (default: text)
--provider string Filter stats to this provider; can be specified multiple times
--run string Filter stats to this run configuration; can be specified multiple times
--model string Filter stats to this model; can be specified multiple times
--suite string Filter stats to this task suite; can be specified multiple times
--category string Filter stats to this task category; can be specified multiple times
--difficulty string Filter stats to this task difficulty; can be specified multiple times
--status string Filter stats to this result status (passed, failed, error, skipped); can be specified multiple times
--tag string Filter stats to results tagged with this value; can be specified multiple times
--tag-mode string How multiple --tag filters combine for stats: all or any (default: all)
Contributing
Contributions are welcome! Please review our CONTRIBUTING.md guidelines for more details.
Getting the Source Code
Clone the repository and install dependencies:
git clone https://github.com/petmal/mindtrial.git
cd mindtrial
go mod download
Running Tests
Execute the unit tests with:
go test -tags=test -race -v ./...
Project Details
/
├── cmd/
│ └── mindtrial/ # Command-line interface and main entry point
│ └── tui/ # Terminal-based UI and interactive mode functionality
├── config/ # Data models and management for configuration and task definitions
├── formatters/ # Output formatting for results
├── pkg/ # Shared packages and utilities
├── providers/ # AI model service provider connectors
│ ├── execution/ # Provider run execution utilities and coordination
│ └── tools/ # Execution engine for external tools used by models
├── runners/ # Task execution and result aggregation
├── stats/ # Derived statistics (filtering, grouping, aggregation) for the stats command
├── taskdata/ # Auxiliary files referenced by tasks in tasks.yaml
├── validators/ # Result validation logic
└── version/ # Application metadata
License
This project is licensed under the Mozilla Public License 2.0 - see the LICENSE file for details.
Directories
¶
| Path | Synopsis |
|---|---|
|
cmd
|
|
|
mindtrial
command
Package main provides the command-line interface and the main entry point for MindTrial.
|
Package main provides the command-line interface and the main entry point for MindTrial. |
|
mindtrial/tui
Package tui provides terminal-based UI for MindTrial CLI.
|
Package tui provides terminal-based UI for MindTrial CLI. |
|
Package config contains the data models representing the structure of configuration and task definition files for the MindTrial application.
|
Package config contains the data models representing the structure of configuration and task definition files for the MindTrial application. |
|
Package formatters provides output formatting functionality for MindTrial results.
|
Package formatters provides output formatting functionality for MindTrial results. |
|
pkg
|
|
|
logging
Package logging provides a structured logging interface compatible with slog levels and common logging utilities for the MindTrial application.
|
Package logging provides a structured logging interface compatible with slog levels and common logging utilities for the MindTrial application. |
|
testutils
Package testutils provides utilities for capturing output, managing test files, logging, and making assertions in tests.
|
Package testutils provides utilities for capturing output, managing test files, logging, and making assertions in tests. |
|
utils
Package utils provides general-purpose utilities for JSON handling and panic recovery.
|
Package utils provides general-purpose utilities for JSON handling and panic recovery. |
|
Package pricing computes estimated costs from recorded token usage and a static price list.
|
Package pricing computes estimated costs from recorded token usage and a static price list. |
|
Package providers implements various AI model service provider connectors supported by MindTrial.
|
Package providers implements various AI model service provider connectors supported by MindTrial. |
|
execution
Package execution provides unified provider execution patterns for the MindTrial application.
|
Package execution provides unified provider execution patterns for the MindTrial application. |
|
tools
Package tools provides implementations for executing tools as part of MindTrial's function calling capabilities.
|
Package tools provides implementations for executing tools as part of MindTrial's function calling capabilities. |
|
Package runners provides interfaces and implementations for executing MindTrial tasks and collecting their results.
|
Package runners provides interfaces and implementations for executing MindTrial tasks and collecting their results. |
|
Package stats computes derived, read-only statistics (pass rates, durations, token and tool-call distributions, and error diagnostics) over MindTrial result sets, filtered and grouped by task/run metadata.
|
Package stats computes derived, read-only statistics (pass rates, durations, token and tool-call distributions, and error diagnostics) over MindTrial result sets, filtered and grouped by task/run metadata. |
|
Package validators provides validation mechanisms for AI model responses.
|
Package validators provides validation mechanisms for AI model responses. |
|
Package version provides information about MindTrial including application name, version, and source code repository.
|
Package version provides information about MindTrial including application name, version, and source code repository. |