Documentation
¶
Overview ¶
Package tools provides implementations for executing tools as part of MindTrial's function calling capabilities.
Index ¶
- Variables
- func IsTaskRuntimeError(err error) bool
- type DockerTool
- type DockerToolExecutor
- func (d *DockerToolExecutor) Close() error
- func (d *DockerToolExecutor) ExecuteTool(ctx context.Context, logger logging.Logger, toolName string, ...) (json.RawMessage, error)
- func (d *DockerToolExecutor) GetCallSummaries() []ToolCallSummary
- func (d *DockerToolExecutor) GetUsageStats() map[string]ToolUsage
- func (d *DockerToolExecutor) IsToolExhausted(toolName string) bool
- func (d *DockerToolExecutor) RegisterTool(tool *DockerTool)
- func (d *DockerToolExecutor) ValidateImage(ctx context.Context, image string) error
- func (d *DockerToolExecutor) ValidateTaskServiceSupport(ctx context.Context) error
- func (d *DockerToolExecutor) ValidateTool(ctx context.Context, cfg config.ToolConfig) error
- type EvaluationTemplateData
- type ExecutionEnvironment
- type ExecutionTemplateData
- type NameTemplateData
- type OutputCapture
- type TaskRuntime
- type TaskRuntimeConfig
- type TextOrData
- type ToolCallContext
- type ToolCallSummary
- type ToolUsage
Constants ¶
This section is empty.
Variables ¶
var ( // ErrTaskRuntimeConfig is returned when task-scoped service configuration is invalid. ErrTaskRuntimeConfig = errors.New("invalid task runtime configuration") // ErrTaskRuntimeUnsupported is returned when the Docker daemon cannot run task-scoped services. ErrTaskRuntimeUnsupported = errors.New("task services are not supported by the Docker daemon") // ErrTaskRuntimeExecution is returned when a service does not become ready or stops unexpectedly. ErrTaskRuntimeExecution = errors.New("task runtime execution failed") // ErrTaskRuntimeInternal is returned for Docker and runtime lifecycle failures. ErrTaskRuntimeInternal = errors.New("task runtime internal error") )
var ( // ErrToolNotAvailable is returned when a requested tool is not available. ErrToolNotAvailable = errors.New("tool not available") // ErrToolExecutionFailed is returned when a tool executes but fails with an error. ErrToolExecutionFailed = errors.New("tool execution failed") // ErrInvalidToolArguments is returned when tool arguments are invalid or don't match the expected schema. ErrInvalidToolArguments = errors.New("invalid tool arguments") // ErrToolInternal is returned for low-level internal errors during tool execution. ErrToolInternal = errors.New("tool internal error") // ErrUnsupportedToolType is returned when an unsupported tool type is encountered. ErrUnsupportedToolType = errors.New("unsupported tool type") // ErrToolMaxCallsExceeded is returned when a tool has exceeded its maximum call limit. ErrToolMaxCallsExceeded = errors.New("tool max calls exceeded") // ErrToolTimeout is returned when a tool execution times out. ErrToolTimeout = errors.New("tool execution timeout") )
Functions ¶
func IsTaskRuntimeError ¶
IsTaskRuntimeError reports whether err is a task runtime failure, which a model cannot resolve by changing its tool calls.
Types ¶
type DockerTool ¶
type DockerTool struct {
// contains filtered or unexported fields
}
func NewDockerTool ¶
func NewDockerTool(cfg *config.ToolConfig, maxCalls *int, timeout *time.Duration, maxMemoryMB *int, cpuPercent *int) *DockerTool
NewDockerTool creates a new Docker tool.
func (*DockerTool) SetReadOnlyFiles ¶
func (t *DockerTool) SetReadOnlyFiles(files map[string][]byte)
SetReadOnlyFiles mounts each file content, keyed by absolute container path, read-only into every container the tool runs.
type DockerToolExecutor ¶
type DockerToolExecutor struct {
// contains filtered or unexported fields
}
DockerToolExecutor executes tools within Docker containers.
func NewDockerToolExecutor ¶
func NewDockerToolExecutor(ctx context.Context) (*DockerToolExecutor, error)
NewDockerToolExecutor creates a standalone Docker tool executor.
func NewToolExecutor ¶
func NewToolExecutor(ctx context.Context, environment ExecutionEnvironment) (*DockerToolExecutor, error)
NewToolExecutor creates a Docker tool executor for environment. A nil environment creates a standalone ephemeral executor.
func (*DockerToolExecutor) Close ¶
func (d *DockerToolExecutor) Close() error
Close cleans up executor-owned resources.
func (*DockerToolExecutor) ExecuteTool ¶
func (d *DockerToolExecutor) ExecuteTool(ctx context.Context, logger logging.Logger, toolName string, args json.RawMessage, data map[string][]byte, callCtx *ToolCallContext) (json.RawMessage, error)
ExecuteTool executes a tool by name with the given arguments and auxiliary data files. callCtx carries optional caller-supplied metadata (see ToolCallContext) to attach to the resulting ToolCallSummary; pass nil if there is none to provide.
func (*DockerToolExecutor) GetCallSummaries ¶ added in v0.20.0
func (d *DockerToolExecutor) GetCallSummaries() []ToolCallSummary
GetCallSummaries returns a log of every recorded invocation attempt across all tools, in the order calls completed, including attempts that never actually ran (e.g. due to invalid arguments or an infrastructure error during setup). This is tracked independently of GetUsageStats; a tool with no entry in GetUsageStats (because its container never once ran) can still appear here. Returns nil if the executor is nil.
func (*DockerToolExecutor) GetUsageStats ¶
func (d *DockerToolExecutor) GetUsageStats() map[string]ToolUsage
GetUsageStats returns aggregate execution statistics for all tools.
func (*DockerToolExecutor) IsToolExhausted ¶ added in v0.15.0
func (d *DockerToolExecutor) IsToolExhausted(toolName string) bool
IsToolExhausted reports whether the named tool has exceeded its maximum call limit. Returns false if the executor is nil or the tool has not been used.
func (*DockerToolExecutor) RegisterTool ¶
func (d *DockerToolExecutor) RegisterTool(tool *DockerTool)
RegisterTool registers a tool with the executor.
func (*DockerToolExecutor) ValidateImage ¶
func (d *DockerToolExecutor) ValidateImage(ctx context.Context, image string) error
ValidateImage ensures a Docker image required by the runtime is available locally.
func (*DockerToolExecutor) ValidateTaskServiceSupport ¶
func (d *DockerToolExecutor) ValidateTaskServiceSupport(ctx context.Context) error
ValidateTaskServiceSupport ensures the Docker daemon can run task-scoped services.
func (*DockerToolExecutor) ValidateTool ¶ added in v0.12.1
func (d *DockerToolExecutor) ValidateTool(ctx context.Context, cfg config.ToolConfig) error
ValidateTool ensures the Docker image referenced by the tool configuration is available locally.
type EvaluationTemplateData ¶
type EvaluationTemplateData struct {
// Seed is shared by every task attempt of the evaluation invocation.
Seed string
}
EvaluationTemplateData describes one evaluation invocation.
type ExecutionEnvironment ¶
type ExecutionEnvironment interface {
NewToolExecutor(context.Context) (*DockerToolExecutor, error)
}
ExecutionEnvironment creates ephemeral executors within one provider attempt.
type ExecutionTemplateData ¶
type ExecutionTemplateData struct {
// Evaluation describes the evaluation invocation.
Evaluation EvaluationTemplateData
// Task identifies the executed task.
Task NameTemplateData
// Provider identifies the provider executing the task.
Provider NameTemplateData
// Run identifies the provider run configuration executing the task.
Run NameTemplateData
}
ExecutionTemplateData identifies the evaluation, task, provider, and run of a task attempt in templates.
type NameTemplateData ¶
type NameTemplateData struct {
// Name is the configured name.
Name string
}
NameTemplateData exposes the configured name of a task, provider, or run.
type OutputCapture ¶ added in v0.20.0
type OutputCapture struct {
// Bytes is the total size of the output stream, regardless of Truncated.
Bytes int64
// Preview is a truncated prefix of the output stream, or nil if not captured
// (e.g. stdout on a successful call) or empty.
Preview *string
// Truncated indicates whether Preview was cut short of the full output.
Truncated bool
}
OutputCapture holds a size-limited preview of a tool call's output stream.
type TaskRuntime ¶
type TaskRuntime struct {
// contains filtered or unexported fields
}
TaskRuntime owns the Docker services and networks of one provider task attempt.
func NewTaskRuntime ¶
func NewTaskRuntime(ctx context.Context, logger logging.Logger, cfg TaskRuntimeConfig) (*TaskRuntime, error)
NewTaskRuntime starts the required services of one provider task attempt, each on its own internal network. If any service fails to start, all resources created so far are removed.
func (*TaskRuntime) Close ¶
func (r *TaskRuntime) Close(ctx context.Context) error
Close removes all service containers and networks. It is safe to call more than once.
func (*TaskRuntime) NewToolExecutor ¶
func (r *TaskRuntime) NewToolExecutor(ctx context.Context) (*DockerToolExecutor, error)
NewToolExecutor creates an ephemeral tool executor attached to this task runtime.
type TaskRuntimeConfig ¶
type TaskRuntimeConfig struct {
// Services lists the available service definitions.
Services []config.ServiceConfig
// RequiredServices names the services started for the attempt.
RequiredServices []string
// ServiceInputs supplies task inputs keyed by service name and input name.
ServiceInputs map[string]map[string]interface{}
// TemplateData is exposed to string service-input templates.
TemplateData ExecutionTemplateData
}
TaskRuntimeConfig describes the services of one provider task attempt.
type TextOrData ¶ added in v0.12.0
TextOrData is a constraint for types that can be written to files.
type ToolCallContext ¶ added in v0.20.0
type ToolCallContext struct {
// CallID is the provider's own identifier for this tool call, if the provider's API
// assigns one (e.g. OpenAI/Anthropic/DeepSeek/Mistral AI/xAI's tool_call/tool_use ID).
// Reusing the provider's own ID - rather than minting an unrelated one - means this same
// ID can also be found in any API error message that references the call. If empty,
// ExecuteTool generates one internally (a ULID) instead, so ToolCallSummary.CallID is
// never empty; note this means CallID's shape/format varies depending on whether - and
// how - the calling provider assigns its own IDs.
CallID string
// ConversationTurn is the 1-based conversation turn this call is being made during, or 0
// if unknown/not applicable.
ConversationTurn int
}
ToolCallContext carries optional caller-supplied metadata to attach to the ToolCallSummary recorded for a single ExecuteTool call. All fields are optional; a nil ToolCallContext (or a zero-value one) simply leaves the corresponding ToolCallSummary fields unset. New fields can be added here in the future without changing ExecuteTool's signature again.
type ToolCallSummary ¶ added in v0.20.0
type ToolCallSummary struct {
// Tool is the name of the tool this call invoked.
Tool string
// CallID is a unique identifier (ULID) for this call. It is also included in the
// prefix of every log line emitted while this call was in progress, so a specific
// invocation can be correlated between this summary and the corresponding log lines.
CallID string
// ConversationTurn is the 1-based conversation turn this call was made during, or 0 if
// unknown/not provided by the caller.
ConversationTurn int
// StartedAt is when this call began (start of setup, before the container ever runs).
StartedAt time.Time
// CompletedAt is when this call finished, successfully or not.
CompletedAt time.Time
// DurationNs is the wall-clock duration of the container's runtime (start through exit),
// the same measurement window as the aggregate ToolUsage.TotalDurationNs. It does not
// include setup (argument parsing, mounts, container creation) or teardown overhead, and
// stays nil when no container process ever ran (e.g. an infrastructure_error during setup).
DurationNs *int64
// WallTimeNs is the wall-clock duration of the entire call attempt, from setup
// through output retrieval - i.e. DurationNs plus setup/teardown overhead. Unlike
// DurationNs, this is always set, even for calls whose container never ran.
WallTimeNs int64
// ExitCode is the container's exit code, or nil if no exit code is known (e.g. setup never
// reached container start, or the run was aborted/cancelled before it could be observed).
ExitCode *int64
// TimedOut indicates the call was aborted due to exceeding its configured timeout.
TimedOut bool
// Status is one of: "success", "nonzero_exit", "empty_output", "timeout",
// "invalid_arguments", "infrastructure_error". Statuses are chosen to be meaningful for
// result analysis: "nonzero_exit", "empty_output", and "invalid_arguments" reflect the
// tool being used incorrectly (a plausible model/tool-usage issue), while
// "infrastructure_error" covers environment/tooling failures (container/filesystem setup,
// Docker runtime errors including cancellation, and log retrieval failures) that the
// model has no influence over and are not informative about the model or tool under test.
Status string
// Stdout is a size-limited capture of the call's standard output, or nil if no output was
// ever captured (e.g. an infrastructure_error before or during log retrieval).
Stdout *OutputCapture
// Stderr is a size-limited capture of the call's standard error, or nil if no output was
// ever captured (e.g. an infrastructure_error before or during log retrieval).
Stderr *OutputCapture
// ErrorMessage is a short explanation of the failure when Status is not "success".
ErrorMessage string
}
ToolCallSummary records the outcome of a single tool invocation.
type ToolUsage ¶
ToolUsage tracks aggregate execution statistics for a tool: CallCount and TotalDurationNs only reflect invocations whose container actually started running (regardless of exit code) - i.e. actual compute time spent, not every invocation attempt. An invocation that fails before the container ever starts (e.g. invalid arguments, or an infrastructure error during setup) does not affect these aggregates. See ToolCallSummary/GetCallSummaries for a complete per-invocation log that does include such attempts.