Documentation
¶
Overview ¶
Package chatcompletionexecutor provides a multi-turn conversation executor for OpenAI-compatible chat completion APIs, including Vertex AI's partner model endpoint.
Prefer responsesexecutor for providers that support the Responses API. Use this package for providers that require Chat Completions. See the Responses executor documentation for capabilities and configuration.
The executor manages the full conversation lifecycle: sending prompts, processing tool calls, recording metrics, and extracting structured results. It mirrors the claudeexecutor and googleexecutor patterns, enabling the metaagent to route to any model available via the OpenAI chat completions API.
Basic Usage ¶
client := openai.NewClient(
option.WithBaseURL(vertexBaseURL),
option.WithHTTPClient(authedClient),
option.WithAPIKey("placeholder"),
)
prompt := promptbuilder.MustParse("Analyze {{.Input}}")
exec, err := chatcompletionexecutor.New[MyRequest, MyResponse](client, prompt,
chatcompletionexecutor.WithModel[MyRequest, MyResponse]("deepseek-ai/deepseek-v3.2-maas"),
chatcompletionexecutor.WithMaxTokens[MyRequest, MyResponse](32768),
chatcompletionexecutor.WithTemperature[MyRequest, MyResponse](0.2),
)
Options ¶
- WithAttribution: set canonical and compatibility route attribution
- WithModel: set the model name (required for Vertex AI partner models)
- WithMaxTokens: set the maximum completion tokens
- WithTemperature: set the sampling temperature (0.0–2.0)
- WithoutTemperature: omit temperature when an explicit route disallows sampling parameters
- WithEffort: set the reasoning effort for reasoning models (xhigh/max clamp to high)
- WithMaxTurns: set the maximum conversation turns before aborting
- WithSystemInstructions: set the system prompt
- WithUserPromptSuffix: append a static prompt to the built user prompt
- WithSubmitResultProvider: register the submit_result tool for structured output
- WithRetryConfig: configure retry behavior for transient API errors
- WithResourceLabels: set labels for observability attribution
Thinking Models ¶
Models that return reasoning_content in their responses (e.g. kimi-k2-thinking-maas) are supported. The executor captures reasoning content into the agent trace automatically.
Submit Result Redirect ¶
When a submit_result tool is configured but the model responds with text instead of calling the tool, the executor sends a redirect message asking the model to call submit_result. Unlike the claudeexecutor, the chatcompletionexecutor does not use a forced tool_choice for the redirect — some models (e.g. reasoning models) return 400 on named tool_choice constraints.
Index ¶
- Constants
- type Interface
- type Option
- func WithAttribution[Request promptbuilder.Bindable, Response any](attribution agenttrace.Attribution) Option[Request, Response]
- func WithEffort[Request promptbuilder.Bindable, Response any](level effort.Level) Option[Request, Response]
- func WithMaxTokens[Request promptbuilder.Bindable, Response any](tokens int64) Option[Request, Response]
- func WithMaxTurns[Request promptbuilder.Bindable, Response any](turns int) Option[Request, Response]
- func WithModel[Request promptbuilder.Bindable, Response any](model string) Option[Request, Response]
- func WithProvider[Request promptbuilder.Bindable, Response any](provider Provider) Option[Request, Response]
- func WithResourceLabels[Request promptbuilder.Bindable, Response any](labels map[string]string) Option[Request, Response]
- func WithResultValidator[Request promptbuilder.Bindable, Response any](v callbacks.ResultValidator[Response]) Option[Request, Response]
- func WithRetryConfig[Request promptbuilder.Bindable, Response any](cfg retry.RetryConfig) Option[Request, Response]
- func WithSubmitResultProvider[Request promptbuilder.Bindable, Response any](provider SubmitResultProvider[Response]) Option[Request, Response]
- func WithSystemInstructions[Request promptbuilder.Bindable, Response any](prompt *promptbuilder.Prompt) Option[Request, Response]
- func WithTemperature[Request promptbuilder.Bindable, Response any](temp float64) Option[Request, Response]
- func WithTokenLimitParameter[Request promptbuilder.Bindable, Response any](parameter TokenLimitParameter) Option[Request, Response]
- func WithToolCallConcurrency[Request promptbuilder.Bindable, Response any](n int) Option[Request, Response]
- func WithUserPromptSuffix[Request promptbuilder.Bindable, Response any](suffix *promptbuilder.Prompt) Option[Request, Response]
- func WithoutTemperature[Request promptbuilder.Bindable, Response any]() Option[Request, Response]
- type Provider
- type SubmitResultProvider
- type TokenLimitParameter
Examples ¶
Constants ¶
const DefaultMaxTurns = 200
DefaultMaxTurns is the default maximum number of conversation turns before aborting.
const DefaultToolCallConcurrency = 10
DefaultToolCallConcurrency is the default bound on how many of a single turn's tool calls run concurrently. Models routinely emit several independent tool calls in one turn (parallel tool calls); dispatching their handlers concurrently cuts wall-clock latency. Override with WithToolCallConcurrency — a value of 1 restores strictly sequential dispatch.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Interface ¶
type Interface[Request promptbuilder.Bindable, Response any] interface { // Execute runs the agent conversation with the given request and tools. Execute(ctx context.Context, request Request, tools map[string]openaistool.Metadata[Response]) (Response, error) }
Interface is the public interface for OpenAI-compatible agent execution.
func New ¶
func New[Request promptbuilder.Bindable, Response any]( client openai.Client, prompt *promptbuilder.Prompt, opts ...Option[Request, Response], ) (Interface[Request, Response], error)
New creates a new OpenAI-compatible executor.
Example ¶
package main
import (
"context"
"fmt"
"github.com/chainguard-dev/clog"
"chainguard.dev/driftlessaf/agents/executor/openai/chatcompletionexecutor"
"chainguard.dev/driftlessaf/agents/promptbuilder"
"github.com/openai/openai-go"
"github.com/openai/openai-go/option"
)
type myRequest struct {
Input string
}
func (r myRequest) Bind(p *promptbuilder.Prompt) (*promptbuilder.Prompt, error) {
return p.BindJSON("input", r.Input)
}
type myResponse struct {
Summary string `json:"summary"`
}
func main() {
ctx := context.Background()
prompt := promptbuilder.MustNewPrompt("Summarize: {{input}}")
client := openai.NewClient(
option.WithAPIKey("placeholder"),
)
exec, err := chatcompletionexecutor.New[myRequest, myResponse](client, prompt,
chatcompletionexecutor.WithModel[myRequest, myResponse]("google/gemini-3.5-flash"),
chatcompletionexecutor.WithMaxTokens[myRequest, myResponse](8192),
chatcompletionexecutor.WithTemperature[myRequest, myResponse](0.1),
)
if err != nil {
clog.FatalContextf(ctx, "%v", err)
}
fmt.Printf("executor created: %v\n", exec != nil)
}
Output: executor created: true
type Option ¶
type Option[Request promptbuilder.Bindable, Response any] func(*executor[Request, Response]) error
Option is a functional option for configuring the executor.
func WithAttribution ¶
func WithAttribution[Request promptbuilder.Bindable, Response any](attribution agenttrace.Attribution) Option[Request, Response]
WithAttribution sets explicit route attribution for metrics, turn spans, and serialized trace turns. It is provider-extensible: callers supply the canonical and legacy provider names directly, without registering them in this executor. All fields are required.
func WithEffort ¶
func WithEffort[Request promptbuilder.Bindable, Response any](level effort.Level) Option[Request, Response]
WithEffort sets the provider-neutral reasoning-effort level, sent as the reasoning_effort request parameter. OpenAI's scale is a subset of the shared one — low, medium, and high — so effort.XHigh and effort.Max clamp to "high". The parameter is only valid for reasoning models; sending it to a non-reasoning model is rejected by the API, so leave it unset for those.
func WithMaxTokens ¶
func WithMaxTokens[Request promptbuilder.Bindable, Response any](tokens int64) Option[Request, Response]
WithMaxTokens sets the maximum completion tokens.
func WithMaxTurns ¶
func WithMaxTurns[Request promptbuilder.Bindable, Response any](turns int) Option[Request, Response]
WithMaxTurns sets the maximum number of conversation turns before aborting.
Example ¶
package main
import (
"context"
"fmt"
"github.com/chainguard-dev/clog"
"chainguard.dev/driftlessaf/agents/executor/openai/chatcompletionexecutor"
"chainguard.dev/driftlessaf/agents/promptbuilder"
"github.com/openai/openai-go"
"github.com/openai/openai-go/option"
)
type myRequest struct {
Input string
}
func (r myRequest) Bind(p *promptbuilder.Prompt) (*promptbuilder.Prompt, error) {
return p.BindJSON("input", r.Input)
}
type myResponse struct {
Summary string `json:"summary"`
}
func main() {
ctx := context.Background()
prompt := promptbuilder.MustNewPrompt("Process: {{input}}")
client := openai.NewClient(
option.WithAPIKey("placeholder"),
)
exec, err := chatcompletionexecutor.New[myRequest, myResponse](client, prompt,
chatcompletionexecutor.WithMaxTurns[myRequest, myResponse](50),
)
if err != nil {
clog.FatalContextf(ctx, "%v", err)
}
fmt.Printf("executor created: %v\n", exec != nil)
}
Output: executor created: true
func WithModel ¶
func WithModel[Request promptbuilder.Bindable, Response any](model string) Option[Request, Response]
WithModel sets the model name.
Example ¶
package main
import (
"context"
"fmt"
"github.com/chainguard-dev/clog"
"chainguard.dev/driftlessaf/agents/executor/openai/chatcompletionexecutor"
"chainguard.dev/driftlessaf/agents/promptbuilder"
"github.com/openai/openai-go"
"github.com/openai/openai-go/option"
)
type myRequest struct {
Input string
}
func (r myRequest) Bind(p *promptbuilder.Prompt) (*promptbuilder.Prompt, error) {
return p.BindJSON("input", r.Input)
}
type myResponse struct {
Summary string `json:"summary"`
}
func main() {
ctx := context.Background()
prompt := promptbuilder.MustNewPrompt("Analyze: {{input}}")
client := openai.NewClient(
option.WithAPIKey("placeholder"),
)
exec, err := chatcompletionexecutor.New[myRequest, myResponse](client, prompt,
chatcompletionexecutor.WithModel[myRequest, myResponse]("deepseek-ai/deepseek-v3.2-maas"),
)
if err != nil {
clog.FatalContextf(ctx, "%v", err)
}
fmt.Printf("executor created: %v\n", exec != nil)
}
Output: executor created: true
func WithProvider ¶
func WithProvider[Request promptbuilder.Bindable, Response any](provider Provider) Option[Request, Response]
WithProvider declares which backend serves the executor's requests so metrics and trace turns remain attributable. The default is ProviderOpenAICompatible, preserving existing behavior.
func WithResourceLabels ¶
func WithResourceLabels[Request promptbuilder.Bindable, Response any](labels map[string]string) Option[Request, Response]
WithResourceLabels sets labels for observability attribution. Automatically includes default labels from environment variables:
- service_name: from K_SERVICE, falling back to CLOUD_RUN_JOB (defaults to "unknown")
- product: from CHAINGUARD_PRODUCT (defaults to "unknown")
- team: from CHAINGUARD_TEAM (defaults to "unknown")
func WithResultValidator ¶
func WithResultValidator[Request promptbuilder.Bindable, Response any](v callbacks.ResultValidator[Response]) Option[Request, Response]
WithResultValidator registers a validator that gates the terminal submit tool. When the model calls the submit tool with a payload that parses, every registered validator runs concurrently against the parsed response; any findings reject the submission back to the model as the tool's result — the loop continues until a submission passes — and a validator error aborts the run. Repeatable: each call appends a validator, and their findings are concatenated in registration order. Only meaningful when a submit tool is configured via WithSubmitResultProvider; without one there is nothing to gate.
func WithRetryConfig ¶
func WithRetryConfig[Request promptbuilder.Bindable, Response any](cfg retry.RetryConfig) Option[Request, Response]
WithRetryConfig sets the retry configuration for transient API errors.
func WithSubmitResultProvider ¶
func WithSubmitResultProvider[Request promptbuilder.Bindable, Response any](provider SubmitResultProvider[Response]) Option[Request, Response]
WithSubmitResultProvider registers the submit_result tool.
func WithSystemInstructions ¶
func WithSystemInstructions[Request promptbuilder.Bindable, Response any](prompt *promptbuilder.Prompt) Option[Request, Response]
WithSystemInstructions sets the system prompt.
func WithTemperature ¶
func WithTemperature[Request promptbuilder.Bindable, Response any](temp float64) Option[Request, Response]
WithTemperature sets the sampling temperature (0.0–2.0).
Example ¶
package main
import (
"context"
"fmt"
"github.com/chainguard-dev/clog"
"chainguard.dev/driftlessaf/agents/executor/openai/chatcompletionexecutor"
"chainguard.dev/driftlessaf/agents/promptbuilder"
"github.com/openai/openai-go"
"github.com/openai/openai-go/option"
)
type myRequest struct {
Input string
}
func (r myRequest) Bind(p *promptbuilder.Prompt) (*promptbuilder.Prompt, error) {
return p.BindJSON("input", r.Input)
}
type myResponse struct {
Summary string `json:"summary"`
}
func main() {
ctx := context.Background()
prompt := promptbuilder.MustNewPrompt("Summarize: {{input}}")
client := openai.NewClient(
option.WithAPIKey("placeholder"),
)
exec, err := chatcompletionexecutor.New[myRequest, myResponse](client, prompt,
chatcompletionexecutor.WithTemperature[myRequest, myResponse](0.5),
)
if err != nil {
clog.FatalContextf(ctx, "%v", err)
}
fmt.Printf("executor created: %v\n", exec != nil)
}
Output: executor created: true
func WithTokenLimitParameter ¶
func WithTokenLimitParameter[Request promptbuilder.Bindable, Response any](parameter TokenLimitParameter) Option[Request, Response]
WithTokenLimitParameter selects the token-limit request field. The default is TokenLimitMaxCompletionTokens, preserving existing behavior.
func WithToolCallConcurrency ¶
func WithToolCallConcurrency[Request promptbuilder.Bindable, Response any](n int) Option[Request, Response]
WithToolCallConcurrency bounds how many of a single turn's tool calls run concurrently when the model emits more than one in a turn (parallel tool calls). Defaults to DefaultToolCallConcurrency.
Tool result messages are always appended in the order the model emitted the calls. A value of 1 forces strictly sequential dispatch. Set it to 1 for agents whose tool handlers mutate shared state (a worktree, a cache) without their own synchronization; concurrent dispatch is otherwise safe because handlers share only the trace, which is concurrency-safe.
Note: some OpenAI-compatible models disable parallel tool calls when strict structured output is in force, in which case the model emits one tool call per turn and this option has no effect.
func WithUserPromptSuffix ¶
func WithUserPromptSuffix[Request promptbuilder.Bindable, Response any](suffix *promptbuilder.Prompt) Option[Request, Response]
WithUserPromptSuffix appends a static, operator-authored prompt to the end of the built user prompt, separated by a blank line. It is the OpenAI-compatible counterpart of the Claude executor's user-prompt-suffix option: agents that share one large payload but vary a small trailing instruction (for example multi-pass reviewers examining one changeset through different lenses) keep the payload in the main prompt and the varying instruction in the suffix. The OpenAI-compatible API has no per-block prompt-cache semantics, so the suffix is simply concatenated and there is no cache-shaping side effect. The suffix must be fully bound by the caller; the request is never bound into it.
func WithoutTemperature ¶
func WithoutTemperature[Request promptbuilder.Bindable, Response any]() Option[Request, Response]
WithoutTemperature omits sampling temperature from provider requests. Explicit routes use this when their effective capabilities narrow sampling parameters out, even if the logical model family normally supports them. The executor also omits temperature by default; WithTemperature opts in.
type Provider ¶
type Provider string
Provider identifies the backend serving an OpenAI-compatible request. Provider attribution is explicit because model slugs are not sufficient to distinguish Vertex-hosted models from external OpenAI-compatible services.
type SubmitResultProvider ¶
type SubmitResultProvider[Response any] func() (openaistool.SubmitMetadata[Response], error)
SubmitResultProvider constructs tool metadata for submit_result.
type TokenLimitParameter ¶
type TokenLimitParameter string
TokenLimitParameter selects the request field used to cap output tokens. OpenAI recommends max_completion_tokens, while some compatible providers, including Baseten Model APIs, require max_tokens.
const ( // TokenLimitMaxCompletionTokens sends the OpenAI max_completion_tokens field. TokenLimitMaxCompletionTokens TokenLimitParameter = "max_completion_tokens" // TokenLimitMaxTokens sends the legacy-compatible max_tokens field. TokenLimitMaxTokens TokenLimitParameter = "max_tokens" )