responsesexecutor

package
v0.10.91 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 16, 2026 License: Apache-2.0 Imports: 27 Imported by: 0

README

Native Responses executor

This executor runs the openai-responses protocol selected by the model router. It receives a configured responses.ResponseService; it doesn't discover credentials, select a provider, or infer a protocol from the model name.

Resolved route → typed Responses binding → native streaming conversation
                                              │
                      completed function calls ↓
                          client tools → result validators → final result
                              │
                              └─ function_call_output + native reasoning
                                 in the next request

Requests use store: false and replay native input/output items instead of previous_response_id. Encrypted reasoning content is preserved on the wire but removed from opt-in trace payloads. Streaming text and argument deltas are consumed through the SDK; only a completed response can dispatch tools. An interrupted stream can't execute a partial function call.

The executor uses the shared terminal-submission gate. Ordinary tools finish before terminal validators run, including when the terminal call appears first in a parallel batch. Only an accepted, validated submission commits a result.

Limits and recovery

  • The defaults are 200 turns, 32,768 output tokens per request, and 10 concurrent tools. Set concurrency to 1 when tool handlers share unsynchronized state.
  • Requests inherit the execution deadline by default. Set RequestTimeout to bound each streaming HTTP attempt, or ExecutionTimeout to change the 30-minute conversation deadline. Neither setting extends a shorter caller deadline. Negative durations are rejected. These are total, not idle, timeouts. For routed agents, use metaagent.Config.ResponsesRequestTimeout and ResponsesExecutionTimeout; other routed protocols reject these settings.
  • Requests and individual tool results are limited to 8 MiB. Streams are limited to 16 MiB and 32,768 events; completed payloads are limited to 8 MiB.
  • A completed response can contain at most 128 function calls. Repeated call IDs and unsupported output/event types fail closed before dispatch.
  • Three consecutive turns without usable tool work or an accepted submission end the execution. This bounds correction of invalid JSON, invalid terminal payloads, unknown tools, and refusals.
  • HTTP 429, 500, 502, 503, and 504 responses can retry twice before a stream starts. Partial streams don't retry: completion and usage may be unknown. The SDK's own retries are disabled. Error diagnostics omit raw provider bodies, messages, URLs, and headers. Before streaming starts, HTTP failures retain allowlisted error codes and request IDs with a canonical UUID or req_ plus 32 hexadecimal digits. Unknown codes and other ID formats are omitted. A recognized access denial suggests checking IAM, model access, and retention eligibility; it does not identify which check failed. Use the request ID for provider-side support.
  • Terminal response.failed, response.incomplete, and error events retain their event type, allowlisted error code or incomplete reason, and validated request and response IDs. Response IDs must be canonical UUIDs or resp_ followed by 32, 48, or 64 hexadecimal digits. Unknown codes, reasons, and ID formats are omitted. These diagnostics appear in execution and turn errors, but never include provider messages or partial output. Even a terminal rate_limit_exceeded or server_error event doesn't retry. Usage may be unavailable when the stream fails.

Choose timeout budgets before comparing models, allowing for reasoning and the configured output-token limit. Deadline expiration remains an execution error, not a quality score or a reason to replay a potentially billable partial stream. The 8 MiB request cap includes the entire replayed conversation and encrypted reasoning. Exceeding it is also a non-retryable execution-limit failure.

Low, medium, and high reasoning effort are sent directly. XHigh and Max clamp to high, matching the existing OpenAI-compatible executor's supported scale. Explicit thinking budgets, sampling controls, explicit prompt-cache boundaries, suspend/resume, hosted tools, and configurable refusal nudges aren't advertised.

Observability and validation

Turn traces include the route's provider, logical model, protocol, provider model ID, input/output tokens, cache-read tokens, and reasoning tokens. Reasoning is already included in output tokens; cache reads are already included in input tokens. Don't add either subset again when calculating totals. The nullable turns.reasoning_tokens schema addition must be deployed before the first Responses run. A recorder that rejects unknown fields can reject the entire trace row if its table still uses the older schema. Verify the deployed table schema before collecting evaluation results. This change doesn't add model prices or dashboard charts.

The HTTP/SSE fixture tests exercise the actual SDK decoder, native continuation, reasoning-state replay, parallel terminal ordering, schema rejection, malformed JSON recovery, cancellation, attribution, and error redaction. They don't prove that an AWS account is entitled to a model or that live retention settings are acceptable. The protected evaluation workflow must check those prerequisites before inference.

Documentation

Overview

Package responsesexecutor runs stateless, streaming Responses conversations with client-side tools and validated terminal submissions. Clients and route attribution are supplied by the caller; this package never loads credentials.

Index

Examples

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type Config

type Config[Response any] struct {
	Model               string
	Attribution         agenttrace.Attribution
	UserPrompt          *promptbuilder.Prompt
	SystemInstructions  *promptbuilder.Prompt
	UserPromptSuffix    *promptbuilder.Prompt
	MaxTurns            int
	ToolCallConcurrency int
	MaxTokens           int64
	Effort              effort.Level
	Submit              submitresult.Options[Response]
	ResultValidators    []callbacks.ResultValidator[Response]
	ResourceLabels      map[string]string

	// RequestTimeout bounds each streaming HTTP attempt. Zero inherits the
	// execution deadline. This is a total timeout, not an idle timeout.
	RequestTimeout time.Duration

	// ExecutionTimeout bounds the conversation, including tools and retries.
	// Zero selects 30 minutes. A shorter caller deadline always wins.
	ExecutionTimeout time.Duration
}

Config configures a credential-free Responses executor. Zero numeric limits select defaults: 200 turns, 32768 output tokens, and 10 concurrent tools. Concurrent tools must synchronize shared state themselves. Reasoning and cache tokens are subsets of output and input tokens, respectively, not extra usage.

type Interface

type Interface[Request promptbuilder.Bindable, Response any] interface {
	Execute(context.Context, Request, map[string]toolcall.Tool[Response]) (Response, error)
}

Interface executes a native Responses conversation with provider-neutral tools.

func New

func New[Request promptbuilder.Bindable, Response any](client responses.ResponseService, cfg Config[Response]) (Interface[Request, Response], error)

New constructs a native Responses executor using an already configured client. It does not access the network or inspect ambient authentication variables.

Example
package main

import (
	"fmt"
	"time"

	"chainguard.dev/driftlessaf/agents/agenttrace"
	"chainguard.dev/driftlessaf/agents/executor/openai/responsesexecutor"
	"chainguard.dev/driftlessaf/agents/promptbuilder"
	"github.com/openai/openai-go/responses"
)

type request struct{}

func (request) Bind(p *promptbuilder.Prompt) (*promptbuilder.Prompt, error) { return p, nil }

func main() {
	type answer struct {
		Summary string `json:"summary"`
	}
	prompt, err := promptbuilder.NewPrompt("Inspect the supplied synthetic fixture and submit your result.")
	if err != nil {
		panic(err)
	}
	// The application supplies a transport-configured service. Construction
	// itself never loads credentials or makes an inference request.
	var service responses.ResponseService
	_, err = responsesexecutor.New[request](service, responsesexecutor.Config[answer]{
		Model: "us.openai.gpt-5.6-sol", UserPrompt: prompt,
		Attribution: agenttrace.Attribution{ProviderName: "aws.bedrock", System: "aws.bedrock", LogicalModel: "gpt-5.6-sol", Protocol: "openai-responses"},
		MaxTurns:    5, MaxTokens: 2048, ToolCallConcurrency: 1,
		RequestTimeout:   15 * time.Minute,
		ExecutionTimeout: time.Hour,
	})
	fmt.Println(err)
}
Output:
<nil>

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL