openai

package
v2.1.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 12, 2026 License: Apache-2.0 Imports: 19 Imported by: 0

Documentation

Overview

Package openai provides OpenAI LLM provider integration.

Package openai provides OpenAI LLM provider integration.

Package openai provides OpenAI LLM provider integration.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI LLM provider integration.

Package openai provides OpenAI Realtime API streaming support.

Package openai provides OpenAI Realtime API streaming support.

Index

Constants

View Source
const (
	// DefaultEmbeddingModel is the default model for embeddings
	DefaultEmbeddingModel = "text-embedding-3-small"

	// EmbeddingModelAda002 is the legacy ada-002 model
	EmbeddingModelAda002 = "text-embedding-ada-002"

	// EmbeddingModel3Small is the newer small model with better performance
	EmbeddingModel3Small = "text-embedding-3-small"

	// EmbeddingModel3Large is the large model with highest quality
	EmbeddingModel3Large = "text-embedding-3-large"
)

Embedding model constants

View Source
const (
	APIModeResponses   = "responses"   // New Responses API (v1/responses)
	APIModeCompletions = "completions" // Legacy Chat Completions API (v1/chat/completions)

)

API mode constants

View Source
const (
	// RealtimeAPIEndpoint is the base WebSocket endpoint for OpenAI Realtime API.
	RealtimeAPIEndpoint = "wss://api.openai.com/v1/realtime"

	// RealtimeBetaHeader is required for the Realtime API.
	RealtimeBetaHeader = "realtime=v1"

	// Default audio configuration for OpenAI Realtime API.
	// OpenAI Realtime uses 24kHz 16-bit PCM mono audio.
	DefaultRealtimeSampleRate = 24000
	DefaultRealtimeChannels   = 1
	DefaultRealtimeBitDepth   = 16
)

Realtime API constants

Variables

View Source
var ErrConnectionLost = errors.New("websocket connection lost")

ErrConnectionLost is the error type emitted on the Response channel when the WebSocket connection drops mid-session (network failure, heartbeat timeout, server-initiated close with a non-graceful code). Callers can check for this with errors.Is to distinguish recoverable connection loss from clean session termination or server-side errors.

Functions

func ParseServerEvent

func ParseServerEvent(data []byte) (interface{}, error)

ParseServerEvent parses a raw JSON message into the appropriate event type.

func RealtimeStreamingCapabilities

func RealtimeStreamingCapabilities() providers.StreamingCapabilities

RealtimeStreamingCapabilities returns the streaming capabilities for OpenAI Realtime API.

Types

type APIMode

type APIMode string

APIMode represents the OpenAI API mode to use

type ClientEvent

type ClientEvent struct {
	EventID string `json:"event_id,omitempty"`
	Type    string `json:"type"`
}

ClientEvent is the base structure for all client events.

type ConversationContent

type ConversationContent struct {
	Type       string `json:"type"` // "input_text", "input_audio", "text", "audio"
	Text       string `json:"text,omitempty"`
	Audio      string `json:"audio,omitempty"`      // Base64-encoded
	Transcript string `json:"transcript,omitempty"` // For audio content
}

ConversationContent represents content within a conversation item.

type ConversationItem

type ConversationItem struct {
	ID        string                `json:"id,omitempty"`
	Type      string                `json:"type"` // "message", "function_call", "function_call_output"
	Status    string                `json:"status,omitempty"`
	Role      string                `json:"role,omitempty"` // "user", "assistant", "system"
	Content   []ConversationContent `json:"content,omitempty"`
	CallID    string                `json:"call_id,omitempty"`   // For function_call_output
	Output    string                `json:"output,omitempty"`    // For function_call_output
	Name      string                `json:"name,omitempty"`      // For function_call
	Arguments string                `json:"arguments,omitempty"` // For function_call
}

ConversationItem represents an item in the conversation.

type ConversationItemCreateEvent

type ConversationItemCreateEvent struct {
	ClientEvent
	PreviousItemID string           `json:"previous_item_id,omitempty"`
	Item           ConversationItem `json:"item"`
}

ConversationItemCreateEvent adds an item to the conversation.

type ConversationItemCreatedEvent

type ConversationItemCreatedEvent struct {
	ServerEvent
	PreviousItemID string           `json:"previous_item_id"`
	Item           ConversationItem `json:"item"`
}

ConversationItemCreatedEvent confirms an item was added.

type ConversationItemInputAudioTranscriptionCompletedEvent

type ConversationItemInputAudioTranscriptionCompletedEvent struct {
	ServerEvent
	ItemID       string `json:"item_id"`
	ContentIndex int    `json:"content_index"`
	Transcript   string `json:"transcript"`
}

ConversationItemInputAudioTranscriptionCompletedEvent provides transcription.

type ConversationItemInputAudioTranscriptionFailedEvent

type ConversationItemInputAudioTranscriptionFailedEvent struct {
	ServerEvent
	ItemID       string      `json:"item_id"`
	ContentIndex int         `json:"content_index"`
	Error        ErrorDetail `json:"error"`
}

ConversationItemInputAudioTranscriptionFailedEvent indicates transcription failed.

type EmbeddingOption

type EmbeddingOption func(*EmbeddingProvider)

EmbeddingOption configures the EmbeddingProvider.

func WithEmbeddingAPIKey

func WithEmbeddingAPIKey(key string) EmbeddingOption

WithEmbeddingAPIKey sets the API key explicitly.

func WithEmbeddingBaseURL

func WithEmbeddingBaseURL(url string) EmbeddingOption

WithEmbeddingBaseURL sets a custom base URL (for Azure or proxies).

func WithEmbeddingHTTPClient

func WithEmbeddingHTTPClient(client *http.Client) EmbeddingOption

WithEmbeddingHTTPClient sets a custom HTTP client.

func WithEmbeddingModel

func WithEmbeddingModel(model string) EmbeddingOption

WithEmbeddingModel sets the embedding model.

func WithEmbeddingPlatformAuth

func WithEmbeddingPlatformAuth() EmbeddingOption

WithEmbeddingPlatformAuth marks the provider as authenticated by its HTTP client's transport (hyperscaler platform auth), so the empty-API-key guard in NewEmbeddingProvider is skipped.

type EmbeddingProvider

type EmbeddingProvider struct {
	*providers.BaseEmbeddingProvider
}

EmbeddingProvider implements embedding generation via OpenAI API.

func NewEmbeddingProvider

func NewEmbeddingProvider(opts ...EmbeddingOption) (*EmbeddingProvider, error)

NewEmbeddingProvider creates an OpenAI embedding provider.

func (*EmbeddingProvider) Embed

Embed generates embeddings for the given texts.

func (*EmbeddingProvider) EstimateCost

func (p *EmbeddingProvider) EstimateCost(tokens int) float64

EstimateCost estimates the cost for embedding the given number of tokens.

type ErrorDetail

type ErrorDetail struct {
	Type    string `json:"type"`
	Code    string `json:"code"`
	Message string `json:"message"`
	Param   string `json:"param,omitempty"`
	EventID string `json:"event_id,omitempty"`
}

ErrorDetail contains error information.

type ErrorEvent

type ErrorEvent struct {
	ServerEvent
	Error ErrorDetail `json:"error"`
}

ErrorEvent indicates an error occurred.

type InputAudioBufferAppendEvent

type InputAudioBufferAppendEvent struct {
	ClientEvent
	Audio string `json:"audio"` // Base64-encoded audio data
}

InputAudioBufferAppendEvent appends audio to the input buffer.

type InputAudioBufferClearEvent

type InputAudioBufferClearEvent struct {
	ClientEvent
}

InputAudioBufferClearEvent clears the audio buffer.

type InputAudioBufferClearedEvent

type InputAudioBufferClearedEvent struct {
	ServerEvent
}

InputAudioBufferClearedEvent confirms audio buffer was cleared.

type InputAudioBufferCommitEvent

type InputAudioBufferCommitEvent struct {
	ClientEvent
}

InputAudioBufferCommitEvent commits the audio buffer for processing.

type InputAudioBufferCommittedEvent

type InputAudioBufferCommittedEvent struct {
	ServerEvent
	PreviousItemID string `json:"previous_item_id"`
	ItemID         string `json:"item_id"`
}

InputAudioBufferCommittedEvent confirms audio buffer was committed.

type InputAudioBufferSpeechStartedEvent

type InputAudioBufferSpeechStartedEvent struct {
	ServerEvent
	AudioStartMs int    `json:"audio_start_ms"`
	ItemID       string `json:"item_id"`
}

InputAudioBufferSpeechStartedEvent indicates speech was detected.

type InputAudioBufferSpeechStoppedEvent

type InputAudioBufferSpeechStoppedEvent struct {
	ServerEvent
	AudioEndMs int    `json:"audio_end_ms"`
	ItemID     string `json:"item_id"`
}

InputAudioBufferSpeechStoppedEvent indicates speech ended.

type Provider

type Provider struct {
	providers.BaseProvider
	// contains filtered or unexported fields
}

OpenAIProvider implements the Provider interface for OpenAI

func NewProvider

func NewProvider(id, model, baseURL string, defaults providers.ProviderDefaults, includeRawOutput bool) *Provider

NewProvider creates a new OpenAI provider

func NewProviderFromConfig

func NewProviderFromConfig(cfg *ProviderConfig) *Provider

NewProviderFromConfig creates a provider from a full config struct.

func NewProviderWithConfig

func NewProviderWithConfig(
	id, model, baseURL string,
	defaults providers.ProviderDefaults,
	includeRawOutput bool,
	additionalConfig map[string]any,
) *Provider

NewProviderWithConfig creates a new OpenAI provider with additional configuration

func NewProviderWithCredential

func NewProviderWithCredential(
	id, model, baseURL string, defaults providers.ProviderDefaults,
	includeRawOutput bool, cred providers.Credential,
	platform string, platformConfig *providers.PlatformConfig,
) *Provider

NewProviderWithCredential creates a new OpenAI provider with explicit credential.

func NewProviderWithCredentialAndConfig

func NewProviderWithCredentialAndConfig(
	id, model, baseURL string, defaults providers.ProviderDefaults,
	includeRawOutput bool, cred providers.Credential, additionalConfig map[string]any,
	platform string, platformConfig *providers.PlatformConfig,
) *Provider

NewProviderWithCredentialAndConfig creates a new OpenAI provider with explicit credential and config.

func (*Provider) CalculateCost

func (p *Provider) CalculateCost(tokensIn, tokensOut, cachedTokens int) types.CostInfo

CalculateCost calculates detailed cost breakdown including optional cached tokens. Thin wrapper over costFromUsage, kept for the Provider interface's legacy signature (cachedTokens here means cache reads only — callers that also have a reasoning-token count, e.g. wire openAIUsage, should call costFromUsage directly instead so reasoning tokens aren't dropped).

func (*Provider) CreateStreamSession

func (p *Provider) CreateStreamSession(
	ctx context.Context,
	req *providers.StreamingInputConfig,
) (providers.StreamInputSession, error)

func (*Provider) EmitsLateInputTranscription

func (p *Provider) EmitsLateInputTranscription() bool

CreateStreamSession creates a new bidirectional streaming session with OpenAI Realtime API.

The session supports real-time audio input/output with the following features: - Bidirectional audio streaming (send and receive audio simultaneously) - Server-side voice activity detection (VAD) for automatic turn detection - Function/tool calling during the streaming session - Input and output audio transcription

Audio Format: OpenAI Realtime API uses 24kHz 16-bit PCM mono audio by default. The session automatically handles base64 encoding/decoding of audio data.

Concurrency bounds: If the provider has `stream_max_concurrent` configured, this call will block until a slot is available (respecting the caller's ctx) or return a rejection error recorded on promptkit_stream_concurrency_rejections_total. The slot is released when the session ends — whether via Close() or via the underlying Done() channel (e.g. context cancellation). This makes Realtime sessions subject to the same Phase 3 back-pressure as SSE streams.

Example usage:

session, err := provider.CreateStreamSession(ctx, &providers.StreamingInputConfig{
    Config: types.StreamingMediaConfig{
        Type:       types.ContentTypeAudio,
        SampleRate: 24000,
        Encoding:   "pcm16",
        Channels:   1,
    },
    SystemInstruction: "You are a helpful assistant.",
})

EmitsLateInputTranscription reports that OpenAI Realtime delivers the user's input transcription asynchronously (Whisper), after the assistant reply has begun — so the pipeline should reorder the transcript when transcription is enabled. Implements providers.LateInputTranscriber.

func (*Provider) GetMultimodalCapabilities

func (p *Provider) GetMultimodalCapabilities() providers.MultimodalCapabilities

GetMultimodalCapabilities returns OpenAI's multimodal capabilities

func (*Provider) GetStreamingCapabilities

func (p *Provider) GetStreamingCapabilities() providers.StreamingCapabilities

GetStreamingCapabilities returns detailed information about OpenAI's streaming support.

func (*Provider) Model

func (p *Provider) Model() string

Model returns the model name/identifier used by this provider.

func (*Provider) Predict

Predict sends a predict request to OpenAI

func (*Provider) PredictStream

func (p *Provider) PredictStream(ctx context.Context, req providers.PredictionRequest) (<-chan providers.StreamChunk, error)

PredictStream streams a predict response from OpenAI.

Bedrock note: AWS Bedrock's `invoke-with-response-stream` returns a binary event-stream protocol with OpenAI-format chunks inside — distinct from the SSE format the rest of the openai stream code expects. Until that path is wired (separate scanner needed), Bedrock falls back to a single non-streaming Predict response surfaced as one terminal chunk on the channel. Callers that need real per-token streaming for openai+bedrock should track that follow-up.

func (*Provider) SupportsPromptCaching

func (p *Provider) SupportsPromptCaching() bool

SupportsPromptCaching reports that OpenAI applies automatic prompt caching (no cache_control needed) for prompts above its minimum, so the pipeline's caching-stall warning applies here. Caching engages only when the request prefix is stable across rounds (see the deterministic tool-order fix).

func (*Provider) SupportsStreamInput

func (p *Provider) SupportsStreamInput() []string

SupportsStreamInput returns the media types supported for streaming input.

func (*Provider) SupportsStreaming

func (p *Provider) SupportsStreaming() bool

SupportsStreaming returns false for audio models on the Completions API because the Chat Completions streaming delta does not include audio data. Audio models only support non-streaming responses (Predict) or the Realtime API for bidirectional audio.

type ProviderConfig

type ProviderConfig struct {
	ID                string
	Model             string
	BaseURL           string
	Defaults          providers.ProviderDefaults
	IncludeRawOutput  bool
	Credential        providers.Credential
	AdditionalConfig  map[string]any
	Platform          string
	PlatformConfig    *providers.PlatformConfig
	UnsupportedParams []string
	Capabilities      []string
}

ProviderConfig holds all configuration for creating an OpenAI provider. Used by CreateProviderFromSpec and other callers that need full control.

type RateLimit

type RateLimit struct {
	Name         string  `json:"name"`
	Limit        int     `json:"limit"`
	Remaining    int     `json:"remaining"`
	ResetSeconds float64 `json:"reset_seconds"`
}

RateLimit contains rate limit details.

type RateLimitsUpdatedEvent

type RateLimitsUpdatedEvent struct {
	ServerEvent
	RateLimits []RateLimit `json:"rate_limits"`
}

RateLimitsUpdatedEvent provides rate limit information.

type RealtimeAudioConfig

type RealtimeAudioConfig struct {
	Input  *RealtimeAudioInput  `json:"input,omitempty"`
	Output *RealtimeAudioOutput `json:"output,omitempty"`
}

RealtimeAudioConfig groups the input and output audio configuration in the GA schema. Either side may be nil to leave server defaults in place.

type RealtimeAudioFormat

type RealtimeAudioFormat struct {
	Type string `json:"type"`
	Rate int    `json:"rate,omitempty"`
}

RealtimeAudioFormat is the GA-shape codec descriptor. For PCM, type is "audio/pcm" and rate is the sample rate in Hz (24000 for gpt-realtime).

type RealtimeAudioInput

type RealtimeAudioInput struct {
	Format        *RealtimeAudioFormat `json:"format,omitempty"`
	TurnDetection *TurnDetectionConfig `json:"turn_detection"`
	Transcription *TranscriptionConfig `json:"transcription,omitempty"`
}

RealtimeAudioInput configures the user-input side of the audio stream. TurnDetection uses a pointer-without-omitempty so the caller can send explicit null to disable server VAD (manual turn control).

type RealtimeAudioOutput

type RealtimeAudioOutput struct {
	Format *RealtimeAudioFormat `json:"format,omitempty"`
	Voice  string               `json:"voice,omitempty"`
	Speed  float64              `json:"speed,omitempty"`
}

RealtimeAudioOutput configures the model-output side of the audio stream.

type RealtimeSession

type RealtimeSession struct {

	// StreamPump provides Response(), BargeIn(), and the barge-in audio drop,
	// shared identically across providers. Embedded so those methods promote.
	*providers.StreamPump
	// contains filtered or unexported fields
}

RealtimeSession implements StreamInputSession for OpenAI Realtime API.

func NewRealtimeSession

func NewRealtimeSession(ctx context.Context, apiKey string, config *RealtimeSessionConfig) (*RealtimeSession, error)

NewRealtimeSession creates a new OpenAI Realtime streaming session.

func (*RealtimeSession) CancelResponse

func (s *RealtimeSession) CancelResponse() error

CancelResponse cancels an in-progress response.

func (*RealtimeSession) ClearAudioBuffer

func (s *RealtimeSession) ClearAudioBuffer() error

ClearAudioBuffer clears the current audio buffer.

func (*RealtimeSession) Close

func (s *RealtimeSession) Close() error

Close closes the session.

func (*RealtimeSession) CommitAudioBuffer

func (s *RealtimeSession) CommitAudioBuffer() error

CommitAudioBuffer commits the current audio buffer for processing.

func (*RealtimeSession) Config

Config returns the session configuration. Callers can use this to create a new session with the same settings after a connection loss:

if errors.Is(session.Error(), openai.ErrConnectionLost) {
    newSession, _ := openai.NewRealtimeSession(ctx, apiKey, session.Config())
}

Note: server-side conversation state is lost on reconnection — this only preserves the client-side configuration.

func (*RealtimeSession) Done

func (s *RealtimeSession) Done() <-chan struct{}

Done returns a channel that's closed when the session ends.

func (*RealtimeSession) EndInput

func (s *RealtimeSession) EndInput()

EndInput signals the end of user input.

In server_vad mode the server auto-commits when it detects end-of-speech in the audio buffer and (with create_response: true) auto-creates a response. Sending manual input_audio_buffer.commit + response.create from the client races with that auto-handling and causes the server to reject audio chunks with a cryptic "Invalid 'audio'... got an invalid value" error (verified via integration bisection). Callers that want server_vad behavior must include trailing silence in the streamed audio so the VAD can detect end-of-speech — pumpTTSChunks already does this via streamSilenceTail.

In manual turn-control mode (TurnDetection == nil, i.e. vad_disabled), the client owns turn boundaries: commit the buffer and trigger a response.

func (*RealtimeSession) Error

func (s *RealtimeSession) Error() error

Error returns any error that occurred during the session.

func (*RealtimeSession) SendChunk

func (s *RealtimeSession) SendChunk(ctx context.Context, chunk *types.MediaChunk) error

SendChunk sends an audio chunk to the server.

func (*RealtimeSession) SendSystemContext

func (s *RealtimeSession) SendSystemContext(ctx context.Context, text string) error

SendSystemContext sends a partial session.update that only modifies instructions, leaving codec/voice/VAD untouched. The GA API merges partial session.update events; type is included because GA requires it on every session.update.

func (*RealtimeSession) SendText

func (s *RealtimeSession) SendText(ctx context.Context, text string) error

SendText sends a text message and triggers a response.

func (*RealtimeSession) SendToolResponse

func (s *RealtimeSession) SendToolResponse(ctx context.Context, toolCallID, result string) error

SendToolResponse sends the result of a tool execution back to the model.

func (*RealtimeSession) SendToolResponses

func (s *RealtimeSession) SendToolResponses(ctx context.Context, responses []providers.ToolResponse) error

SendToolResponses sends multiple tool results at once (for parallel tool calls).

func (*RealtimeSession) TriggerResponse

func (s *RealtimeSession) TriggerResponse(config *ResponseConfig) error

TriggerResponse manually triggers a response from the model.

type RealtimeSessionConfig

type RealtimeSessionConfig struct {
	// Model specifies the model to use (e.g., "gpt-4o-realtime-preview").
	Model string

	// Modalities specifies the input/output modalities.
	// Valid values: "text", "audio"
	// Default: ["text", "audio"]
	Modalities []string

	// Instructions is the system prompt for the session.
	Instructions string

	// Voice selects the voice for audio output.
	// Options: "alloy", "echo", "fable", "onyx", "nova", "shimmer"
	// Default: "alloy"
	Voice string

	// InputAudioFormat specifies the format for input audio.
	// Options: "pcm16", "g711_ulaw", "g711_alaw"
	// Default: "pcm16"
	InputAudioFormat string

	// OutputAudioFormat specifies the format for output audio.
	// Options: "pcm16", "g711_ulaw", "g711_alaw"
	// Default: "pcm16"
	OutputAudioFormat string

	// InputAudioTranscription configures transcription of input audio.
	// If nil, input transcription is disabled.
	InputAudioTranscription *TranscriptionConfig

	// TurnDetection configures server-side voice activity detection.
	// If nil, VAD is disabled and turn management is manual.
	TurnDetection *TurnDetectionConfig

	// Tools defines available functions for the session.
	Tools []RealtimeToolDefinition

	// OutputSampleRate is the sample rate in Hz for output audio.
	// Default: 24000 (OpenAI Realtime API outputs 24kHz PCM).
	OutputSampleRate int

	// Temperature controls randomness (0.6-1.2, default 0.8).
	Temperature float64

	// MaxResponseOutputTokens limits response length.
	// Use "inf" for unlimited, or a specific number.
	MaxResponseOutputTokens interface{}
}

RealtimeSessionConfig configures a new OpenAI Realtime streaming session.

func DefaultRealtimeSessionConfig

func DefaultRealtimeSessionConfig() RealtimeSessionConfig

DefaultRealtimeSessionConfig returns sensible defaults for a Realtime session.

type RealtimeToolDef

type RealtimeToolDef struct {
	Type        string                 `json:"type"`
	Name        string                 `json:"name"`
	Description string                 `json:"description,omitempty"`
	Parameters  map[string]interface{} `json:"parameters,omitempty"`
}

RealtimeToolDef is the tool definition format for session config.

type RealtimeToolDefinition

type RealtimeToolDefinition struct {
	// Type is always "function" for function tools.
	Type string `json:"type"`

	// Name is the function name.
	Name string `json:"name"`

	// Description explains what the function does.
	Description string `json:"description,omitempty"`

	// Parameters is the JSON Schema for function parameters.
	Parameters map[string]interface{} `json:"parameters,omitempty"`
}

RealtimeToolDefinition defines a function available in the session.

type RealtimeWebSocket

type RealtimeWebSocket struct {
	// contains filtered or unexported fields
}

RealtimeWebSocket manages WebSocket connections for OpenAI Realtime API. It delegates transport concerns to the shared streaming.Conn.

func NewRealtimeWebSocket

func NewRealtimeWebSocket(model, apiKey string) *RealtimeWebSocket

NewRealtimeWebSocket creates a new WebSocket manager for OpenAI Realtime API.

func (*RealtimeWebSocket) Close

func (ws *RealtimeWebSocket) Close() error

Close closes the WebSocket connection gracefully.

func (*RealtimeWebSocket) Conn

func (ws *RealtimeWebSocket) Conn() *streaming.Conn

Conn returns the underlying streaming.Conn for use by the session layer.

func (*RealtimeWebSocket) ConnectWithRetry

func (ws *RealtimeWebSocket) ConnectWithRetry(ctx context.Context) error

ConnectWithRetry attempts to connect with exponential backoff.

func (*RealtimeWebSocket) IsClosed

func (ws *RealtimeWebSocket) IsClosed() bool

IsClosed returns whether the WebSocket is closed.

func (*RealtimeWebSocket) Receive

func (ws *RealtimeWebSocket) Receive(ctx context.Context) ([]byte, error)

Receive reads a message from the WebSocket with context support.

func (*RealtimeWebSocket) ReceiveLoop

func (ws *RealtimeWebSocket) ReceiveLoop(ctx context.Context, msgCh chan<- []byte) error

ReceiveLoop continuously reads messages and sends them to the provided channel.

func (*RealtimeWebSocket) Send

func (ws *RealtimeWebSocket) Send(msg interface{}) error

Send sends a message to the WebSocket.

func (*RealtimeWebSocket) StartHeartbeat

func (ws *RealtimeWebSocket) StartHeartbeat(ctx context.Context, interval time.Duration)

StartHeartbeat starts a goroutine that sends ping messages periodically.

type ResponseAudioDeltaEvent

type ResponseAudioDeltaEvent struct {
	ServerEvent
	ResponseID   string `json:"response_id"`
	ItemID       string `json:"item_id"`
	OutputIndex  int    `json:"output_index"`
	ContentIndex int    `json:"content_index"`
	Delta        string `json:"delta"` // Base64-encoded audio
}

ResponseAudioDeltaEvent provides streaming audio.

type ResponseAudioDoneEvent

type ResponseAudioDoneEvent struct {
	ServerEvent
	ResponseID   string `json:"response_id"`
	ItemID       string `json:"item_id"`
	OutputIndex  int    `json:"output_index"`
	ContentIndex int    `json:"content_index"`
}

ResponseAudioDoneEvent indicates audio streaming completed.

type ResponseAudioTranscriptDeltaEvent

type ResponseAudioTranscriptDeltaEvent struct {
	ServerEvent
	ResponseID   string `json:"response_id"`
	ItemID       string `json:"item_id"`
	OutputIndex  int    `json:"output_index"`
	ContentIndex int    `json:"content_index"`
	Delta        string `json:"delta"`
}

ResponseAudioTranscriptDeltaEvent provides streaming transcript.

type ResponseAudioTranscriptDoneEvent

type ResponseAudioTranscriptDoneEvent struct {
	ServerEvent
	ResponseID   string `json:"response_id"`
	ItemID       string `json:"item_id"`
	OutputIndex  int    `json:"output_index"`
	ContentIndex int    `json:"content_index"`
	Transcript   string `json:"transcript"`
}

ResponseAudioTranscriptDoneEvent indicates transcript completed.

type ResponseCancelEvent

type ResponseCancelEvent struct {
	ClientEvent
}

ResponseCancelEvent cancels an in-progress response.

type ResponseConfig

type ResponseConfig struct {
	Modalities        []string          `json:"modalities,omitempty"`
	Instructions      string            `json:"instructions,omitempty"`
	Voice             string            `json:"voice,omitempty"`
	OutputAudioFormat string            `json:"output_audio_format,omitempty"`
	Tools             []RealtimeToolDef `json:"tools,omitempty"`
	ToolChoice        interface{}       `json:"tool_choice,omitempty"`
	Temperature       float64           `json:"temperature,omitempty"`
	MaxOutputTokens   interface{}       `json:"max_output_tokens,omitempty"`
}

ResponseConfig configures a response.

type ResponseContentPartAddedEvent

type ResponseContentPartAddedEvent struct {
	ServerEvent
	ResponseID   string              `json:"response_id"`
	ItemID       string              `json:"item_id"`
	OutputIndex  int                 `json:"output_index"`
	ContentIndex int                 `json:"content_index"`
	Part         ConversationContent `json:"part"`
}

ResponseContentPartAddedEvent indicates content was added.

type ResponseContentPartDoneEvent

type ResponseContentPartDoneEvent struct {
	ServerEvent
	ResponseID   string              `json:"response_id"`
	ItemID       string              `json:"item_id"`
	OutputIndex  int                 `json:"output_index"`
	ContentIndex int                 `json:"content_index"`
	Part         ConversationContent `json:"part"`
}

ResponseContentPartDoneEvent indicates content part completed.

type ResponseCreateEvent

type ResponseCreateEvent struct {
	ClientEvent
	Response *ResponseConfig `json:"response,omitempty"`
}

ResponseCreateEvent triggers a response from the model.

type ResponseCreatedEvent

type ResponseCreatedEvent struct {
	ServerEvent
	Response ResponseInfo `json:"response"`
}

ResponseCreatedEvent indicates a response is starting.

type ResponseDoneEvent

type ResponseDoneEvent struct {
	ServerEvent
	Response ResponseInfo `json:"response"`
}

ResponseDoneEvent indicates a response completed.

type ResponseFunctionCallArgumentsDeltaEvent

type ResponseFunctionCallArgumentsDeltaEvent struct {
	ServerEvent
	ResponseID  string `json:"response_id"`
	ItemID      string `json:"item_id"`
	OutputIndex int    `json:"output_index"`
	CallID      string `json:"call_id"`
	Delta       string `json:"delta"`
}

ResponseFunctionCallArgumentsDeltaEvent provides streaming function args.

type ResponseFunctionCallArgumentsDoneEvent

type ResponseFunctionCallArgumentsDoneEvent struct {
	ServerEvent
	ResponseID  string `json:"response_id"`
	ItemID      string `json:"item_id"`
	OutputIndex int    `json:"output_index"`
	CallID      string `json:"call_id"`
	Name        string `json:"name"`
	Arguments   string `json:"arguments"`
}

ResponseFunctionCallArgumentsDoneEvent indicates function args completed.

type ResponseInfo

type ResponseInfo struct {
	ID            string             `json:"id"`
	Object        string             `json:"object"`
	Status        string             `json:"status"`
	StatusDetails interface{}        `json:"status_details"`
	Output        []ConversationItem `json:"output"`
	Usage         *UsageInfo         `json:"usage"`
}

ResponseInfo contains response details.

type ResponseOutputItemAddedEvent

type ResponseOutputItemAddedEvent struct {
	ServerEvent
	ResponseID  string           `json:"response_id"`
	OutputIndex int              `json:"output_index"`
	Item        ConversationItem `json:"item"`
}

ResponseOutputItemAddedEvent indicates an output item was added.

type ResponseOutputItemDoneEvent

type ResponseOutputItemDoneEvent struct {
	ServerEvent
	ResponseID  string           `json:"response_id"`
	OutputIndex int              `json:"output_index"`
	Item        ConversationItem `json:"item"`
}

ResponseOutputItemDoneEvent indicates an output item completed.

type ResponseTextDeltaEvent

type ResponseTextDeltaEvent struct {
	ServerEvent
	ResponseID   string `json:"response_id"`
	ItemID       string `json:"item_id"`
	OutputIndex  int    `json:"output_index"`
	ContentIndex int    `json:"content_index"`
	Delta        string `json:"delta"`
}

ResponseTextDeltaEvent provides streaming text.

type ResponseTextDoneEvent

type ResponseTextDoneEvent struct {
	ServerEvent
	ResponseID   string `json:"response_id"`
	ItemID       string `json:"item_id"`
	OutputIndex  int    `json:"output_index"`
	ContentIndex int    `json:"content_index"`
	Text         string `json:"text"`
}

ResponseTextDoneEvent indicates text streaming completed.

type ServerEvent

type ServerEvent struct {
	EventID string `json:"event_id"`
	Type    string `json:"type"`
}

ServerEvent is the base structure for all server events.

type SessionConfig

type SessionConfig struct {
	// Type distinguishes a speech-to-speech session ("realtime") from a
	// transcription-only session ("transcription"). Required by GA.
	Type             string               `json:"type,omitempty"`
	Instructions     string               `json:"instructions,omitempty"`
	OutputModalities []string             `json:"output_modalities,omitempty"`
	Audio            *RealtimeAudioConfig `json:"audio,omitempty"`
	Tools            []RealtimeToolDef    `json:"tools,omitempty"`
	ToolChoice       interface{}          `json:"tool_choice,omitempty"`
	MaxOutputTokens  interface{}          `json:"max_output_tokens,omitempty"`
}

SessionConfig is the session configuration sent in session.update. Shape matches the GA Realtime API: output_modalities replaces modalities; codec, voice, VAD, and transcription all moved into a nested audio.{input,output} object; the legacy beta-flat shape (with top-level voice, input_audio_format, etc.) is no longer accepted by gpt-realtime and other GA models.

type SessionCreatedEvent

type SessionCreatedEvent struct {
	ServerEvent
	Session SessionInfo `json:"session"`
}

SessionCreatedEvent is sent when the session is established.

type SessionInfo

type SessionInfo struct {
	ID                      string               `json:"id"`
	Object                  string               `json:"object"`
	Model                   string               `json:"model"`
	Modalities              []string             `json:"modalities"`
	Instructions            string               `json:"instructions"`
	Voice                   string               `json:"voice"`
	InputAudioFormat        string               `json:"input_audio_format"`
	OutputAudioFormat       string               `json:"output_audio_format"`
	InputAudioTranscription *TranscriptionConfig `json:"input_audio_transcription"`
	TurnDetection           *TurnDetectionConfig `json:"turn_detection"`
	Tools                   []RealtimeToolDef    `json:"tools"`
	Temperature             float64              `json:"temperature"`
	MaxResponseOutputTokens interface{}          `json:"max_response_output_tokens"`
}

SessionInfo contains session details.

type SessionUpdateEvent

type SessionUpdateEvent struct {
	ClientEvent
	Session SessionConfig `json:"session"`
}

SessionUpdateEvent updates session configuration.

type SessionUpdatedEvent

type SessionUpdatedEvent struct {
	ServerEvent
	Session SessionInfo `json:"session"`
}

SessionUpdatedEvent confirms a session update.

type ToolProvider

type ToolProvider struct {
	*Provider
}

ToolProvider extends OpenAIProvider with tool support

func NewToolProvider

func NewToolProvider(
	id, model, baseURL string,
	defaults providers.ProviderDefaults,
	includeRawOutput bool,
	additionalConfig map[string]any,
	unsupportedParams []string,
) *ToolProvider

NewToolProvider creates a new OpenAI provider with tool support

func NewToolProviderWithCredential

func NewToolProviderWithCredential(
	id, model, baseURL string, defaults providers.ProviderDefaults,
	includeRawOutput bool, additionalConfig map[string]any, cred providers.Credential,
	platform string, platformConfig *providers.PlatformConfig,
	unsupportedParams []string,
) *ToolProvider

NewToolProviderWithCredential creates an OpenAI tool provider with explicit credential.

func (*ToolProvider) BuildTooling

func (p *ToolProvider) BuildTooling(descriptors []*providers.ToolDescriptor) (providers.ProviderTools, error)

BuildTooling converts tool descriptors to OpenAI format. By default, strict mode is enabled for reliable argument generation. Set additional_config.strict_tools: false in provider config to disable.

func (*ToolProvider) PredictStreamWithTools

func (p *ToolProvider) PredictStreamWithTools(
	ctx context.Context,
	req providers.PredictionRequest,
	tools interface{},
	toolChoice string,
) (<-chan providers.StreamChunk, error)

PredictStreamWithTools performs a streaming predict request with tool support.

Bedrock note: same fallback as PredictStream — Bedrock's streaming endpoint uses binary event-stream framing distinct from SSE; we run a single non-streaming call and surface it as one terminal chunk.

func (*ToolProvider) PredictWithTools

PredictWithTools performs a prediction request with tool support

type TranscriptionConfig

type TranscriptionConfig struct {
	// Model specifies the transcription model.
	// Default: "whisper-1"
	Model string `json:"model,omitempty"`
}

TranscriptionConfig configures audio transcription.

type TurnDetectionConfig

type TurnDetectionConfig struct {
	// Type specifies the VAD type.
	// Options: "server_vad", "semantic_vad"
	Type string `json:"type"`

	// Threshold is the activation threshold (0.0-1.0).
	// Default: 0.5
	Threshold float64 `json:"threshold,omitempty"`

	// PrefixPaddingMs is audio padding before speech in milliseconds.
	// Default: 300
	PrefixPaddingMs int `json:"prefix_padding_ms,omitempty"`

	// SilenceDurationMs is silence duration to detect end of speech.
	// Default: 500
	SilenceDurationMs int `json:"silence_duration_ms,omitempty"`

	// CreateResponse determines if a response is automatically created
	// when speech ends. Default: true
	CreateResponse bool `json:"create_response,omitempty"`
}

TurnDetectionConfig configures server-side VAD.

type UsageInfo

type UsageInfo struct {
	TotalTokens       int `json:"total_tokens"`
	InputTokens       int `json:"input_tokens"`
	OutputTokens      int `json:"output_tokens"`
	InputTokenDetails struct {
		CachedTokens int `json:"cached_tokens"`
		TextTokens   int `json:"text_tokens"`
		AudioTokens  int `json:"audio_tokens"`
	} `json:"input_token_details"`
	OutputTokenDetails struct {
		TextTokens  int `json:"text_tokens"`
		AudioTokens int `json:"audio_tokens"`
	} `json:"output_token_details"`
}

UsageInfo contains token usage information.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL