Documentation
¶
Overview ¶
Package speechkit provides the public SDK for embedding SpeechKit voice capture, transcription, and assist/voice-agent pipelines into host applications.
Surface ¶
The kernel exposes three strict interaction modes:
- Dictation — speech to text only, no AI rewriting.
- Assist — speech (or text) to a one-shot result, with optional TTS.
- Voice Agent — realtime audio-to-audio dialogue.
The Mode enum carries two further constants that are capability surfaces, not interaction modes: ModeTTS exposes Text-to-Speech as a model-selection axis (its IntelligenceVoiceOutput contract is strictly text in, audio out), and ModeNone means no mode is selected.
Each mode is constructed via a small subpackage so host apps depend only on what they use:
- github.com/kombifyio/SpeechKit/pkg/speechkit/dictation
- github.com/kombifyio/SpeechKit/pkg/speechkit/assist
- github.com/kombifyio/SpeechKit/pkg/speechkit/voiceagent
- github.com/kombifyio/SpeechKit/pkg/speechkit/agentkit (tool registry, session memory, lifecycle hooks for Voice Agent hosts)
- github.com/kombifyio/SpeechKit/pkg/speechkit/client (HTTP client for talking to a remote SpeechKit Server)
Central types in this package ¶
Runtime owns shared state and the event channel that host apps read from. Engine is the full voice pipeline; RecordingController and TranscriptionWorker can be composed independently for custom pipelines. [Catalog] exposes the provider/model/mode metadata for setup UIs and readiness checks; RuntimePolicy lets the host pin profiles and gate fallbacks.
Stability ¶
pkg/speechkit is the OSS public surface. Symbols here follow semver from v1.0 onward. Before v1.0 the surface may still evolve — see CHANGELOG.md and the release notes for breaking-change calls.
Package-level documentation lives in doc.go.
Index ¶
- Constants
- Variables
- func CountTerminalSentences(text string) int
- func DefaultProviderAuthRequirement(profile ProviderProfile) string
- func DefaultProviderTransport(profile ProviderProfile) string
- func JoinTranscriptFragments(parts ...string) string
- func LiveInjectFragment(previousSession uint64, previousTail, next string, session uint64) (fragment, tail string, nextSession uint64)
- func LowConfidenceWords(words []WordConfidence, threshold float64) (terms []string, minConfidence float64)
- func MissingFreshnessReports(rows []ProviderModelDescriptor) []string
- func NeedsLiveInjectSpace(prev, next string) bool
- func NormalizeProviderID(provider string) string
- func NormalizeProviderProfileID(profileID string) string
- func PCMDurationSecs(pcm []byte) float64
- func PCMToWAV(pcm []byte) []byte
- func ProviderCredentialTarget(profile ProviderProfile) string
- func ProviderIDForExecutionMode(mode ExecutionMode) string
- func ProviderIDForProfile(profile ProviderProfile) string
- func ProviderProfileRequiresCredential(profile ProviderProfile) bool
- func RecordOutcome(ctx context.Context, name string, err error, attrs ...Attr)
- func StaleFreshnessReports(rows []ProviderModelDescriptor, now time.Time) []string
- func ValidateDefaultCatalog() error
- func ValidateModeSettingsForPolicy(profiles []ProviderProfile, settings ModeSettings, policy RuntimePolicy) error
- func ValidateProfileForMode(profile ProviderProfile, mode Mode) error
- func ValidateRuntimePolicy(profiles []ProviderProfile, policy RuntimePolicy) error
- type AssistRequest
- type AssistResult
- type AssistService
- type AssistSetting
- type AssistSurfaceDecision
- type Attr
- type AudioData
- type AudioIdleObserver
- type AudioRecorder
- type AudioSegment
- type Capability
- type Command
- type CommandBus
- type CommandType
- type CommitObserver
- type Completion
- type CustomizationAction
- type DictationRun
- type DictationSegmenter
- func (s *DictationSegmenter) CollectStopSegments(fullPCM []byte) ([]AudioSegment, error)
- func (s *DictationSegmenter) DrainReadySegments() []AudioSegment
- func (s *DictationSegmenter) FeedPCM(pcm []byte) error
- func (s *DictationSegmenter) IdleAudio() (time.Duration, time.Time)
- func (s *DictationSegmenter) IdleSince() time.Time
- func (s *DictationSegmenter) SetMinIntermediateSegment(d time.Duration)
- type DictationService
- type DictationSetting
- type DictationStream
- type DictationStreamEvent
- type DictationStreamOptions
- type DictationStreamProvider
- type DictationStreamSink
- type DictationStreamSinkOptions
- type Engine
- type Event
- type EventType
- type ExecutionMode
- type Hooks
- type IdleObserver
- type IntelligenceKind
- type JobSubmitter
- type LiveCommitFlusher
- type LiveCommitPolicy
- type Metadata
- type Modality
- type Mode
- type ModeBehavior
- type ModeContract
- type ModeSetting
- type ModeSettings
- type ModelLifecycle
- type ModelVariant
- type Persistence
- type PooledPCMRecorder
- type ProviderDefault
- type ProviderFeature
- type ProviderFeatureSupport
- type ProviderKind
- type ProviderMatrixRow
- type ProviderModelDescriptor
- type ProviderProfile
- type ProviderSupportKind
- type QuickNoteStore
- type Readiness
- type ReadinessAction
- type ReadinessArtifact
- type ReadinessRequirement
- type ReadySegmentCollector
- type RecordingCancelOptions
- type RecordingController
- func (c *RecordingController) Cancel(opts RecordingCancelOptions) error
- func (c *RecordingController) IsCapturing() bool
- func (c *RecordingController) IsRecording() bool
- func (c *RecordingController) SetDictationStream(provider DictationStreamProvider, sink DictationStreamSink)
- func (c *RecordingController) SetFragmentSegments(enabled bool)
- func (c *RecordingController) SetIdleWatchInterval(d time.Duration)
- func (c *RecordingController) SetStreamSegments(enabled bool)
- func (c *RecordingController) Start(opts RecordingStartOptions) error
- func (c *RecordingController) Stop(opts RecordingStopOptions) error
- type RecordingObserver
- type RecordingStartOptions
- type RecordingStopOptions
- type Runtime
- func (r *Runtime) Close()
- func (r *Runtime) Commands() CommandBus
- func (r *Runtime) Events() <-chan Event
- func (r *Runtime) Publish(event Event) bool
- func (r *Runtime) SetState(snapshot Snapshot)
- func (r *Runtime) Start(ctx context.Context) error
- func (r *Runtime) State() Snapshot
- func (r *Runtime) Stop(ctx context.Context) error
- func (r *Runtime) UpdateState(update func(*Snapshot)) Snapshot
- type RuntimePolicy
- type SegmentCollector
- type SegmentCollectorFactory
- type ServerConnectionSetting
- type ServerConnectionTarget
- type Snapshot
- type Submission
- type Transcriber
- type Transcript
- type TranscriptInterceptor
- type TranscriptOutput
- type TranscriptSegmentKey
- type TranscriptSessionLedger
- type TranscriptTransformer
- type TranscriptionDraftObserver
- type TranscriptionJob
- type TranscriptionObserver
- type TranscriptionRunner
- type TranscriptionStore
- type TranscriptionWorker
- func (w *TranscriptionWorker) Close()
- func (w *TranscriptionWorker) HandleDictationStreamEvent(ctx context.Context, event DictationStreamEvent, ...) error
- func (w *TranscriptionWorker) Start(ctx context.Context)
- func (w *TranscriptionWorker) Submit(job TranscriptionJob) error
- func (w *TranscriptionWorker) Wait()
- type TranscriptionWorkerConfig
- type VoiceActivityDetector
- type VoiceAgentService
- type VoiceAgentSession
- type VoiceAgentSessionSummary
- type VoiceAgentSetting
- type VoiceAgentTurn
- type WordConfidence
Constants ¶
const ( AudioSampleRate = audio.SampleRate AudioChannels = audio.Channels AudioBitsPerSample = audio.BitsPerSample AudioBytesPerSample = audio.BytesPerSample )
const ( DefaultDictationPause = 1500 * time.Millisecond DefaultDictationMinSegment = 1200 * time.Millisecond DefaultDictationMinIntermediateSegment = 6 * time.Second DefaultDictationParagraphPause = 4 * time.Second DefaultDictationPadding = 480 * time.Millisecond DefaultDictationOverlap = 200 * time.Millisecond )
const ( // LiveCommitImmediate pastes each provider-final as soon as it arrives. LiveCommitImmediate = "immediate" // LiveCommitPhrase waits for one sentence (or a short pause) before paste. LiveCommitPhrase = "phrase" // LiveCommitPassage waits for about two sentences (or a longer pause) so // live field injection reads as prose rather than breath-sized fragments. LiveCommitPassage = "passage" )
const ( ModelAssemblyAIUniversal35ProRealtime = "universal-3-5-pro" ModelAssemblyAIU3RTPro = "u3-rt-pro" ModelAssemblyAIVoiceAgent = "assemblyai-voice-agent" ModelDeepgramFluxGeneralEN = "flux-general-en" ModelDeepgramFluxGeneralMulti = "flux-general-multi" // ModelDeepgramFluxTTSDefaultEN is Deepgram's default Flux TTS voice. Flux // TTS is English-only; Aura-2 remains the multilingual speak leg. ModelDeepgramFluxTTSDefaultEN = "flux-kit-en" ModelDeepgramNova3 = "nova-3" ModelGroqWhisperLargeV3 = "whisper-large-v3" ModelGroqWhisperLargeV3Turbo = "whisper-large-v3-turbo" ModelGemini35LiveTranslatePreview = "gemini-3.5-live-translate-preview" ModelGemini31FlashLivePreview = "gemini-3.1-flash-live-preview" ModelGemini25FlashNativeAudioPreview = "gemini-2.5-flash-native-audio-preview-12-2025" ModelOpenAIGPT4OTranscribe = "gpt-4o-transcribe" ModelOpenAIGPT4OMiniTranscribe = "gpt-4o-mini-transcribe" ModelOpenAIGPT4OTranscribeDiarize = "gpt-4o-transcribe-diarize" ModelOpenAIRealtime2 = "gpt-realtime-2" ModelOpenAIRealtime21 = "gpt-realtime-2.1" ModelOpenAIRealtime21Mini = "gpt-realtime-2.1-mini" )
const ( OutcomeEmptyFinalTranscript = "empty_final_transcript" OutcomePCMQueueDrop = "pcm_queue_drop" OutcomeAssistEmptySpeak = "assist_empty_speak" )
Outcome names for RecordOutcome. Keep these stable — backends and alerts key on the string, not on log message text.
const ( ProviderAuthNone = "none" ProviderAuthAPIKey = "api_key" ProviderAuthToken = "token" ProviderAuthHostDependencies = "host_dependencies" ProviderAuthOptionalAPIKey = "optional_api_key" ProviderTransportLocal = "local" ProviderTransportHTTP = "http" ProviderTransportHTTPS = "https" ProviderTransportWebSocket = "websocket" ProviderTransportPipeline = "pipeline" )
const ( CaptureChannelMicrophone = "mic" CaptureChannelSystem = "system" )
Capture channels name the audio source behind a recording. A host that records one session from several sources at once — meeting capture takes the microphone and the system loopback in parallel — labels each controller so the resulting transcripts stay attributable.
const DefaultLocalBuiltInLLMModel = "ggml-org/gemma-4-E2B-it-GGUF:Q8_0"
const DefaultMinPCMBytes = 3200
const DefaultProcessingMessage = "Recording stopped · Transcribing"
const EmptyFinalTranscriptMessage = "No speech recognized · check the configured language"
EmptyFinalTranscriptMessage is shown when a provider returns a successful final transcript containing no text.
This is a named outcome rather than a silent drop because it is the visible half of a real data-loss bug: a provider answers HTTP 200 with a zero-length transcript when the pinned language does not match the speech, so the user's words disappear with nothing to alert on. The message names the most likely cause, since that is the one the user can act on.
const ModelFreshnessSLA = 7 * 24 * time.Hour
ModelFreshnessSLA is the maximum age of LastVerifiedAt before a default or recommended model row is considered stale.
const ReadinessSchemaVersion = "provider-readiness.v1"
Variables ¶
var ( ErrMissingRunner = errors.New("speechkit: transcription worker requires a runner") ErrMissingTranscriber = errors.New("speechkit: transcription runner requires a transcriber") ErrWorkerClosed = errors.New("speechkit: transcription worker is closed") ErrWorkerQueueFull = errors.New("speechkit: transcription worker queue is full") )
ErrCommandHandlerUnavailable is returned by CommandBus.Dispatch when no command handler has been configured on the Runtime.
var ErrUnsupportedAudioFormat = errors.New("speechkit: unsupported audio format for this dictation stream provider")
ErrUnsupportedAudioFormat is returned by a DictationStreamProvider when the requested speaker.AudioFormat is one its realtime API cannot accept — e.g. AssemblyAI's v3 streaming API exposes no channel parameter and decodes the socket as a single channel, so stereo would transcribe as braided garbage rather than fail. Providers wrap it with %w.
Format support is per-provider, not a property of the format (Deepgram serves stereo natively), so this is deliberately not a protocol-level validation: the router's fallback loop treats it like any other start failure and tries the next candidate. Only when no provider can serve the format does it reach a caller, who should errors.Is it to report a format problem rather than invite a blind retry of the same doomed format.
Functions ¶
func CountTerminalSentences ¶ added in v0.61.20
CountTerminalSentences counts fragments that already look like finished sentences so live commit can wait for a short paragraph instead of a breath.
func DefaultProviderAuthRequirement ¶ added in v0.48.0
func DefaultProviderAuthRequirement(profile ProviderProfile) string
DefaultProviderAuthRequirement describes the credential class a host must satisfy before a provider profile can run. It is intentionally semantic: hosts map the value to their own env vars or secret stores.
func DefaultProviderTransport ¶ added in v0.48.0
func DefaultProviderTransport(profile ProviderProfile) string
DefaultProviderTransport exposes the dominant runtime transport class for a profile. Native realtime providers use websocket; cascaded voice providers use pipeline; batch/provider APIs use HTTPS/HTTP/local.
func JoinTranscriptFragments ¶ added in v0.61.20
JoinTranscriptFragments concatenates live transcript slices with a single separating space when the next slice would otherwise glue onto the previous word or sentence.
func LiveInjectFragment ¶ added in v0.61.20
func LiveInjectFragment(previousSession uint64, previousTail, next string, session uint64) (fragment, tail string, nextSession uint64)
LiveInjectFragment returns the text that should be pasted for this fragment so consecutive live injects keep a word gap without rewriting earlier text.
func LowConfidenceWords ¶ added in v0.46.0
func LowConfidenceWords(words []WordConfidence, threshold float64) (terms []string, minConfidence float64)
LowConfidenceWords returns the distinct word texts whose per-word confidence is below threshold, together with the minimum confidence observed across all words. A threshold <= 0 disables detection. Words without per-word data (providers that do not expose it) yield (nil, 0). The returned terms are the raw STT tokens, so callers can match them against the (possibly rewritten) display text without depending on character offsets.
func MissingFreshnessReports ¶ added in v0.60.0
func MissingFreshnessReports(rows []ProviderModelDescriptor) []string
MissingFreshnessReports lists default/recommended registry rows that still lack LastVerifiedAt.
func NeedsLiveInjectSpace ¶ added in v0.61.20
NeedsLiveInjectSpace reports whether two adjacent live fragments need a separating space when pasted one after the other.
func NormalizeProviderID ¶ added in v0.48.0
func NormalizeProviderProfileID ¶ added in v0.42.0
NormalizeProviderProfileID maps legacy profile IDs to their current canonical IDs while preserving unknown custom IDs.
func PCMDurationSecs ¶ added in v0.24.0
PCMDurationSecs returns the duration of 16kHz S16 mono PCM audio in seconds.
func ProviderCredentialTarget ¶ added in v0.48.0
func ProviderCredentialTarget(profile ProviderProfile) string
func ProviderIDForExecutionMode ¶ added in v0.48.0
func ProviderIDForExecutionMode(mode ExecutionMode) string
func ProviderIDForProfile ¶ added in v0.48.0
func ProviderIDForProfile(profile ProviderProfile) string
func ProviderProfileRequiresCredential ¶ added in v0.48.0
func ProviderProfileRequiresCredential(profile ProviderProfile) bool
func RecordOutcome ¶ added in v0.60.41
RecordOutcome attaches a named framework result to the active span.
This is the vendor-neutral error/outcome seam: callers keep using slog for operators, but user-visible failures also land on the trace so a configured OTLP backend can surface them. With no TracerProvider installed (the local-only default) this is a zero-cost no-op.
func StaleFreshnessReports ¶ added in v0.60.41
func StaleFreshnessReports(rows []ProviderModelDescriptor, now time.Time) []string
StaleFreshnessReports lists default/recommended rows whose LastVerifiedAt is missing, unparsable, or older than ModelFreshnessSLA relative to now.
func ValidateDefaultCatalog ¶ added in v0.24.0
func ValidateDefaultCatalog() error
ValidateDefaultCatalog verifies the framework invariant that every strict mode exposes all four provider groups and every visible profile satisfies its mode contract. v0.37 added ModeTTS as a model-selection axis with the same four-provider-group invariant (Local Built-in via Piper, Local Provider via Kokoro/openedai-speech, Cloud Provider via Hugging Face Parler, Direct Provider via OpenAI + Google).
func ValidateModeSettingsForPolicy ¶ added in v0.24.0
func ValidateModeSettingsForPolicy(profiles []ProviderProfile, settings ModeSettings, policy RuntimePolicy) error
ValidateModeSettingsForPolicy checks mode selections against a RuntimePolicy.
func ValidateProfileForMode ¶ added in v0.24.0
func ValidateProfileForMode(profile ProviderProfile, mode Mode) error
ValidateProfileForMode checks the stable v23 mode capability contract.
func ValidateRuntimePolicy ¶ added in v0.24.0
func ValidateRuntimePolicy(profiles []ProviderProfile, policy RuntimePolicy) error
ValidateRuntimePolicy checks that a policy references existing profiles and does not require a profile that violates its mode contract.
Types ¶
type AssistRequest ¶ added in v0.24.0
type AssistRequest struct {
Text string `json:"text"`
Locale string `json:"locale,omitempty"`
Selection string `json:"selection,omitempty"`
Context string `json:"context,omitempty"`
EditableTarget bool `json:"editableTarget,omitempty"`
ProviderProfileID string `json:"providerProfileId,omitempty"`
SessionKey string `json:"sessionKey,omitempty"`
SpeakerOptions speaker.Options `json:"speakerOptions,omitempty"`
Speakers *speaker.DiarizationResult `json:"speakers,omitempty"`
}
AssistRequest is the mode-scoped input for Assist integrations.
type AssistResult ¶ added in v0.24.0
type AssistResult struct {
Text string `json:"text"`
SpeakText string `json:"speakText,omitempty"`
Action string `json:"action,omitempty"`
Kind string `json:"kind,omitempty"`
Surface AssistSurfaceDecision `json:"surface"`
ShortcutID string `json:"shortcutId,omitempty"`
Locale string `json:"locale,omitempty"`
MessageID localization.MessageID `json:"messageId,omitempty"`
ReasonCode string `json:"reasonCode,omitempty"`
Audio *AudioData `json:"audio,omitempty"`
Format string `json:"format,omitempty"`
Speakers *speaker.DiarizationResult `json:"speakers,omitempty"`
}
AssistResult is the public one-shot output contract for Assist Mode.
type AssistService ¶ added in v0.24.0
type AssistService interface {
Process(context.Context, AssistRequest) (AssistResult, error)
}
AssistService is the mode-scoped SDK contract for one-shot utilities and work-product generation.
type AssistSetting ¶ added in v0.24.0
type AssistSetting struct {
ModeSetting
TTSEnabled bool `json:"ttsEnabled"`
UtilityRegistry string `json:"utilityRegistry,omitempty"`
}
type AssistSurfaceDecision ¶ added in v0.24.0
type AssistSurfaceDecision string
AssistSurfaceDecision describes where an Assist result should be presented.
const ( AssistSurfacePanel AssistSurfaceDecision = "panel" AssistSurfaceInsert AssistSurfaceDecision = "insert" AssistSurfaceReplace AssistSurfaceDecision = "replace" AssistSurfaceActionAck AssistSurfaceDecision = "action_ack" AssistSurfaceSilent AssistSurfaceDecision = "silent" )
type Attr ¶ added in v0.64.3
Attr is one key/value pair attached to a recorded outcome. It exists so RecordOutcome does not put OpenTelemetry in an embedder's own signatures: a host records outcomes with SpeechKit's own vocabulary, and whether a tracing backend is installed stays SpeechKit's business.
func Float64Attr ¶ added in v0.64.3
func StringAttr ¶ added in v0.64.3
StringAttr, Int64Attr, Float64Attr and BoolAttr build an Attr of the matching type. They read better at a call site than a struct literal and keep the value's type explicit.
type AudioData ¶ added in v0.40.1
type AudioData []byte
AudioData carries optional synthesized audio without making AssistResult non-comparable for existing SDK consumers.
func NewAudioData ¶ added in v0.40.1
type AudioIdleObserver ¶ added in v0.51.2
AudioIdleObserver is implemented by SegmentCollectors that can report silence in audio time: the cumulative duration of processed silent frames since the last detected speech, plus the wall-clock time of the most recently processed frame. When a collector satisfies this interface it is preferred over IdleObserver, because audio-anchored silence is immune to CPU-starvation stalls — when frame delivery stalls, the silence counter freezes instead of counting wall-clock seconds and auto-stopping mid-dictation.
type AudioRecorder ¶
AudioRecorder is the hardware abstraction for microphone capture.
type AudioSegment ¶ added in v0.24.0
AudioSegment is a transcribable utterance extracted from a dictation recording. PCM is raw 16kHz S16 mono audio.
func FallbackDictationSegments ¶
func FallbackDictationSegments(fullPCM []byte) []AudioSegment
FallbackDictationSegments wraps all of fullPCM in a single segment. Used when VAD-based segmentation is unavailable or produces no output.
type Capability ¶ added in v0.24.0
type Capability string
Capability is a mode capability declared by a provider profile.
const ( CapabilityTranscription Capability = "transcription" CapabilitySTT Capability = "stt" CapabilityAudioInput Capability = "audio_input" CapabilityLLM Capability = "llm" CapabilityTTS Capability = "tts" CapabilityRealtimeAudio Capability = "realtime_audio" CapabilityPipelineFallback Capability = "pipeline_fallback" CapabilityToolCalling Capability = "tool_calling" CapabilityDictionaryPrompt Capability = "dictionary_prompt" CapabilityDictionaryNativeHints Capability = "dictionary_native_hints" CapabilityWordsPrompt Capability = "words_prompt" CapabilityWordsNativeHints Capability = "words_native_hints" CapabilityPostSTTReplacements Capability = "post_stt_replacements" CapabilitySessionSummary Capability = "session_summary" CapabilityTranscript Capability = "transcript" CapabilityInterruptions Capability = "interruptions" CapabilitySessionResume Capability = "session_resume" CapabilityNativeContextPrompt Capability = "native_context_prompt" CapabilityNativeKeyterms Capability = "native_keyterms" CapabilityNativeDictationStream Capability = "native_dictation_stream" CapabilityLanguageHints Capability = "language_hints" CapabilitySpeakerStreaming Capability = "speaker_streaming" CapabilityPrivacyRedaction Capability = "privacy_redaction" CapabilityVoiceFocus Capability = "voice_focus" CapabilityMedicalDomain Capability = "medical_domain" CapabilityReasoningEffort Capability = "reasoning_effort" CapabilityTranslation Capability = "translation" CapabilityTranscriptionOnly Capability = "transcription_only" CapabilitySpeakerDiarization Capability = "speaker_diarization" CapabilitySpeakerIdentification Capability = "speaker_identification" CapabilitySpeakerAttribution Capability = "speaker_attribution" CapabilitySpeakerEnrollment Capability = "speaker_enrollment" )
func RequiredCapabilities ¶ added in v0.24.0
func RequiredCapabilities(mode Mode, nativeRealtime bool) []Capability
RequiredCapabilities returns the minimum capability set for a profile to satisfy a mode contract.
type Command ¶
type Command struct {
Type CommandType
Text string
NoteID int64
Target string
Metadata map[string]string
}
Command is a request dispatched through the CommandBus.
type CommandBus ¶
CommandBus delivers Command values to the registered handler.
type CommandType ¶
type CommandType string
CommandType identifies the action a Command requests.
const ( CommandShowDashboard CommandType = "dashboard.show" CommandStartDictation CommandType = "dictation.start" CommandStopDictation CommandType = "dictation.stop" CommandStartMode CommandType = "mode.start" CommandStopMode CommandType = "mode.stop" CommandSetActiveMode CommandType = "mode.set_active" CommandOpenQuickNote CommandType = "quicknote.open" CommandOpenQuickCapture CommandType = "quicknote.capture.open" CommandCloseQuickCapture CommandType = "quicknote.capture.close" CommandArmQuickNoteRecording CommandType = "quicknote.record.arm" CommandCopyLastTranscription CommandType = "transcription.copy_last" CommandInsertLastTranscription CommandType = "transcription.insert_last" CommandSummarizeSelection CommandType = "selection.summarize" )
type CommitObserver ¶
type CommitObserver interface {
OnCommit(completion Completion)
}
CommitObserver is notified after each successful TranscriptionRunner.Commit.
type Completion ¶
type Completion struct {
Transcript Transcript
QuickNoteCommitted bool
QuickNoteCreated bool
QuickNoteID int64
TranscriptionPersisted bool
AudioDurationMs int64
}
Completion describes the outcome of a TranscriptionRunner.Commit call.
type CustomizationAction ¶ added in v0.47.0
type CustomizationAction struct {
ReplacementID string `json:"replacement_id,omitempty"`
Kind string `json:"kind,omitempty"`
Intent string `json:"intent,omitempty"`
Text string `json:"text,omitempty"`
Template string `json:"template,omitempty"`
Payload map[string]any `json:"payload,omitempty"`
MatchedText string `json:"matched_text,omitempty"`
Count int `json:"count,omitempty"`
}
CustomizationAction describes a command/snippet/template action produced by Words and Replacements v2. Known command intents can be executed by hosts; unknown intents remain structured metadata for event/API consumers.
type DictationRun ¶ added in v0.24.0
type DictationRun struct {
ID string `json:"id,omitempty"`
Transcript Transcript `json:"transcript"`
StartedAt time.Time `json:"startedAt,omitempty"`
CompletedAt time.Time `json:"completedAt,omitempty"`
ProviderProfile string `json:"providerProfile,omitempty"`
DictionaryTerms []string `json:"dictionaryTerms,omitempty"`
AudioDurationMs int64 `json:"audioDurationMs,omitempty"`
ProcessingTimeMs int64 `json:"processingTimeMs,omitempty"`
Speakers *speaker.DiarizationResult `json:"speakers,omitempty"`
}
DictationRun is the public record produced by a completed Dictation request. Hosts may persist it directly or map it into their own history model.
type DictationSegmenter ¶
type DictationSegmenter struct {
// contains filtered or unexported fields
}
DictationSegmenter implements SegmentCollector using VAD-based pause detection to split continuous speech into discrete segments.
func NewDictationSegmenter ¶
func NewDictationSegmenter(detector VoiceActivityDetector, pauseThreshold time.Duration) *DictationSegmenter
func (*DictationSegmenter) CollectStopSegments ¶
func (s *DictationSegmenter) CollectStopSegments(fullPCM []byte) ([]AudioSegment, error)
func (*DictationSegmenter) DrainReadySegments ¶ added in v0.47.0
func (s *DictationSegmenter) DrainReadySegments() []AudioSegment
DrainReadySegments returns pause-bounded intermediate segments that were completed during FeedPCM calls. It leaves the active utterance in place so a later Stop() only flushes the remaining tail.
func (*DictationSegmenter) FeedPCM ¶
func (s *DictationSegmenter) FeedPCM(pcm []byte) error
func (*DictationSegmenter) IdleAudio ¶ added in v0.51.2
func (s *DictationSegmenter) IdleAudio() (time.Duration, time.Time)
IdleAudio reports silence measured in audio time rather than wall-clock time: the cumulative duration of processed silent frames since the last detected speech, plus the wall-clock time of the most recently processed PCM frame. While speech is in progress the silence duration is zero.
Audio-time anchoring makes silence-based auto-stop robust against CPU starvation: when frame delivery stalls, the silence counter freezes instead of counting real seconds against a stale timestamp.
Satisfies the AudioIdleObserver contract consumed by RecordingController; preferred over [IdleSince] when available.
func (*DictationSegmenter) IdleSince ¶ added in v0.35.21
func (s *DictationSegmenter) IdleSince() time.Time
IdleSince returns the wall-clock time at which the segmenter most recently transitioned out of speech (or, for a fresh session that has not yet seen speech, the construction time). Returns the zero value when speech is currently being captured — the poller treats zero as "user is actively speaking, silence timer should reset."
Satisfies the IdleObserver contract consumed by RecordingController to drive silence-based auto-stop.
func (*DictationSegmenter) SetMinIntermediateSegment ¶ added in v0.47.0
func (s *DictationSegmenter) SetMinIntermediateSegment(d time.Duration)
SetMinIntermediateSegment configures the minimum active utterance duration that can be emitted before Stop(). Shorter utterances stay merged across natural pauses so dictation does not over-fragment; <=0 emits on every pause-bounded segment.
type DictationService ¶ added in v0.24.0
type DictationService interface {
Start(context.Context) error
Stop(context.Context) (DictationRun, error)
}
DictationService is the mode-scoped SDK contract for text-only dictation.
type DictationSetting ¶ added in v0.24.0
type DictationSetting struct {
ModeSetting
DictionaryEnabled bool `json:"dictionaryEnabled"`
}
type DictationStream ¶ added in v0.48.0
type DictationStream interface {
SendPCM(ctx context.Context, pcm []byte) error
Finalize(ctx context.Context) error
Receive(ctx context.Context) (DictationStreamEvent, error)
Close() error
}
DictationStream is a provider-neutral realtime dictation session.
type DictationStreamEvent ¶ added in v0.48.0
type DictationStreamEvent struct {
Sequence int64
SessionID uint64
SegmentID uint64
ProviderItemID string
Text string
IsFinal bool
Language string
Provider string
Model string
Confidence float64
Words []WordConfidence
Speakers *speaker.DiarizationResult
}
DictationStreamEvent is the provider-neutral event emitted by DictationStream.Receive.
func (DictationStreamEvent) Transcript ¶ added in v0.48.0
func (e DictationStreamEvent) Transcript() Transcript
type DictationStreamOptions ¶ added in v0.48.0
type DictationStreamOptions struct {
SessionID uint64
ProviderProfileID string
Language string
Model string
InterimResults bool
EndpointingMs int
TurnDetection bool
Keyterms []string
PromptHint string
Diarization bool
}
DictationStreamOptions configures provider-native live transcription for Dictation and meeting transcription. It is intentionally separate from speaker.Options because plain dictation must not require diarization.
type DictationStreamProvider ¶ added in v0.48.0
type DictationStreamProvider interface {
StartDictationStream(ctx context.Context, opts DictationStreamOptions, format speaker.AudioFormat) (DictationStream, error)
}
DictationStreamProvider is implemented by STT providers that can consume raw PCM frames and emit draft/final transcript events without waiting for a completed WAV upload.
type DictationStreamSink ¶ added in v0.48.0
type DictationStreamSink interface {
HandleDictationStreamEvent(ctx context.Context, event DictationStreamEvent, opts DictationStreamSinkOptions) error
}
DictationStreamSink consumes provider-native live dictation events. Drafts are allowed to update UI state, but only final events may reach output or persistence.
type DictationStreamSinkOptions ¶ added in v0.48.0
type DictationStreamSinkOptions struct {
Target any
QuickNote bool
QuickNoteID int64
Language string
RecordingSessionID int64
// CaptureChannel names the capture source feeding this stream, and
// CaptureEpoch is the wall clock the session's timeline is measured from.
// Sinks stamp final events with the elapsed offset so parallel channels of
// one meeting can be interleaved. A zero epoch disables the timeline.
CaptureChannel string
CaptureEpoch time.Time
}
DictationStreamSinkOptions carries host metadata needed to commit final provider-stream events through the same path as batch transcription.
type Engine ¶
type Engine interface {
Start(context.Context) error
Stop(context.Context) error
Events() <-chan Event
Commands() CommandBus
State() Snapshot
}
Engine is the interface implemented by a full SpeechKit voice pipeline.
type Event ¶
type Event struct {
Type EventType
Time time.Time
Message string
Text string
Provider string
Mode string
SessionID string
QuickNote bool
Err error
Shortcut string
Metadata *Metadata
CustomizationActions []CustomizationAction
}
Event is a notification published to the event channel returned by Runtime.Events. Consumers should switch on Type and inspect the relevant fields.
type EventType ¶
type EventType string
EventType identifies the kind of event published to the event channel.
const ( EventStateChanged EventType = "state.changed" EventRecordingStarted EventType = "recording.started" EventProcessingStarted EventType = "processing.started" EventTranscriptionDraft EventType = "transcription.draft" EventTranscriptionReady EventType = "transcription.ready" EventTranscriptCommitted EventType = "transcription.committed" EventQuickNoteModeArmed EventType = "quicknote.mode_armed" EventQuickNoteUpdated EventType = "quicknote.updated" EventWarningRaised EventType = "warning.raised" EventErrorRaised EventType = "error.raised" EventShortcutMatched EventType = "shortcut.matched" EventWakeFired EventType = "wake.fired" EventSkillExecuted EventType = "skill.executed" EventCompanionSessionStarted EventType = "companion.session.started" EventCompanionSessionEnded EventType = "companion.session.ended" EventVoiceAgentTurnFinalized EventType = "voiceagent.turn.finalized" EventTTSStarted EventType = "tts.started" EventTTSFinished EventType = "tts.finished" EventCustomizationAction EventType = "customization.action" )
type ExecutionMode ¶ added in v0.24.0
type ExecutionMode string
ExecutionMode describes the technical runtime behind a provider profile.
const ( ExecutionModeLocal ExecutionMode = "local" ExecutionModeSelfHostedHTTP ExecutionMode = "self_hosted_http" ExecutionModeHFRouted ExecutionMode = "hf_routed" ExecutionModeOpenAI ExecutionMode = "openai_api" ExecutionModeGroq ExecutionMode = "groq_api" ExecutionModeGoogle ExecutionMode = "google_api" ExecutionModeDeepgram ExecutionMode = "deepgram_api" ExecutionModeAssemblyAI ExecutionMode = "assemblyai_api" ExecutionModeOllama ExecutionMode = "ollama_local" ExecutionModeOpenRouter ExecutionMode = "openrouter_api" )
type Hooks ¶
type Hooks struct {
Start func(context.Context) error
Stop func(context.Context) error
HandleCommand func(context.Context, Command) error
}
Hooks are the lifecycle callbacks wired into a Runtime. Nil hooks are silently skipped.
type IdleObserver ¶ added in v0.35.21
IdleObserver is implemented by SegmentCollectors that want to drive silence-based auto-stop. Returning the zero value tells the watcher "user is actively speaking; reset the timer." Returning a non-zero time tells the watcher "user has been silent since T."
type IntelligenceKind ¶ added in v0.24.0
type IntelligenceKind string
IntelligenceKind names the mode-specific intelligence contract.
const ( IntelligenceUser IntelligenceKind = "user" IntelligenceUtility IntelligenceKind = "utility" IntelligenceBrainstorming IntelligenceKind = "brainstorming" // IntelligenceVoiceOutput is the contract for the TTS mode: render // generated text to audio. No user intelligence, no utility tools, // no brainstorming — strictly text in, audio out. IntelligenceVoiceOutput IntelligenceKind = "voice_output" )
type JobSubmitter ¶
type JobSubmitter interface {
Submit(TranscriptionJob) error
}
JobSubmitter accepts a TranscriptionJob for async processing.
type LiveCommitFlusher ¶ added in v0.61.20
LiveCommitFlusher drains a grouped live-commit buffer. Recording stop calls this so a trailing sentence is not left behind when the stream ends.
type LiveCommitPolicy ¶ added in v0.61.20
LiveCommitPolicy groups provider-native finals before they reach a sink. Overlay drafts still pass through immediately.
func NormalizeLiveCommitPolicy ¶ added in v0.61.20
func NormalizeLiveCommitPolicy(mode string) LiveCommitPolicy
NormalizeLiveCommitPolicy maps a host mode onto hold/sentence defaults. Empty or unknown modes disable grouping so existing hosts stay immediate.
type Metadata ¶ added in v0.40.1
Metadata carries optional event key/value data without making Event non-comparable for existing SDK consumers.
func NewMetadata ¶ added in v0.40.1
type Modality ¶ added in v0.62.1
type Modality string
Modality classifies what a catalog entry does, independent of the three user-facing modes. Every profile a user can select maps onto a Mode as well; support entries a host needs but a user never picks — embeddings, rerankers, utility models — only have a Modality.
func ModalityForMode ¶ added in v0.62.1
ModalityForMode returns the modality a user-facing mode runs as, or "" for a mode that has none.
type Mode ¶ added in v0.24.0
type Mode string
Mode identifies one of SpeechKit's strict product modes.
const ( ModeNone Mode = "none" ModeDictation Mode = "dictation" ModeAssist Mode = "assist" ModeVoiceAgent Mode = "voice_agent" // ModeTTS exposes Text-to-Speech as a first-class model-selection axis // alongside the three product modes. The host activates an Assist or // Voice-Agent session and the TTS profile selected here drives which // provider speaks the response. v0.37 introduced this alongside the // Voice-Companion hands-free flow — Thalia + Companion Live need a // stable place to pin a TTS voice across deployments. ModeTTS Mode = "tts" )
func ModeForModality ¶ added in v0.62.1
ModeForModality returns the user-facing mode a modality is selectable in, or ModeNone for support modalities a user never picks directly.
func NormalizeMode ¶ added in v0.24.0
type ModeBehavior ¶ added in v0.24.0
type ModeBehavior string
ModeBehavior describes how much mode-specific intelligence a host enables.
const ( // ModeBehaviorClean keeps a mode on its core contract, such as strict STT // for Dictation or deterministic utility handling for Assist. ModeBehaviorClean ModeBehavior = "clean" // ModeBehaviorIntelligence allows optional intelligence layers such as // LLM utility handling, TTS, summaries, or realtime tool use. ModeBehaviorIntelligence ModeBehavior = "intelligence" )
type ModeContract ¶ added in v0.24.0
type ModeContract struct {
Mode Mode `json:"mode"`
Intelligence IntelligenceKind `json:"intelligence"`
Input string `json:"input"`
Output string `json:"output"`
Allowed []Capability `json:"allowed"`
Forbidden []Capability `json:"forbidden"`
}
ModeContract documents what a mode may and may not do. Hosts can use this to validate custom adapters before exposing them to users.
func DefaultModeContracts ¶ added in v0.24.0
func DefaultModeContracts() []ModeContract
type ModeSetting ¶ added in v0.24.0
type ModeSetting struct {
Enabled bool `json:"enabled"`
Hotkey string `json:"hotkey,omitempty"`
HotkeyBehavior string `json:"hotkeyBehavior,omitempty"`
PrimaryProfileID string `json:"primaryProfileId,omitempty"`
FallbackProfileID string `json:"fallbackProfileId,omitempty"`
// ModeSource is "local" (default) or "server". When "server", this mode
// runs against the speechkit-server pointed to by ServerConnection
// instead of the in-process Framework kernel. Empty/missing is treated
// as "local" for backwards compatibility with pre-0.26 hosts.
ModeSource string `json:"modeSource,omitempty"`
}
ModeSetting is the public per-mode configuration shape used by the SDK and the versioned HTTP control plane.
type ModeSettings ¶ added in v0.24.0
type ModeSettings struct {
Dictation DictationSetting `json:"dictation"`
Assist AssistSetting `json:"assist"`
VoiceAgent VoiceAgentSetting `json:"voiceAgent"`
ServerConnection ServerConnectionSetting `json:"serverConnection"`
}
type ModelLifecycle ¶ added in v0.47.0
type ModelLifecycle string
const ( ModelLifecycleGA ModelLifecycle = "ga" ModelLifecyclePreview ModelLifecycle = "preview" ModelLifecycleLegacy ModelLifecycle = "legacy" ModelLifecycleDeprecated ModelLifecycle = "deprecated" )
type ModelVariant ¶ added in v0.24.0
type ModelVariant struct {
ID string `json:"id"`
Name string `json:"name"`
ModelID string `json:"modelId"`
Description string `json:"description,omitempty"`
Recommended bool `json:"recommended,omitempty"`
}
ModelVariant is a concrete model choice inside a provider profile group.
type Persistence ¶
type Persistence interface {
QuickNoteStore
TranscriptionStore
}
Persistence combines QuickNoteStore and TranscriptionStore.
type PooledPCMRecorder ¶ added in v0.51.2
type PooledPCMRecorder interface {
SetPooledPCMHandler(func(buf []byte, release func()))
}
PooledPCMRecorder is optionally implemented by AudioRecorders whose backend leases per-frame buffers from a pool instead of allocating a fresh copy per frame (~33 allocations/sec during capture). When the recorder satisfies this interface the controller installs the pool-aware handler and releases each buffer as soon as the frame has been fed to the collector/stream — the controller never retains a frame. Structurally matches internal/audio's SetPooledPCMHandler.
type ProviderDefault ¶ added in v0.48.0
type ProviderDefault struct {
Provider string `json:"provider"`
DisplayName string `json:"displayName"`
Mode Mode `json:"mode"`
ProfileID string `json:"profileId"`
ModelID string `json:"modelId,omitempty"`
ProviderKind ProviderKind `json:"providerKind"`
ExecutionMode ExecutionMode `json:"executionMode,omitempty"`
Support ProviderSupportKind `json:"support"`
Capabilities []Capability `json:"capabilities,omitempty"`
NativeOptions []string `json:"nativeOptions,omitempty"`
AuthRequirement string `json:"authRequirement,omitempty"`
CredentialRequired bool `json:"credentialRequired"`
CredentialTarget string `json:"credentialTarget,omitempty"`
Transport string `json:"transport,omitempty"`
EvidenceURL string `json:"evidenceUrl,omitempty"`
Default bool `json:"default,omitempty"`
Recommended bool `json:"recommended,omitempty"`
Experimental bool `json:"experimental,omitempty"`
Variants []ModelVariant `json:"variants,omitempty"`
}
func DefaultProviderDefaults ¶ added in v0.48.0
func DefaultProviderDefaults() []ProviderDefault
func FindProviderDefault ¶ added in v0.48.0
func FindProviderDefault(provider string, mode Mode) (ProviderDefault, bool)
func ProviderDefaultsFor ¶ added in v0.48.0
func ProviderDefaultsFor(provider string) []ProviderDefault
type ProviderFeature ¶ added in v0.48.0
type ProviderFeature string
const ( ProviderFeatureDictation ProviderFeature = "dictation" ProviderFeatureDictationStreaming ProviderFeature = "dictation_streaming" ProviderFeatureLongTranscription ProviderFeature = "long_transcription" ProviderFeatureSpeakerDiarization ProviderFeature = "speaker_diarization" ProviderFeatureSpeakerIdentification ProviderFeature = "speaker_identification" ProviderFeatureAssist ProviderFeature = "assist" ProviderFeatureRealtimeVoice ProviderFeature = "realtime_voice" ProviderFeatureTTS ProviderFeature = "tts" )
type ProviderFeatureSupport ¶ added in v0.48.0
type ProviderFeatureSupport struct {
Feature ProviderFeature `json:"feature"`
Support ProviderSupportKind `json:"support"`
Mode Mode `json:"mode,omitempty"`
ProfileID string `json:"profileId,omitempty"`
ModelID string `json:"modelId,omitempty"`
NativeOptions []string `json:"nativeOptions,omitempty"`
EvidenceURL string `json:"evidenceUrl,omitempty"`
}
type ProviderKind ¶ added in v0.24.0
type ProviderKind string
ProviderKind is the product-facing provider group shown for every mode.
const ( ProviderKindLocalBuiltIn ProviderKind = "local_built_in" ProviderKindLocalProvider ProviderKind = "local_provider" ProviderKindCloudProvider ProviderKind = "cloud_provider" ProviderKindDirectProvider ProviderKind = "direct_provider" )
func ProviderKindsForMode ¶ added in v0.24.0
func ProviderKindsForMode(mode Mode) []ProviderKind
type ProviderMatrixRow ¶ added in v0.48.0
type ProviderMatrixRow struct {
Provider string `json:"provider"`
DisplayName string `json:"displayName"`
Profiles []ProviderDefault `json:"profiles"`
Features []ProviderFeatureSupport `json:"features"`
}
func DefaultProviderMatrix ¶ added in v0.48.0
func DefaultProviderMatrix() []ProviderMatrixRow
func FindProviderMatrixRow ¶ added in v0.48.0
func FindProviderMatrixRow(provider string) (ProviderMatrixRow, bool)
func (ProviderMatrixRow) Feature ¶ added in v0.48.0
func (r ProviderMatrixRow) Feature(feature ProviderFeature) (ProviderFeatureSupport, bool)
type ProviderModelDescriptor ¶ added in v0.47.0
type ProviderModelDescriptor struct {
Provider string `json:"provider"`
ModelID string `json:"modelId"`
ProfileID string `json:"profileId,omitempty"`
Mode Mode `json:"mode"`
Name string `json:"name"`
Lifecycle ModelLifecycle `json:"lifecycle"`
Default bool `json:"default,omitempty"`
Recommended bool `json:"recommended,omitempty"`
SourceURL string `json:"sourceUrl"`
// Freshness metadata (kombify-SpeechKit-glnc). Dates are calendar days
// (YYYY-MM-DD) from vendor documentation. LastVerifiedAt is the day the
// row was last checked against those docs; TestDefaultModelRegistryFreshnessSLA
// fails when a default/recommended row is missing it or older than
// ModelFreshnessSLA.
ReleasedAt string `json:"releasedAt,omitempty"`
DeprecatedAt string `json:"deprecatedAt,omitempty"`
SunsetAt string `json:"sunsetAt,omitempty"`
LastVerifiedAt string `json:"lastVerifiedAt,omitempty"`
MultilanguageCapable bool `json:"multilanguageCapable,omitempty"`
}
ProviderModelDescriptor is the public source-of-truth row for model IDs that SpeechKit treats as framework defaults or first-class live-provider choices.
func DefaultModelRegistry ¶ added in v0.47.0
func DefaultModelRegistry() []ProviderModelDescriptor
func FindModelDescriptor ¶ added in v0.47.0
func FindModelDescriptor(provider, modelID string) (ProviderModelDescriptor, bool)
type ProviderProfile ¶ added in v0.24.0
type ProviderProfile struct {
ID string `json:"id"`
Mode Mode `json:"mode"`
// Modality is what the entry does. ProviderProfileWithDefaults derives it
// from Mode when unset, and derives Mode from it for support entries that
// carry no mode.
Modality Modality `json:"modality,omitempty"`
Name string `json:"name"`
ProviderKind ProviderKind `json:"providerKind"`
ExecutionMode ExecutionMode `json:"executionMode,omitempty"`
Provider string `json:"provider,omitempty"`
ModelID string `json:"modelId,omitempty"`
Lifecycle ModelLifecycle `json:"lifecycle,omitempty"`
Source string `json:"source,omitempty"`
Description string `json:"description,omitempty"`
License string `json:"license,omitempty"`
Capabilities []Capability `json:"capabilities,omitempty"`
SupportedLocales []string `json:"supportedLocales,omitempty"`
NativeOptions []string `json:"nativeOptions,omitempty"`
AuthRequirement string `json:"authRequirement,omitempty"`
Transport string `json:"transport,omitempty"`
EvidenceURL string `json:"evidenceUrl,omitempty"`
AdapterKind string `json:"adapterKind,omitempty"`
Variants []ModelVariant `json:"variants,omitempty"`
AllowInference bool `json:"inferenceAllowed,omitempty"`
Default bool `json:"default,omitempty"`
Recommended bool `json:"recommended,omitempty"`
Experimental bool `json:"experimental,omitempty"`
}
ProviderProfile is the public catalog entry host applications can present or activate. ProviderKind is the stable user-facing grouping; ExecutionMode is the technical adapter underneath it.
func DefaultProviderProfiles ¶ added in v0.24.0
func DefaultProviderProfiles() []ProviderProfile
DefaultProviderProfiles returns the built-in framework provider catalog for the three strict SpeechKit modes. The Windows desktop host adapts this public catalog into its internal runtime model; the catalog itself belongs to the reusable framework layer.
func FilterProviderProfiles ¶ added in v0.24.0
func FilterProviderProfiles(profiles []ProviderProfile, policy RuntimePolicy) []ProviderProfile
FilterProviderProfiles returns the profiles visible under policy.
func ProfilesForMode ¶ added in v0.24.0
func ProfilesForMode(mode Mode) []ProviderProfile
func ProviderProfileWithDefaults ¶ added in v0.48.0
func ProviderProfileWithDefaults(profile ProviderProfile) ProviderProfile
ProviderProfileWithDefaults returns a copy with framework-standard provider metadata filled in. Explicit profile metadata wins; missing provider, credential, and transport fields are derived from the canonical provider id, execution mode, and mode capabilities.
func (ProviderProfile) HasCapability ¶ added in v0.24.0
func (p ProviderProfile) HasCapability(capability Capability) bool
type ProviderSupportKind ¶ added in v0.48.0
type ProviderSupportKind string
const ( ProviderSupportUnsupported ProviderSupportKind = "unsupported" ProviderSupportPlanned ProviderSupportKind = "planned" ProviderSupportCascaded ProviderSupportKind = "cascaded" ProviderSupportRouted ProviderSupportKind = "routed" ProviderSupportNative ProviderSupportKind = "native" )
type QuickNoteStore ¶
type QuickNoteStore interface {
SaveQuickNote(ctx context.Context, text, language, provider string, durationMs, latencyMs int64, audioData []byte) (int64, error)
GetQuickNoteText(ctx context.Context, id int64) (string, error)
UpdateQuickNote(ctx context.Context, id int64, text string) error
UpdateQuickNoteCapture(ctx context.Context, id int64, text, provider string, durationMs, latencyMs int64, audioData []byte) error
}
QuickNoteStore persists and retrieves Quick Note records.
type Readiness ¶ added in v0.24.0
type Readiness struct {
SchemaVersion string `json:"schemaVersion,omitempty"`
ProfileID string `json:"profileId"`
Mode Mode `json:"mode"`
ProviderKind ProviderKind `json:"providerKind"`
ExecutionMode ExecutionMode `json:"executionMode,omitempty"`
ModelID string `json:"modelId,omitempty"`
Source string `json:"source,omitempty"`
Active bool `json:"active"`
Default bool `json:"default"`
Configured bool `json:"configured"`
CredentialsReady bool `json:"credentialsReady"`
RuntimeReady bool `json:"runtimeReady"`
CapabilityReady bool `json:"capabilityReady"`
Ready bool `json:"ready"`
Missing []string `json:"missing,omitempty"`
Requirements []ReadinessRequirement `json:"requirements,omitempty"`
Actions []ReadinessAction `json:"actions,omitempty"`
Artifacts []ReadinessArtifact `json:"artifacts,omitempty"`
}
Readiness describes whether a provider profile can be used right now.
type ReadinessAction ¶ added in v0.24.0
type ReadinessAction struct {
ID string `json:"id"`
Label string `json:"label"`
Kind string `json:"kind"`
Target string `json:"target,omitempty"`
}
ReadinessAction describes the next setup command a host can expose when a requirement is not ready.
type ReadinessArtifact ¶ added in v0.24.0
type ReadinessArtifact struct {
ID string `json:"id"`
Name string `json:"name"`
Kind string `json:"kind"`
SizeLabel string `json:"sizeLabel,omitempty"`
SizeBytes int64 `json:"sizeBytes,omitempty"`
Available bool `json:"available"`
Selected bool `json:"selected"`
RuntimeReady bool `json:"runtimeReady,omitempty"`
RuntimeProblem string `json:"runtimeProblem,omitempty"`
Recommended bool `json:"recommended,omitempty"`
}
ReadinessArtifact describes downloadable or pullable model artifacts tied to a provider profile. Local Built-in profiles use this to expose concrete model choices through the same readiness API as credentials and runtime checks.
type ReadinessRequirement ¶ added in v0.24.0
type ReadinessRequirement struct {
ID string `json:"id"`
Label string `json:"label"`
Category string `json:"category"`
Required bool `json:"required"`
Ready bool `json:"ready"`
Missing string `json:"missing,omitempty"`
}
ReadinessRequirement is a machine-readable setup check for a provider profile. Hosts can render these checks directly instead of hard-coding provider-specific setup rules.
type ReadySegmentCollector ¶ added in v0.47.0
type ReadySegmentCollector interface {
SegmentCollector
DrainReadySegments() []AudioSegment
}
ReadySegmentCollector is implemented by collectors that can hand completed pause-bounded segments to the transcription queue before recording stops.
type RecordingCancelOptions ¶ added in v0.46.0
type RecordingCancelOptions struct {
Label string
}
RecordingCancelOptions controls cancellation of an active recording without submitting captured audio. It is for host-level interruptions such as a mode switch, where the old buffer must be discarded rather than transcribed.
type RecordingController ¶
type RecordingController struct {
// contains filtered or unexported fields
}
RecordingController manages the start/stop lifecycle of a single recording session and hands audio segments to the submission queue.
func NewRecordingController ¶
func NewRecordingController(recorder AudioRecorder, submitter JobSubmitter, observer RecordingObserver, segmenterFactory SegmentCollectorFactory) *RecordingController
func (*RecordingController) Cancel ¶ added in v0.46.0
func (c *RecordingController) Cancel(opts RecordingCancelOptions) error
Cancel stops the active recorder and discards the captured audio. Hosts use this when the user switches modes mid-capture; submitting the old buffer would deliver stale speech through the newly selected mode.
func (*RecordingController) IsCapturing ¶ added in v0.51.2
func (c *RecordingController) IsCapturing() bool
IsCapturing reports whether the microphone is physically open. Unlike [IsRecording] it excludes the stop/drain window, so hosts can distinguish "user is still dictating" (suppress post-capture UI states) from "capture ended, transcription in flight" (terminal states must display).
func (*RecordingController) IsRecording ¶
func (c *RecordingController) IsRecording() bool
func (*RecordingController) SetDictationStream ¶ added in v0.48.0
func (c *RecordingController) SetDictationStream(provider DictationStreamProvider, sink DictationStreamSink)
SetDictationStream configures the optional provider-native live dictation path. Hosts can leave this unset to keep the public full-capture default.
func (*RecordingController) SetFragmentSegments ¶ added in v0.46.0
func (c *RecordingController) SetFragmentSegments(enabled bool)
SetFragmentSegments controls whether Stop() submits the VAD-derived segments as the STT source (true) or the full captured audio as a single submission (false, default). The segmenter still drives silence-based auto-stop either way; this only changes what audio is sent to transcription.
func (*RecordingController) SetIdleWatchInterval ¶ added in v0.35.21
func (c *RecordingController) SetIdleWatchInterval(d time.Duration)
SetIdleWatchInterval overrides the polling interval used by the silence-based auto-stop watcher. Tests use this to keep the unit tests fast (e.g. 5ms polling). Production should never touch this.
func (*RecordingController) SetStreamSegments ¶ added in v0.47.0
func (c *RecordingController) SetStreamSegments(enabled bool)
SetStreamSegments controls whether completed pause-bounded segments are submitted during recording instead of waiting until Stop(). The default is false. This is intended for live-ish dictation surfaces; hosts that need the strongest protection against VAD excision should keep the full-capture default.
func (*RecordingController) Start ¶
func (c *RecordingController) Start(opts RecordingStartOptions) error
func (*RecordingController) Stop ¶
func (c *RecordingController) Stop(opts RecordingStopOptions) error
type RecordingObserver ¶
type RecordingStartOptions ¶
type RecordingStartOptions struct {
// Context scopes provider-native streaming sessions. When nil,
// context.Background is used.
Context context.Context
Label string
Target any
Language string
QuickNote bool
QuickNoteID int64
// RecordingSessionID links final transcript commits to a persisted
// long-running dictation or meeting session owned by the host.
RecordingSessionID int64
// CaptureChannel names the audio source this controller records, so a host
// running several controllers over one recording session — meeting capture
// records the microphone and the system loopback at the same time — can tell
// the resulting transcripts apart. See CaptureChannel*.
CaptureChannel string
// CaptureEpoch is the wall clock that transcript timestamps are measured
// from. Hosts recording one session across several controllers pass the same
// epoch to all of them so the transcripts interleave on a single timeline.
// Zero means "start of this recording".
CaptureEpoch time.Time
// StreamSegments enables live-ish dictation for this recording session.
// Completed pause-bounded segments are queued before Stop(); Stop() then
// flushes only pending/remaining tail segments. Leave false for Assist and
// Voice Agent fallback capture, where a single full turn is the safer unit.
StreamSegments bool
// ProviderStream enables provider-native realtime dictation when the host
// configured a DictationStreamProvider and DictationStreamSink. If stream
// startup fails, the controller keeps using StreamSegments/full-capture
// fallback behavior.
ProviderStream bool
// DictationStreamOptions are passed to the native provider stream. SessionID
// and Language are filled from the active recording when left empty.
DictationStreamOptions DictationStreamOptions
// LiveCommitMode groups provider-finals before field injection.
// Empty keeps immediate commit (tests and hosts that do not opt in).
// Desktop dictation defaults to LiveCommitPassage.
LiveCommitMode string
// IdleTimeout, when greater than zero AND the underlying collector
// implements [IdleObserver], arms a watcher that calls
// OnIdleTimeoutCallback once the user has been silent for this long.
// Zero (default) disables the watcher — typical for hold-to-talk
// hotkey sessions that already terminate on KeyUp.
IdleTimeout time.Duration
// OnIdleTimeoutCallback fires once if IdleTimeout elapses without
// observed speech. Wired by the host to dispatch a Stop command so
// the dictate session ends after a silence window. The watcher
// guarantees at-most-one invocation per Start() call.
OnIdleTimeoutCallback func()
}
type RecordingStopOptions ¶
type Runtime ¶
type Runtime struct {
// contains filtered or unexported fields
}
Runtime manages shared observable state and event delivery for a SpeechKit session. Create one with NewRuntime and wire it into the host application via Runtime.Events and Runtime.Commands.
func NewRuntime ¶
func (*Runtime) Commands ¶
func (r *Runtime) Commands() CommandBus
func (*Runtime) Events ¶
Events returns the runtime's event channel. The channel is buffered with 64 slots; Runtime.Publish never blocks and drops events once the buffer is full, so consumers must drain the channel promptly to avoid losing events.
func (*Runtime) Publish ¶
Publish delivers event to the channel returned by Runtime.Events. It never blocks: the event channel is buffered with 64 slots, and when the buffer is full (or the runtime is closed, or event.Type is empty) the event is silently dropped. The bool return reports whether the event was actually delivered to the buffer. Slow consumers must drain the events channel promptly, or events will be lost.
func (*Runtime) UpdateState ¶
type RuntimePolicy ¶ added in v0.24.0
type RuntimePolicy struct {
EnabledModes []Mode `json:"enabledModes,omitempty"`
AllowedProfiles []string `json:"allowedProfiles,omitempty"`
FixedProfiles map[Mode]string `json:"fixedProfiles,omitempty"`
AllowFallbacks bool `json:"allowFallbacks,omitempty"`
ModeBehaviors map[Mode]ModeBehavior `json:"modeBehaviors,omitempty"`
}
RuntimePolicy constrains which parts of the SpeechKit framework a host application exposes. Empty EnabledModes or AllowedProfiles mean "all".
type SegmentCollector ¶
type SegmentCollector interface {
FeedPCM([]byte) error
CollectStopSegments(fullPCM []byte) ([]AudioSegment, error)
}
SegmentCollector accumulates real-time PCM frames and splits them into dictation segments when recording stops.
type SegmentCollectorFactory ¶
type SegmentCollectorFactory func() SegmentCollector
type ServerConnectionSetting ¶ added in v0.26.0
type ServerConnectionSetting struct {
Enabled bool `json:"enabled"`
ActiveTargetID string `json:"activeTargetId,omitempty"`
URL string `json:"url"`
BearerTokenEnv string `json:"bearerTokenEnv,omitempty"`
AuthMode string `json:"authMode,omitempty"`
BetaInstallIDEnv string `json:"betaInstallIdEnv,omitempty"`
BetaInstallSecretEnv string `json:"betaInstallSecretEnv,omitempty"`
BearerTokenSet bool `json:"bearerTokenSet"`
FallbackToLocal bool `json:"fallbackToLocal"`
RequestTimeoutSec int `json:"requestTimeoutSec"`
Targets []ServerConnectionTarget `json:"targets,omitempty"`
}
ServerConnectionSetting exposes the [server_connection] config section to the control-plane API + frontend. The bearer token is never sent across this boundary — only the env var name + connection metadata.
type ServerConnectionTarget ¶ added in v0.31.0
type ServerConnectionTarget struct {
ID string `json:"id"`
Label string `json:"label"`
URL string `json:"url"`
AuthMode string `json:"authMode"`
BearerTokenEnv string `json:"bearerTokenEnv,omitempty"`
BetaInstallIDEnv string `json:"betaInstallIdEnv,omitempty"`
BetaInstallSecretEnv string `json:"betaInstallSecretEnv,omitempty"`
BearerTokenSet bool `json:"bearerTokenSet"`
FallbackToLocal bool `json:"fallbackToLocal"`
RequestTimeoutSec int `json:"requestTimeoutSec"`
}
type Snapshot ¶
type Snapshot struct {
Status string
Text string
Level float64
Hotkey string
ActiveMode string
Providers []string
ActiveProfiles map[string]string
Transcriptions int
QuickNoteMode bool
QuickCaptureMode bool
LastTranscriptionText string
}
Snapshot is a point-in-time copy of the Runtime's observable state. All slice and map fields are safe to read without holding any lock.
type Submission ¶
type Submission struct {
PCM []byte
WAV []byte
DurationSecs float64
Language string
Prefix string
QuickNote bool
QuickNoteID int64
SessionID uint64
SegmentID uint64
// RecordingSessionID is copied into the final Transcript and Completion so
// host observers can attach committed text to a long-running session after
// persistence/output succeeds.
RecordingSessionID int64
// CaptureChannel, CapturedStartMs and CapturedEndMs carry the capture
// source and wall-clock placement through to the Transcript. See the
// matching Transcript fields.
CaptureChannel string
CapturedStartMs int64
CapturedEndMs int64
// ProviderItemID carries provider-native turn/item IDs for realtime
// streams. Segment-batch jobs leave it empty and rely on SessionID+SegmentID.
ProviderItemID string
SegmentFinal bool
// QueuedAt is set by TranscriptionWorker.Submit when the segment enters
// the worker queue. Hosts can prefill it when replaying externally queued
// work, but ordinary callers should leave it zero.
QueuedAt time.Time
}
Submission carries a single audio segment and its metadata into the transcription pipeline.
type Transcriber ¶
type Transcriber interface {
Transcribe(ctx context.Context, audio []byte, durationSecs float64, language string) (Transcript, error)
}
Transcriber converts raw WAV audio into a Transcript.
type Transcript ¶
type Transcript struct {
Text string
Language string
Duration time.Duration
Provider string
Model string
Confidence float64
// Session metadata is set for progressive dictation/meeting pipelines.
// Draft transcripts may be replaced by later revisions; only final
// segment IDs are committed by the worker ledger.
SessionID uint64
SegmentID uint64
ProviderItemID string
SegmentFinal bool
// RecordingSessionID links this transcript to a persisted long-running
// dictation or meeting session in the host store.
RecordingSessionID int64
// CaptureChannel names the capture source this transcript came from (see
// CaptureChannel*). Sessions that record more than one source at once —
// meeting capture runs the microphone and the system loopback in parallel —
// use it to keep the two apart. Empty for single-source captures.
CaptureChannel string
// CapturedStartMs and CapturedEndMs place this transcript on the capture
// session's wall-clock timeline, relative to RecordingStartOptions.
// CaptureEpoch. Both are zero when the host did not request a timeline.
CapturedStartMs int64
CapturedEndMs int64
// Words carries per-word acoustic confidence when available (Deepgram,
// AssemblyAI). Used to surface likely-misrecognized terms; nil otherwise.
Words []WordConfidence
Speakers *speaker.DiarizationResult
CustomizationActions []CustomizationAction `json:"customization_actions,omitempty"`
}
Transcript holds the result of a single transcription call.
type TranscriptInterceptor ¶
type TranscriptInterceptor interface {
Intercept(ctx context.Context, transcript Transcript, target any) (bool, error)
}
TranscriptInterceptor can handle a transcript before it reaches the normal output path. Return (true, nil) to signal that the transcript was consumed.
type TranscriptOutput ¶
type TranscriptOutput interface {
Deliver(ctx context.Context, transcript Transcript, target any) error
}
TranscriptOutput delivers a completed Transcript to the host application (e.g. clipboard injection or text-field paste).
type TranscriptSegmentKey ¶ added in v0.48.0
type TranscriptSegmentKey struct {
CaptureChannel string
SessionID uint64
SegmentID uint64
ProviderItemID string
}
TranscriptSegmentKey uniquely identifies a final transcript unit inside a progressive dictation or meeting session.
CaptureChannel is part of the identity because a meeting records several sources at once and each source numbers its own segments from one. Without it the two channels collide on their first segment and one of them is discarded as a duplicate of the other.
func (TranscriptSegmentKey) IsZero ¶ added in v0.48.0
func (k TranscriptSegmentKey) IsZero() bool
type TranscriptSessionLedger ¶ added in v0.48.0
type TranscriptSessionLedger struct {
// contains filtered or unexported fields
}
TranscriptSessionLedger suppresses duplicate final commits for progressive transcription. It tracks both in-flight and completed segments so repeated Stop/Finalize/provider events cannot paste the same text twice.
func NewTranscriptSessionLedger ¶ added in v0.48.0
func NewTranscriptSessionLedger() *TranscriptSessionLedger
func (*TranscriptSessionLedger) Begin ¶ added in v0.48.0
func (l *TranscriptSessionLedger) Begin(key TranscriptSegmentKey) bool
func (*TranscriptSessionLedger) Commit ¶ added in v0.48.0
func (l *TranscriptSessionLedger) Commit(key TranscriptSegmentKey)
func (*TranscriptSessionLedger) Release ¶ added in v0.48.0
func (l *TranscriptSessionLedger) Release(key TranscriptSegmentKey)
type TranscriptTransformer ¶ added in v0.45.0
type TranscriptTransformer interface {
Transform(ctx context.Context, transcript Transcript) (Transcript, error)
}
TranscriptTransformer can apply final post-STT changes after all audio segments have been transcribed and merged, but before command routing or user-visible output.
type TranscriptionDraftObserver ¶ added in v0.48.0
type TranscriptionDraftObserver interface {
OnTranscriptDraft(transcript Transcript)
}
TranscriptionDraftObserver is optionally implemented by observers that can surface live provider draft text. Drafts are never passed to output handlers.
type TranscriptionJob ¶
type TranscriptionJob struct {
Submission
Segments []Submission
Target any
}
TranscriptionJob pairs a Submission with its delivery target.
func (TranscriptionJob) Clone ¶
func (j TranscriptionJob) Clone() TranscriptionJob
type TranscriptionObserver ¶
type TranscriptionObserver interface {
OnState(status, text string)
OnLog(message, kind string)
OnTranscriptCommitted(transcript Transcript, quickNote bool)
}
TranscriptionObserver receives real-time status and log updates from a TranscriptionWorker during processing.
type TranscriptionRunner ¶
type TranscriptionRunner struct {
// contains filtered or unexported fields
}
TranscriptionRunner transcribes audio submissions and persists results. Create one with NewTranscriptionRunner.
func NewTranscriptionRunner ¶
func NewTranscriptionRunner(transcriber Transcriber, store Persistence) *TranscriptionRunner
NewTranscriptionRunner creates a TranscriptionRunner backed by the given transcriber and persistence store. Either argument may be nil.
func (*TranscriptionRunner) Commit ¶
func (r *TranscriptionRunner) Commit(ctx context.Context, submission Submission, transcript Transcript) (Completion, error)
func (*TranscriptionRunner) WithObserver ¶
func (r *TranscriptionRunner) WithObserver(observer CommitObserver) *TranscriptionRunner
type TranscriptionStore ¶
type TranscriptionStore interface {
SaveTranscription(ctx context.Context, text, language, provider, model string, durationMs, latencyMs int64, audioData []byte) error
}
TranscriptionStore persists completed dictation transcriptions.
type TranscriptionWorker ¶
type TranscriptionWorker struct {
// contains filtered or unexported fields
}
TranscriptionWorker processes TranscriptionJob values from an internal queue on a single goroutine. Start it with TranscriptionWorker.Start and submit work with TranscriptionWorker.Submit.
func NewTranscriptionWorker ¶
func NewTranscriptionWorker(cfg TranscriptionWorkerConfig) (*TranscriptionWorker, error)
func (*TranscriptionWorker) Close ¶
func (w *TranscriptionWorker) Close()
func (*TranscriptionWorker) HandleDictationStreamEvent ¶ added in v0.48.0
func (w *TranscriptionWorker) HandleDictationStreamEvent(ctx context.Context, event DictationStreamEvent, opts DictationStreamSinkOptions) error
HandleDictationStreamEvent routes provider-native dictation events through the same final-commit path as batch transcription. Interim/draft events are UI/status-only and must never call output or persistence.
func (*TranscriptionWorker) Start ¶
func (w *TranscriptionWorker) Start(ctx context.Context)
func (*TranscriptionWorker) Submit ¶
func (w *TranscriptionWorker) Submit(job TranscriptionJob) error
func (*TranscriptionWorker) Wait ¶
func (w *TranscriptionWorker) Wait()
type TranscriptionWorkerConfig ¶
type TranscriptionWorkerConfig struct {
Timeout time.Duration
QueueSize int
Runner *TranscriptionRunner
Output TranscriptOutput
Interceptor TranscriptInterceptor
Transformer TranscriptTransformer
Observer TranscriptionObserver
Ledger *TranscriptSessionLedger
// LowConfidenceThreshold flags recognized words below this acoustic
// confidence (0..1) so the host can surface likely-misrecognized terms.
// <= 0 disables the check. Only providers that expose per-word confidence
// (Deepgram, AssemblyAI) produce data here.
LowConfidenceThreshold float64
}
TranscriptionWorkerConfig configures a TranscriptionWorker. Runner is required; all other fields are optional.
type VoiceActivityDetector ¶ added in v0.24.0
VoiceActivityDetector is the public VAD contract consumed by DictationSegmenter. It intentionally matches SpeechKit's internal Silero detector shape without exposing internal packages.
type VoiceAgentService ¶ added in v0.24.0
type VoiceAgentService interface {
Start(context.Context) error
Stop(context.Context) (VoiceAgentSession, error)
SendText(context.Context, string) error
CurrentSession(context.Context) (VoiceAgentSession, error)
}
VoiceAgentService is the mode-scoped SDK contract for realtime dialogue.
type VoiceAgentSession ¶ added in v0.24.0
type VoiceAgentSession struct {
ID string `json:"id,omitempty"`
StartedAt time.Time `json:"startedAt,omitempty"`
EndedAt time.Time `json:"endedAt,omitempty"`
Locale string `json:"locale,omitempty"`
ProviderProfileID string `json:"providerProfileId,omitempty"`
RuntimeKind string `json:"runtimeKind,omitempty"`
Turns []VoiceAgentTurn `json:"turns,omitempty"`
Summary VoiceAgentSessionSummary `json:"summary"`
Speakers *speaker.DiarizationResult `json:"speakers,omitempty"`
}
VoiceAgentSession is the public record for a live dialogue.
type VoiceAgentSessionSummary ¶ added in v0.24.0
type VoiceAgentSessionSummary struct {
Title string `json:"title,omitempty"`
Summary string `json:"summary"`
Ideas []string `json:"ideas,omitempty"`
Decisions []string `json:"decisions,omitempty"`
OpenQuestions []string `json:"openQuestions,omitempty"`
NextSteps []string `json:"nextSteps,omitempty"`
RawText string `json:"rawText,omitempty"`
}
VoiceAgentSessionSummary is the structured handoff produced when a Voice Agent session ends.
type VoiceAgentSetting ¶ added in v0.24.0
type VoiceAgentSetting struct {
ModeSetting
SessionSummary bool `json:"sessionSummary"`
PipelineFallback bool `json:"pipelineFallback"`
CloseBehavior string `json:"closeBehavior,omitempty"`
AgentProfileID string `json:"agentProfileId,omitempty"`
AgentSequenceID string `json:"agentSequenceId,omitempty"`
}
type VoiceAgentTurn ¶ added in v0.24.0
type VoiceAgentTurn struct {
Role string `json:"role"`
Text string `json:"text"`
CreatedAt time.Time `json:"createdAt,omitempty"`
SpeakerLabel string `json:"speakerLabel,omitempty"`
PersonID string `json:"personId,omitempty"`
DisplayName string `json:"displayName,omitempty"`
SpeakerConfidence float64 `json:"speakerConfidence,omitempty"`
}
VoiceAgentTurn is one finalized turn in a realtime or fallback dialogue.
type WordConfidence ¶ added in v0.46.0
WordConfidence is a recognized word with the provider's per-word acoustic confidence in [0,1]. It mirrors stt.WordConfidence (the kernel keeps Transcript decoupled from the stt package; stt.AsTranscriber is the public bridge that maps between them). nil when the provider does not expose word-level confidence.
Source Files
¶
Directories
¶
| Path | Synopsis |
|---|---|
|
Package agentbridge defines the framework-neutral seam through which a SpeechKit host fronts an external coding agent (adopted 2026-08-10, AI-VOICE-SPEECHKIT-TARGET.md "External Coding Agent Bridge").
|
Package agentbridge defines the framework-neutral seam through which a SpeechKit host fronts an external coding agent (adopted 2026-08-10, AI-VOICE-SPEECHKIT-TARGET.md "External Coding Agent Bridge"). |
|
codex
Package codex drives the official OpenAI Codex binary as an agentbridge.Agent.
|
Package codex drives the official OpenAI Codex binary as an agentbridge.Agent. |
|
voicetools
Package voicetools binds an agentbridge.Agent to the voice agent's tool surface with the "Call GPT" semantics (owner decision 2026-08-10, AI-VOICE-SPEECHKIT-TARGET.md External Coding Agent Bridge):
|
Package voicetools binds an agentbridge.Agent to the voice agent's tool surface with the "Call GPT" semantics (owner decision 2026-08-10, AI-VOICE-SPEECHKIT-TARGET.md External Coding Agent Bridge): |
|
Package agentkit provides a small Go harness for building SpeechKit Voice Agent hosts.
|
Package agentkit provides a small Go harness for building SpeechKit Voice Agent hosts. |
|
Package assist provides an embeddable Assist Mode service.
|
Package assist provides an embeddable Assist Mode service. |
|
genkitadapter
Package genkitadapter keeps Genkit-specific Assist wiring out of the core public assist package.
|
Package genkitadapter keeps Genkit-specific Assist wiring out of the core public assist package. |
|
skills
Package skills exposes SpeechKit's Voice-Companion skill catalog — Time, Date, Math, Weather, Timer, Reminder, Wikipedia, plus a fail-closed Home Assistant boundary — as a public assist.ToolMatcher + assist.ToolExecutor pair, ready to plug into an assist.Service.
|
Package skills exposes SpeechKit's Voice-Companion skill catalog — Time, Date, Math, Weather, Timer, Reminder, Wikipedia, plus a fail-closed Home Assistant boundary — as a public assist.ToolMatcher + assist.ToolExecutor pair, ready to plug into an assist.Service. |
|
toolbridge
Package toolbridge adapts Assist-mode tools (assist.ToolMatcher / assist.ToolExecutor — the deterministic skill layer, e.g.
|
Package toolbridge adapts Assist-mode tools (assist.ToolMatcher / assist.ToolExecutor — the deterministic skill layer, e.g. |
|
Package audio provides the shared PCM audio primitives for the SpeechKit capture format: 16kHz S16 mono constants, WAV framing, duration math, and RMS level estimation.
|
Package audio provides the shared PCM audio primitives for the SpeechKit capture format: 16kHz S16 mono constants, WAV framing, duration math, and RMS level estimation. |
|
capture
Package capture is the public microphone / system-audio capture layer of the SpeechKit framework: backend registry (RegisterBackend, Open), capture Session contract, device enumeration (ListCaptureDevices, ListOutputDevices), and the pooled PCM frame buffers (FramePool).
|
Package capture is the public microphone / system-audio capture layer of the SpeechKit framework: backend registry (RegisterBackend, Open), capture Session contract, device enumeration (ListCaptureDevices, ListOutputDevices), and the pooled PCM frame buffers (FramePool). |
|
Package client provides a typed HTTP client for talking to a remote SpeechKit Server (the `cmd/speechkit-server` Linux container or any compatible deployment).
|
Package client provides a typed HTTP client for talking to a remote SpeechKit Server (the `cmd/speechkit-server` Linux container or any compatible deployment). |
|
Package companion provides small composers for hands-free SpeechKit hosts.
|
Package companion provides small composers for hands-free SpeechKit hosts. |
|
Package customize defines SpeechKit's public Words/Replacements contract.
|
Package customize defines SpeechKit's public Words/Replacements contract. |
|
Package deviceagent implements the credential-minimal LAN-side SpeechKit device-agent client and its versioned wire contract.
|
Package deviceagent implements the credential-minimal LAN-side SpeechKit device-agent client and its versioned wire contract. |
|
Package dictation provides an embeddable strict Dictation runtime.
|
Package dictation provides an embeddable strict Dictation runtime. |
|
Package hostconfig turns a SpeechKit TOML configuration file into the public SDK types an embedding host drives the framework with: a speechkit.ModeSettings (which modes are on, their hotkeys and selected provider profiles) and a permissive speechkit.RuntimePolicy (which modes the host exposes and whether fallbacks are allowed).
|
Package hostconfig turns a SpeechKit TOML configuration file into the public SDK types an embedding host drives the framework with: a speechkit.ModeSettings (which modes are on, their hotkeys and selected provider profiles) and a permissive speechkit.RuntimePolicy (which modes the host exposes and whether fallbacks are allowed). |
|
internal
|
|
|
Package lifecycle owns mode start/stop orchestration and refcounted shared dependencies for SpeechKit hosts.
|
Package lifecycle owns mode start/stop orchestration and refcounted shared dependencies for SpeechKit hosts. |
|
Package localization resolves stable SpeechKit message IDs against the repository-owned locale catalogs.
|
Package localization resolves stable SpeechKit message IDs against the repository-owned locale catalogs. |
|
Package netsec provides centralized network security primitives used by every HTTP-based provider in SpeechKit (STT, TTS, LLM, downloads).
|
Package netsec provides centralized network security primitives used by every HTTP-based provider in SpeechKit (STT, TTS, LLM, downloads). |
|
Package procguard ties long-lived child processes to the lifetime of the process that spawned them.
|
Package procguard ties long-lived child processes to the lifetime of the process that spawned them. |
|
Package provideropts defines SpeechKit's provider-neutral voice option vocabulary and the manifest/resolve types used by concrete provider adapters.
|
Package provideropts defines SpeechKit's provider-neutral voice option vocabulary and the manifest/resolve types used by concrete provider adapters. |
|
Package speaker defines SpeechKit's public speaker diarization and attribution contracts.
|
Package speaker defines SpeechKit's public speaker diarization and attribution contracts. |
|
Package storage defines the public storage-backend contract: backend capabilities and metadata, install/device/user/tenant scopes with their enforcement policies, and the configuration shape hosts use to construct a backend.
|
Package storage defines the public storage-backend contract: backend capabilities and metadata, install/device/user/tenant scopes with their enforcement policies, and the configuration shape hosts use to construct a backend. |
|
Package stt defines the SpeechKit speech-to-text provider interface and houses the concrete provider implementations: whisper.cpp (local built-in), HuggingFace, OpenAI, Groq, Google, an OpenAI-compatible adapter (covers Ollama and other compatible servers), and the self-hosted VPS adapter.
|
Package stt defines the SpeechKit speech-to-text provider interface and houses the concrete provider implementations: whisper.cpp (local built-in), HuggingFace, OpenAI, Groq, Google, an OpenAI-compatible adapter (covers Ollama and other compatible servers), and the self-hosted VPS adapter. |
|
allproviders
Package allproviders is the batteries-included STT assembly layer: it knows every provider SpeechKit ships and turns a host's resolved credentials into a ready stt.Router.
|
Package allproviders is the batteries-included STT assembly layer: it knows every provider SpeechKit ships and turns a host's resolved credentials into a ready stt.Router. |
|
sttcontract
Package sttcontract provides a reusable conformance suite that every stt.STTProvider implementation is expected to satisfy.
|
Package sttcontract provides a reusable conformance suite that every stt.STTProvider implementation is expected to satisfy. |
|
vps
Package vps is the self-hosted whisper-server provider for SpeechKit: an OpenAI-compatible endpoint a user runs themselves, so audio never reaches a commercial provider.
|
Package vps is the self-hosted whisper-server provider for SpeechKit: an OpenAI-compatible endpoint a user runs themselves, so audio never reaches a commercial provider. |
|
Package tts exposes the embeddable SpeechKit text-to-speech surface.
|
Package tts exposes the embeddable SpeechKit text-to-speech surface. |
|
ttscontract
Package ttscontract provides a reusable conformance suite that every tts.Provider implementation is expected to satisfy.
|
Package ttscontract provides a reusable conformance suite that every tts.Provider implementation is expected to satisfy. |
|
Package ttsroute holds the single source of truth that maps a Voice-Output profile ID (e.g.
|
Package ttsroute holds the single source of truth that maps a Voice-Output profile ID (e.g. |
|
Package voiceagent provides an embeddable Voice Agent service.
|
Package voiceagent provides an embeddable Voice Agent service. |
|
cascaded
Package cascaded implements a turn-based STT -> LLM -> TTS voice agent provider.
|
Package cascaded implements a turn-based STT -> LLM -> TTS voice agent provider. |
|
live
Package live exposes the low-level Voice Agent realtime-protocol types.
|
Package live exposes the low-level Voice Agent realtime-protocol types. |
|
live/livecontract
Package livecontract provides reusable conformance checks for LiveProvider implementations.
|
Package livecontract provides reusable conformance checks for LiveProvider implementations. |
|
local
Package local implements voiceagent.Provider on top of an in-process live session — realtime voice agents (Deepgram Voice Agent, Gemini Live, OpenAI Realtime, AssemblyAI, cascaded) without a speechkit-server.
|
Package local implements voiceagent.Provider on top of an in-process live session — realtime voice agents (Deepgram Voice Agent, Gemini Live, OpenAI Realtime, AssemblyAI, cascaded) without a speechkit-server. |
|
Package wakeword exposes embeddable SpeechKit wake-word contracts.
|
Package wakeword exposes embeddable SpeechKit wake-word contracts. |
|
sherpa
Package sherpa exposes the sherpa-onnx wake-word detector adapter.
|
Package sherpa exposes the sherpa-onnx wake-word detector adapter. |