Documentation
¶
Overview ¶
Package llm handles LLM-based generation tasks: prompt rendering, template embedding, runner invocation, and output parsing. Finders (read side) and handlers (write side) both consume this package — it owns the "call an LLM with a structured prompt" concern.
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func ComputePromptHash ¶
ComputePromptHash returns the hex-encoded SHA-256 hash of the rendered prompt. Uses the first 16 bytes (32 hex chars) — enough for collision avoidance.
func FormatEntryForPrompt ¶
FormatEntryForPrompt formats an entry as readable text for inclusion in a prompt.
Types ¶
type Embedder ¶ added in v0.4.0
type Embedder interface {
// EmbedDocuments embeds texts as index-side passages, applying the
// configured DocumentTemplate. Used by the indexer at build /
// lazy-fill time.
EmbedDocuments(ctx context.Context, texts []string) ([][]float32, error)
// EmbedQueries embeds texts as retrieval-side queries, applying the
// configured QueryTemplate. Used by the search finder at query time.
EmbedQueries(ctx context.Context, texts []string) ([][]float32, error)
Dimensions() int
Fingerprint() string
// BatchSize is the per-call input cap the embedder targets. Drivers
// the indexer's outer bucketing so each Embed call corresponds to
// exactly one HTTP round-trip — progress callbacks fire per batch
// instead of "all-at-once at the end" of a giant cross-entry call.
// Larger inputs still embed correctly; the embedder splits
// internally and the indexer just sees a slower individual call.
BatchSize() int
}
Embedder turns a batch of texts into dense vectors. Implementations are thin transport adapters around an embedding service (OpenAI-compatible `/v1/embeddings`, Ollama `/api/embed`); construction is via embed.New.
The interface splits embedding into a document-side and a query-side because instruction-tuned encoders (Qwen3, E5, Nomic, BGE) want asymmetric prefixes — `passage:` on documents, `query:` on queries, or `Instruct: …\nQuery:…` only on queries. The split makes the call-site intent explicit so the wrong template can't silently slip onto the wrong side. Untemplated models (OpenAI text-embedding-3) treat both methods as equivalent.
Dimensions returns the vector length the implementation produces — must be stable across calls so the index can validate row shape.
Fingerprint is an opaque identifier for the (provider + model + truncation + document template) tuple. Used by the search index to detect rows whose embedding is stale after a configuration change. The format is implementation-defined, but implementations must guarantee that two embedders that produce distribution-comparable document vectors share a fingerprint, and two that don't, don't. The query template deliberately does NOT factor into the fingerprint — query template changes affect retrieval quality but never invalidate indexed embeddings, so they're a free-tweak knob (see EmbeddingConfig).
type LLMMetadata ¶
type LLMMetadata struct {
TotalCostUSD float64
InputTokens int
OutputTokens int
CacheReadTokens int
CacheCreateTokens int
NumTurns int
Duration time.Duration
DurationAPI time.Duration
Models map[string]ModelUsage
}
LLMMetadata holds agent-neutral per-call metrics from the LLM provider.
type ModelUsage ¶
type ModelUsage struct {
InputTokens int
OutputTokens int
CacheReadTokens int
CacheCreateTokens int
CostUSD float64
}
ModelUsage holds per-model token and cost metrics.
type PreflightResult ¶
type PreflightResult struct {
Findings []Finding
}
PreflightResult holds the parsed findings from a pre-flight validator run. An empty Findings slice means the validator reported no findings.
func Preflight ¶
func Preflight(ctx context.Context, runner Runner, entry *model.Entry, graph *model.Graph, configuredLanguage string) (*PreflightResult, error)
Preflight runs the pre-flight validator against the given entry and graph. Returns the parsed result regardless of finding severity. Returns an error only for infrastructure failures (runner error, template error, parse error).
Participant validation is handled by mechanical checks in the finders layer (see finders.mechanicalPreflight) — the LLM participant-drift rubric is retired per plan d-cpt-d34 AC 9.
configuredLanguage is the graph authoring language (locale code) from `.sdd/config.yaml` (empty string when unset — English default). It feeds the language-drift check which flags entries whose description prose does not match the configured language.
func (*PreflightResult) HasBlocking ¶
func (r *PreflightResult) HasBlocking() bool
HasBlocking reports whether any finding blocks entry creation. Currently only SeverityHigh blocks.
type Request ¶ added in v0.2.0
Request carries the two-part prompt submitted to a Runner. SystemPrompt holds the stable portion (instructions, structural rules) — providers that support prompt caching treat this as the cacheable prefix. UserPrompt holds the per-call variable portion (entry content, refs). Runners that can't distinguish system from user concatenate them with SystemPrompt first.
func RenderSummaryPrompt ¶
RenderSummaryPrompt renders the summary prompt for an entry. Returns a Request with the full rendered prompt in UserPrompt; the system/user split is introduced when templates are refactored (see the plan decision).
func (Request) Combined ¶ added in v0.2.0
Combined returns SystemPrompt followed by UserPrompt separated by a blank line when both are non-empty. Runners without native system-prompt support use this to flatten the Request into a single payload. The hash used for summary-skip detection is computed over this combined form so changes to either half invalidate the cached summary.
type RunResult ¶
type RunResult struct {
Text string
Meta *LLMMetadata
}
RunResult holds the LLM response text and optional metadata.
type Runner ¶
Runner executes a structured LLM request and returns the response with metadata. The implementation decides which model and transport to use. Injected so tests can substitute fakes.
type Severity ¶
type Severity string
Severity classifies a pre-flight finding. Mirrored in the query package; templates describe severity in purely semantic terms.
type SummarizeResult ¶
SummarizeResult holds the generated summary and its prompt hash.
func Summarize ¶
func Summarize(ctx context.Context, runner Runner, entry *model.Entry, graph *model.Graph, force bool) (*SummarizeResult, error)
Summarize generates a summary for a single entry using the LLM runner. Returns nil if the entry's stored SummaryHash matches the computed hash (skip). Set force to regenerate regardless.
Directories
¶
| Path | Synopsis |
|---|---|
|
Package claude implements llm.Runner by invoking the Claude CLI.
|
Package claude implements llm.Runner by invoking the Claude CLI. |
|
Package embed provides Embedder implementations for OpenAI-compatible (`/v1/embeddings`) and Ollama (`/api/embeddings`) endpoints, plus a factory that dispatches by configured provider and wraps remote providers with a rate.Limiter (paralleling the chat-runner factory).
|
Package embed provides Embedder implementations for OpenAI-compatible (`/v1/embeddings`) and Ollama (`/api/embeddings`) endpoints, plus a factory that dispatches by configured provider and wraps remote providers with a rate.Limiter (paralleling the chat-runner factory). |
|
Package factory resolves an llm.Runner from model.LLMConfig.
|
Package factory resolves an llm.Runner from model.LLMConfig. |
|
Package gollm implements llm.Runner on top of github.com/teilomillet/gollm, providing a unified adapter for Anthropic API, OpenAI, Ollama, and other providers supported by gollm.
|
Package gollm implements llm.Runner on top of github.com/teilomillet/gollm, providing a unified adapter for Anthropic API, OpenAI, Ollama, and other providers supported by gollm. |