Documentation
¶
Overview ¶
Package inference defines a single generic Provider interface for every classification-shaped inference call PromptKit makes against a vendor API (zero-shot text classification, topic control, moderation, and similar label-scoring tasks).
The design rule is one (role, type) = one vendor API: a Provider is a thin codec over that API's wire shape (Request in, Response out) and nothing more. It does not know what task it is serving — "is this on-topic", "is this toxic", "what emotion is this" are all the same shape of call (candidate labels in, scored labels out) to the vendor, so they share one interface here. Task semantics — which labels to ask for, how to interpret the highest-scoring label, what threshold makes a guardrail fail closed — live entirely in the callers (the eval/guardrail handlers), never in a Provider implementation.
Index ¶
- Variables
- func RegisterFactory(providerType string, f Factory)
- func RegisteredTypes() []string
- func ResolveCredential(ctx context.Context, providerType string, cfgDir string, ...) (credentials.Credential, error)
- func WithEmitter(ctx context.Context, e *events.Emitter) context.Context
- func WithRegistry(ctx context.Context, r *Registry) context.Context
- type Factory
- type LabelScore
- type Provider
- type ProviderSpec
- type Registry
- type Request
- type Response
- type Usage
Constants ¶
This section is empty.
Variables ¶
var ( // ErrModelLoading is returned when the provider reports the model is // still loading after retries are exhausted. ErrModelLoading = errors.New("inference: model still loading after retries") // ErrModelNotSupported is returned when the provider rejects the // configured model as unsupported by the inference path in use. ErrModelNotSupported = errors.New("inference: model not supported by the configured inference path") // ErrLabelsRequired is returned by providers that need candidate labels // (Request.Labels) but received none. ErrLabelsRequired = errors.New("inference: this provider needs candidate labels") )
Functions ¶
func RegisterFactory ¶
RegisterFactory registers a factory for the given provider type. Typically called from a per-backend package init().
func RegisteredTypes ¶
func RegisteredTypes() []string
RegisteredTypes returns the inference provider types with a registered factory, sorted. Use it to check a configured type before CreateFromSpec rather than constructing and parsing the error.
func ResolveCredential ¶
func ResolveCredential( ctx context.Context, providerType string, cfgDir string, cred *credentials.CredentialConfig, ) (credentials.Credential, error)
ResolveCredential is a thin wrapper around base.ResolveCredential, matching tts.ResolveCredential / stt.ResolveCredential / classify.ResolveCredential.
func WithEmitter ¶
WithEmitter returns ctx with e attached. Instrument uses the attached emitter to publish inference.call.completed / inference.call.failed events; a Provider wrapped by Instrument is a no-op passthrough for telemetry when the context carries no emitter.
Types ¶
type Factory ¶
Factory builds a Provider from a spec. Per-backend packages register one via init() so this package never imports them.
type LabelScore ¶
LabelScore is a single candidate label and the provider's score for it.
type Provider ¶
Provider is implemented by every inference backend: one vendor API, codec'd to and from the shapes above.
func CreateFromSpec ¶
func CreateFromSpec(spec ProviderSpec) (Provider, error)
CreateFromSpec builds a Provider for the spec's Type.
func Instrument ¶
Instrument wraps p so every Infer call times the request and, when ctx carries an events.Emitter (see WithEmitter), publishes an inference.call.completed or inference.call.failed event carrying the provider id, provider type, model, duration and cost. id identifies the configured provider instance (e.g. "hf", "openai-moderation"); providerType is the vendor type backing it (e.g. "huggingface", "openai").
The wrapped Infer's result is returned unchanged; instrumentation never alters the response or error. With no emitter on ctx, Instrument is a transparent passthrough. Instrument(nil, ...) returns nil.
type ProviderSpec ¶
type ProviderSpec = base.CapabilitySpec
ProviderSpec is the runtime form of an inference-provider declaration. Aliased to base.CapabilitySpec so the field shape is shared with the TTS, STT, embedding, and image factories (id/type/model/base_url/credential/ additional_config).
type Registry ¶
type Registry struct {
// contains filtered or unexported fields
}
Registry holds named Provider instances keyed by an id supplied at config time (e.g. "hf", "openai-moderation"). Callers look up by id; the id-to-provider mapping is the only thing a handler config needs to know.
Providers register themselves into a Registry at engine startup; the registry then travels with context.Context to handlers via WithRegistry / FromContext.
func FromContext ¶
FromContext returns the Registry attached to ctx, or nil if none. Callers that don't find a registry fall back to a "skipped" result rather than failing — inference is an optional feature of the runtime, not a hard dependency.
func (*Registry) Get ¶
Get resolves id, or the default when id is "". Returns a non-nil error when nothing matches.
func (*Registry) Register ¶
Register adds p under id. Duplicate ids return an error. The first registration made against this Registry becomes the default.
func (*Registry) SetDefault ¶
SetDefault names the provider used when a caller passes an empty id to Get. The id must already be registered.
type Request ¶
type Request struct {
// Model overrides the provider's configured model.
Model string
// Inputs is the content under judgment; roles are preserved.
Inputs []types.Message
// Labels lists candidate answers; empty means the model's own label set.
Labels []string
// Prompt is instruction text. Fixed-head backends reject a non-empty Prompt.
Prompt string
// Params carries API-level tweaks (e.g. "multi_label": true).
Params map[string]any
}
Request is the vendor-agnostic shape of an inference call. Not every field applies to every provider: a fixed-head classifier (e.g. a dedicated moderation endpoint) has no use for Labels or Prompt and rejects a non-empty Prompt; a zero-shot provider requires Labels.
Directories
¶
| Path | Synopsis |
|---|---|
|
Package all blank-imports every inference provider so their factories self-register.
|
Package all blank-imports every inference provider so their factories self-register. |
|
Package huggingface implements inference.Provider over the HuggingFace Inference API: HF's serverless router (router.huggingface.co/hf-inference) or a dedicated HF Inference Endpoint.
|
Package huggingface implements inference.Provider over the HuggingFace Inference API: HF's serverless router (router.huggingface.co/hf-inference) or a dedicated HF Inference Endpoint. |
|
Package openai implements inference.Provider over any OpenAI-compatible chat-completions endpoint that returns token logprobs — OpenAI itself, NVIDIA's hosted NemoGuard models, vLLM, LiteLLM and most gateways.
|
Package openai implements inference.Provider over any OpenAI-compatible chat-completions endpoint that returns token logprobs — OpenAI itself, NVIDIA's hosted NemoGuard models, vLLM, LiteLLM and most gateways. |
|
Package systemone implements inference.Provider over the "typed decision" wire protocol: POST {base}/systemone with a shared state and a named typed question, answered with a probability distribution rather than generated text.
|
Package systemone implements inference.Provider over the "typed decision" wire protocol: POST {base}/systemone with a shared state and a named typed question, answered with a probability distribution rather than generated text. |