inference

package
v2.9.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 30, 2026 License: Apache-2.0 Imports: 11 Imported by: 0

Documentation

Overview

Package inference defines a single generic Provider interface for every classification-shaped inference call PromptKit makes against a vendor API (zero-shot text classification, topic control, moderation, and similar label-scoring tasks).

The design rule is one (role, type) = one vendor API: a Provider is a thin codec over that API's wire shape (Request in, Response out) and nothing more. It does not know what task it is serving — "is this on-topic", "is this toxic", "what emotion is this" are all the same shape of call (candidate labels in, scored labels out) to the vendor, so they share one interface here. Task semantics — which labels to ask for, how to interpret the highest-scoring label, what threshold makes a guardrail fail closed — live entirely in the callers (the eval/guardrail handlers), never in a Provider implementation.

Index

Constants

This section is empty.

Variables

View Source
var (
	// ErrModelLoading is returned when the provider reports the model is
	// still loading after retries are exhausted.
	ErrModelLoading = errors.New("inference: model still loading after retries")
	// ErrModelNotSupported is returned when the provider rejects the
	// configured model as unsupported by the inference path in use.
	ErrModelNotSupported = errors.New("inference: model not supported by the configured inference path")
	// ErrLabelsRequired is returned by providers that need candidate labels
	// (Request.Labels) but received none.
	ErrLabelsRequired = errors.New("inference: this provider needs candidate labels")
)

Functions

func RegisterFactory

func RegisterFactory(providerType string, f Factory)

RegisterFactory registers a factory for the given provider type. Typically called from a per-backend package init().

func RegisteredTypes

func RegisteredTypes() []string

RegisteredTypes returns the inference provider types with a registered factory, sorted. Use it to check a configured type before CreateFromSpec rather than constructing and parsing the error.

func ResolveCredential

func ResolveCredential(
	ctx context.Context,
	providerType string,
	cfgDir string,
	cred *credentials.CredentialConfig,
) (credentials.Credential, error)

ResolveCredential is a thin wrapper around base.ResolveCredential, matching tts.ResolveCredential / stt.ResolveCredential / classify.ResolveCredential.

func WithEmitter

func WithEmitter(ctx context.Context, e *events.Emitter) context.Context

WithEmitter returns ctx with e attached. Instrument uses the attached emitter to publish inference.call.completed / inference.call.failed events; a Provider wrapped by Instrument is a no-op passthrough for telemetry when the context carries no emitter.

func WithRegistry

func WithRegistry(ctx context.Context, r *Registry) context.Context

WithRegistry returns ctx with the Registry attached.

Types

type Factory

type Factory = base.Factory[Provider]

Factory builds a Provider from a spec. Per-backend packages register one via init() so this package never imports them.

type LabelScore

type LabelScore struct {
	Label string
	Score float64
}

LabelScore is a single candidate label and the provider's score for it.

type Provider

type Provider interface {
	Infer(ctx context.Context, req Request) (Response, error)
}

Provider is implemented by every inference backend: one vendor API, codec'd to and from the shapes above.

func CreateFromSpec

func CreateFromSpec(spec ProviderSpec) (Provider, error)

CreateFromSpec builds a Provider for the spec's Type.

func Instrument

func Instrument(p Provider, id, providerType string) Provider

Instrument wraps p so every Infer call times the request and, when ctx carries an events.Emitter (see WithEmitter), publishes an inference.call.completed or inference.call.failed event carrying the provider id, provider type, model, duration and cost. id identifies the configured provider instance (e.g. "hf", "openai-moderation"); providerType is the vendor type backing it (e.g. "huggingface", "openai").

The wrapped Infer's result is returned unchanged; instrumentation never alters the response or error. With no emitter on ctx, Instrument is a transparent passthrough. Instrument(nil, ...) returns nil.

type ProviderSpec

type ProviderSpec = base.CapabilitySpec

ProviderSpec is the runtime form of an inference-provider declaration. Aliased to base.CapabilitySpec so the field shape is shared with the TTS, STT, embedding, and image factories (id/type/model/base_url/credential/ additional_config).

type Registry

type Registry struct {
	// contains filtered or unexported fields
}

Registry holds named Provider instances keyed by an id supplied at config time (e.g. "hf", "openai-moderation"). Callers look up by id; the id-to-provider mapping is the only thing a handler config needs to know.

Providers register themselves into a Registry at engine startup; the registry then travels with context.Context to handlers via WithRegistry / FromContext.

func FromContext

func FromContext(ctx context.Context) *Registry

FromContext returns the Registry attached to ctx, or nil if none. Callers that don't find a registry fall back to a "skipped" result rather than failing — inference is an optional feature of the runtime, not a hard dependency.

func NewRegistry

func NewRegistry() *Registry

NewRegistry returns an empty Registry.

func (*Registry) Get

func (r *Registry) Get(id string) (Provider, error)

Get resolves id, or the default when id is "". Returns a non-nil error when nothing matches.

func (*Registry) IDs

func (r *Registry) IDs() []string

IDs returns every registered provider id, sorted.

func (*Registry) Register

func (r *Registry) Register(id string, p Provider) error

Register adds p under id. Duplicate ids return an error. The first registration made against this Registry becomes the default.

func (*Registry) SetDefault

func (r *Registry) SetDefault(id string) error

SetDefault names the provider used when a caller passes an empty id to Get. The id must already be registered.

type Request

type Request struct {
	// Model overrides the provider's configured model.
	Model string
	// Inputs is the content under judgment; roles are preserved.
	Inputs []types.Message
	// Labels lists candidate answers; empty means the model's own label set.
	Labels []string
	// Prompt is instruction text. Fixed-head backends reject a non-empty Prompt.
	Prompt string
	// Params carries API-level tweaks (e.g. "multi_label": true).
	Params map[string]any
}

Request is the vendor-agnostic shape of an inference call. Not every field applies to every provider: a fixed-head classifier (e.g. a dedicated moderation endpoint) has no use for Labels or Prompt and rejects a non-empty Prompt; a zero-shot provider requires Labels.

type Response

type Response struct {
	Model  string
	Scores []LabelScore // probabilities, highest first
	Usage  Usage
	Raw    string // provider's unparsed answer, for diagnostics
}

Response is the vendor-agnostic shape of an inference result.

func (Response) Score

func (r Response) Score(label string) (float64, bool)

Score returns the probability for label (case-insensitive) and whether it was present in the response.

type Usage

type Usage struct {
	InputTokens int
	Cost        float64 // USD, when the API reports it
}

Usage carries billing/telemetry data for one Infer call.

Directories

Path Synopsis
Package all blank-imports every inference provider so their factories self-register.
Package all blank-imports every inference provider so their factories self-register.
Package huggingface implements inference.Provider over the HuggingFace Inference API: HF's serverless router (router.huggingface.co/hf-inference) or a dedicated HF Inference Endpoint.
Package huggingface implements inference.Provider over the HuggingFace Inference API: HF's serverless router (router.huggingface.co/hf-inference) or a dedicated HF Inference Endpoint.
Package openai implements inference.Provider over any OpenAI-compatible chat-completions endpoint that returns token logprobs — OpenAI itself, NVIDIA's hosted NemoGuard models, vLLM, LiteLLM and most gateways.
Package openai implements inference.Provider over any OpenAI-compatible chat-completions endpoint that returns token logprobs — OpenAI itself, NVIDIA's hosted NemoGuard models, vLLM, LiteLLM and most gateways.
Package systemone implements inference.Provider over the "typed decision" wire protocol: POST {base}/systemone with a shared state and a named typed question, answered with a probability distribution rather than generated text.
Package systemone implements inference.Provider over the "typed decision" wire protocol: POST {base}/systemone with a shared state and a named typed question, answered with a probability distribution rather than generated text.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL