Documentation
¶
Overview ¶
Package huggingface implements inference.Provider over the HuggingFace Inference API: HF's serverless router (router.huggingface.co/hf-inference) or a dedicated HF Inference Endpoint.
A single Provider serves three shapes of call, chosen by the Request's content: a candidate-label list routes to HF's zero-shot-classification pipeline, an audio or image part in Inputs routes to the model's own classification endpoint with the raw media bytes, and plain text with no labels routes to the model's own classification endpoint with a JSON body. Embeddings are a separate provider (role: embedding) — HF's feature-extraction API has its own wire shape and does not belong here.
Index ¶
Constants ¶
const DefaultBaseURL = "https://router.huggingface.co/hf-inference"
DefaultBaseURL is the canonical HF Inference API endpoint. HF deprecated the older `api-inference.huggingface.co/models/{id}` URL in favor of the Inference Providers router; `hf-inference` is the provider that serves the same free-tier serverless models. Override via Config.BaseURL for HF Inference Endpoints (dedicated paid hosts) or to pin a different provider (e.g. replicate, fireworks).
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Config ¶
type Config struct {
// APIKey is the HF token.
APIKey string
// BaseURL overrides DefaultBaseURL. Used for HF Inference Endpoints
// (e.g. https://my-endpoint.xxx.endpoints.huggingface.cloud) and the
// Inference Providers routing layer.
BaseURL string
// Dedicated, when true, treats BaseURL as a fully-specified inference
// endpoint and skips the /models/{model_id} suffix that the public
// Inference API requires. Set this when pointing at HF Inference
// Endpoints (the paid dedicated host shape).
Dedicated bool
// Model is the provider's configured default model, used when a
// Request doesn't carry its own Model. Mirrors openai.Config.Model.
Model string
// HTTPClient lets the caller provide a custom transport (test
// httptest server, timeouts, retry middleware). Default is a
// 60s-timeout client.
HTTPClient *http.Client
}
Config configures a Provider.
type Provider ¶
type Provider struct {
// contains filtered or unexported fields
}
Provider answers inference.Request calls against the HF Inference API.
func (*Provider) ApplyHTTPTuning ¶ added in v2.12.0
func (p *Provider) ApplyHTTPTuning(t base.HTTPTuning) error
ApplyHTTPTuning replaces the provider's HTTP client with a tuned copy, implementing base.HTTPTunable so a provider file's headers, request_timeout and http_transport reach this backend.
func (*Provider) Infer ¶
Infer routes req to the right HF pipeline: a media part (audio or image) in Inputs goes to the model's classification endpoint as raw bytes; otherwise the last Inputs message's text goes to the model's own pipeline (Labels empty) or the zero-shot-classification pipeline (Labels set).