Documentation
¶
Overview ¶
Package openai implements inference.Provider over any OpenAI-compatible chat-completions endpoint that returns token logprobs — OpenAI itself, NVIDIA's hosted NemoGuard models, vLLM, LiteLLM and most gateways.
Each Infer call becomes one single-token completion: the candidate labels are the allowed answer set, the model is asked for exactly one of them, and the response is the softmax of the labels' logprobs, renormalized over the allowed set — read from the endpoint's top_logprobs rather than a dedicated logit-scoring API. Like any such approach the numbers are uncalibrated.
Index ¶
Constants ¶
const ( // NemoGuardModel is NVIDIA's hosted topic-control chat model, served // through an OpenAI-compatible chat-completions endpoint. NemoGuardModel = "nvidia/llama-3.1-nemoguard-8b-topic-control" )
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Config ¶
type Config struct {
BaseURL string
APIKey string
Model string // the chat model that answers; overridable per-Request
Timeout time.Duration
HTTPClient *http.Client
// RetryPolicy governs retries of transient failures (429, 502-504,
// network errors). Nil uses providers.DefaultRetryPolicy().
RetryPolicy *pipeline.RetryPolicy
}
Config configures a Provider.
type Provider ¶
type Provider struct {
// contains filtered or unexported fields
}
Provider answers inference.Request calls using a chat model's logprobs over a single-token completion.
func (*Provider) HTTPTimeout ¶
HTTPTimeout reports the per-call timeout this Provider will apply. Exported for tests: the timeout is only observable otherwise by waiting for it to expire.