constant

package
v1.11.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jul 11, 2026 License: MIT Imports: 1 Imported by: 0

Documentation

Index

Constants

This section is empty.

Variables

View Source
var (
	OLLAMA    = LLMProvider{Name: "ollama", ApiUrl: "http://localhost:11434", Models: []Model{QWEN_3_6}}
	ANTHROPIC = LLMProvider{Name: "anthropic", ApiUrl: "https://api.anthropic.com", Models: []Model{SONNET_4_6, OPUS_4_8}}
	DEEPSEEK  = LLMProvider{Name: "deepseek", ApiUrl: "https://api.deepseek.com", Models: []Model{DEEPSEEK_V4_FLASH, DEEPSEEK_V4_PRO}}
	// OPENAI — all currently listed models are reasoning-class (gpt-5 / o-series).
	// If a non-reasoning model (gpt-4*, gpt-3.5*) is added here later, update
	// pkg/llm/openai/client.go isReasoningModel to match.
	OPENAI = LLMProvider{Name: "openai", ApiUrl: "https://api.openai.com", Models: []Model{GPT_5_4_MINI, GPT_5_5}}
	// GLM — Zhipu AI / z.ai, reached over its Anthropic-compatible endpoint
	// (same wire format as ANTHROPIC; pkg/llm/glm copies that engine and swaps
	// auth to Authorization: Bearer). ApiUrl is the z.ai Anthropic gateway, so
	// pkg/llm/glm hits ApiUrl + "/v1/messages".
	GLM = LLMProvider{Name: "glm", ApiUrl: "https://api.z.ai/api/anthropic", Models: []Model{GLM_4_6, GLM_5_2}}
	// QWEN — Alibaba Cloud Tongyi/Qwen via the DashScope OpenAI-compatible
	// endpoint (same wire format as OPENAI/DEEPSEEK; pkg/llm/qwen copies the
	// deepseek engine, Bearer auth, thinking via enable_thinking). ApiUrl is the
	// international (Singapore) gateway; pkg/llm/qwen hits ApiUrl + "/chat/completions".
	// China users override api_url to https://dashscope.aliyuncs.com/compatible-mode/v1.
	QWEN = LLMProvider{Name: "qwen", ApiUrl: "https://dashscope-intl.aliyuncs.com/compatible-mode/v1", Models: []Model{QWEN3_7_PLUS, QWEN3_7_MAX}}
)

Use var for struct "constants" Models are ordered by price

View Source
var MODEL_CONTEXT_SIZE = map[Model]int{
	QWEN_3_6:          262_144,
	SONNET_4_6:        500_000,
	OPUS_4_8:          500_000,
	DEEPSEEK_V4_FLASH: 1_000_000,
	DEEPSEEK_V4_PRO:   1_000_000,
	GPT_5_4_MINI:      400_000,
	GPT_5_5:           1_050_000,
	GLM_4_6:           200_000,
	GLM_5_2:           1_000_000,
	QWEN3_7_PLUS:      1_000_000,
	QWEN3_7_MAX:       1_000_000,
}

for completion

View Source
var MODEL_PRICING = map[Model]Pricing{

	QWEN_3_6: {},

	SONNET_4_6: {Input: 3.00, Output: 15.00, CacheRead: 0.30, CacheWrite: 3.75},
	OPUS_4_8:   {Input: 5.00, Output: 25.00, CacheRead: 0.50, CacheWrite: 6.25},

	DEEPSEEK_V4_FLASH: {Input: 0.14, Output: 0.28, CacheRead: 0.14, CacheWrite: 0.14},
	DEEPSEEK_V4_PRO:   {Input: 1.74, Output: 3.48, CacheRead: 1.74, CacheWrite: 1.74},

	GPT_5_4_MINI: {Input: 0.75, Output: 4.50, CacheRead: 0.75, CacheWrite: 0.75},
	GPT_5_5:      {Input: 5.00, Output: 30.00, CacheRead: 5.00, CacheWrite: 5.00},

	GLM_4_6: {Input: 0.43, Output: 1.74, CacheRead: 0.11, CacheWrite: 0.43},
	GLM_5_2: {Input: 1.40, Output: 4.40, CacheRead: 0.26, CacheWrite: 1.40},

	QWEN3_7_PLUS: {Input: 0.40, Output: 1.60, CacheRead: 0.08, CacheWrite: 0.40},
	QWEN3_7_MAX:  {Input: 2.50, Output: 7.50, CacheRead: 2.50, CacheWrite: 2.50},
}

MODEL_PRICING is the per-model USD rate card, keyed like MODEL_CONTEXT_SIZE. Rates are USD per 1M tokens and were verified against each provider's published list price via web search in June 2026 — treat them as sane defaults for the cost HUD, not invoices; an operator on negotiated rates can edit them here. A model absent from this map is "unpriced" and the HUD hides its dollar figure rather than guessing.

Functions

func CostOf added in v1.8.0

func CostOf(m Model, in, out, cacheRead, cacheWrite int) (cost float64, ok bool)

CostOf prices a usage slice for model m against MODEL_PRICING. ok is false when m has no rate card — the caller should then hide the cost rather than render a misleading $0.00. A model priced at zero (e.g. a local Ollama model) returns (0, true): genuinely free, not unknown.

Types

type AgentStatus

type AgentStatus string
const (
	INIT         AgentStatus = "init"
	DRAINING     AgentStatus = "draining"
	THINKING     AgentStatus = "thinking"
	TEXTING      AgentStatus = "texting"
	EXECUTING    AgentStatus = "executing"
	INTERRUPTED  AgentStatus = "interrupted"
	IDLE         AgentStatus = "idle"
	MAX_ITERS    AgentStatus = "max_iters"
	SAVING       AgentStatus = "saving"
	COMPACTING   AgentStatus = "compacting"
	READY_REPORT AgentStatus = "ready_report"
	SHUTDOWN     AgentStatus = "shutdown"
	CRUSHED      AgentStatus = "crushed"
)

func (AgentStatus) String

func (a AgentStatus) String() string

type LLMProvider

type LLMProvider struct {
	Name   string
	ApiUrl string
	Models []Model
}

func GetAllProviders

func GetAllProviders() []LLMProvider

func GetProvider

func GetProvider(name string) (LLMProvider, bool)

func ProviderOfModel added in v1.4.3

func ProviderOfModel(m Model) (LLMProvider, bool)

ProviderOfModel returns the provider that lists m in its Models. Lets a caller holding only a model id (e.g. a per-member `model:` pin in a swarm profile.yml) derive the provider to route through.

func (LLMProvider) ModelForLevel

func (p LLMProvider) ModelForLevel(level int) Model

ModelForLevel returns the model to use at a given capability tier:

  • level 1 (default) → the smallest / cheapest model the provider lists (Models[0]). Suitable for routine work; "normal" tier.
  • level 2 → the largest model the provider lists (Models[len-1]). Suitable for hard reasoning; "big" tier. More expensive — the LLM should only pick this when the task genuinely needs deeper thinking.

Providers with only one configured model collapse both tiers to that model. Levels outside {1, 2} clamp to the nearest valid tier.

type Model

type Model string
const (
	// OLLAMA
	QWEN_3_6 Model = "qwen3.6"

	// ANTHROPIC
	SONNET_4_6 Model = "claude-sonnet-4-6"
	OPUS_4_8   Model = "claude-opus-4-8"

	// DEEPSEEK
	DEEPSEEK_V4_FLASH Model = "deepseek-v4-flash"
	DEEPSEEK_V4_PRO   Model = "deepseek-v4-pro"

	// OPENAI
	GPT_5_4_MINI Model = "gpt-5.4-mini"
	GPT_5_5      Model = "gpt-5.5"

	// GLM (Zhipu / z.ai). glm-4.6 is the cheaper coding model (normal tier),
	// glm-5.2 the flagship (big tier). Both are reached over the Anthropic-
	// compatible endpoint; vision input requires a vision-capable GLM model.
	GLM_4_6 Model = "glm-4.6"
	GLM_5_2 Model = "glm-5.2"

	// QWEN (Alibaba DashScope). qwen3.7-plus is the cheaper tier, qwen3.7-max the
	// flagship (1M ctx, native extended-thinking). OpenAI-compatible wire.
	QWEN3_7_PLUS Model = "qwen3.7-plus"
	QWEN3_7_MAX  Model = "qwen3.7-max"
)

func GetModel

func GetModel(name string) (Model, bool)

GetModel resolves a string to its typed Model constant. Mirrors GetProvider; the canonical model registry is MODEL_CONTEXT_SIZE.

type Pricing added in v1.8.0

type Pricing struct {
	Input      float64 // uncached input
	Output     float64 // output — hidden reasoning tokens are a subset, billed here
	CacheRead  float64 // cached-prefix re-read
	CacheWrite float64 // cached-prefix first write
}

Pricing is a model's USD rate card, every rate expressed per 1,000,000 tokens. A zero-value Pricing means "genuinely free" (a locally hosted Ollama model costs nothing to run) — distinct from a model that is absent from MODEL_PRICING, which is "unpriced / unknown" (see CostOf).

CacheRead / CacheWrite price Anthropic-style prompt-cache tokens: a READ re-uses a previously cached prefix at a steep discount, while the first WRITE of that prefix carries a surcharge. Providers without prompt caching report zero cache tokens (see llm.Usage), so for them these rates are never exercised and simply mirror Input.

func (Pricing) CostUSD added in v1.8.0

func (p Pricing) CostUSD(in, out, cacheRead, cacheWrite int) float64

CostUSD prices one usage slice. The arguments mirror llm.Usage: in is the FULL input-token count, of which cacheRead + cacheWrite are a subset (Anthropic's accounting), and out is the output-token count (hidden reasoning tokens are already inside out). Cached tokens are billed at their own rates and carved out of the uncached-input pool so they are never double-charged.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL