crewpick

package
v0.2.2-rc.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 18, 2026 License: Apache-2.0 Imports: 4 Imported by: 0

Documentation

Overview

Package crewpick picks one model per seat for a three-seat crew by reading the pareto front of expected task bill against crew quality.

It deliberately imports nothing from the rest of the tree and touches no disk and no network: the caller owns the candidate list and the seat shapes, and the same candidates give the same front however often they are asked about and in whatever order they arrive.

The law it implements: every (worker, high, mastermind) combination the candidates can field is scored by the bill its seats' token volumes run up and by the mean of its three seat qualities. A candidate is in the running for a seat only when it meets the family and the seat's needs and its quality reaches Floor of that seat's best in the family, and a crew is refused outright when its high seat comes from the worker's own vendor. The crews that no other crew beats on both bill and quality survive as the front, sorted by bill; Presets and AtKnob read picks off that front.

A candidate that publishes some but not all of its three indexes is scored on the ones it publishes: each missing index is estimated from the call's candidates that publish both it and one of its measured ones — the donors — never above the call's largest measured value of that index, and the result says which of a pick's indexes were estimated. A candidate that publishes none of its indexes, or no price at all, is out of the running.

Index

Constants

View Source
const Floor = 0.80

Floor is the share of a seat's best quality in the family that a candidate must reach to stay in the running for that seat. A model that cannot do the job is not cheap, whatever it costs.

View Source
const PriorWeightAt = 30

PriorWeightAt is the observation count at which a pool's own rating for a model carries the same weight in a seat's quality as the catalog's published indexes: at N observations the rating is worth half the seat and the catalog the other half. A rating with no observations never enters.

Variables

This section is empty.

Functions

func DefaultShapes

func DefaultShapes() map[Seat]SeatShape

DefaultShapes returns the three seat shapes crews are picked for. The worker takes half its quality from the agentic index and reads barely anything but cached input; the high seat judges the three indexes evenly, needs tools and images both, and sees the same token mix on a sliver of the volume; the mastermind reads intelligence and agentic only, on a terse ten-in-one mix with no cache and a fraction of the volume.

func SeatCost

func SeatCost(c Candidate, shape SeatShape) float64

SeatCost is the seat's expected cost of a candidate, in dollars per 1M input-equivalent tokens: the prompt price blended with the cache-read price by the seat's cache share, plus the completion price spread over the seat's input-to-output ratio. A candidate that publishes no cache-read price pays the prompt price on the cache share too.

func SeatQuality

func SeatQuality(c Candidate, shape SeatShape, pool []Candidate) float64

SeatQuality is the candidate's quality for the seat on a 0-100 scale: each index counted against the pool's best for that index and weighted by the seat's weights. The pool sets the scale, so the same candidate scores differently in a stronger pool. A missing index is estimated from the pool's donors exactly as in Front. A row that is not a candidate at all, and any row against a pool with no candidates to raise a scale, score zero.

func SeatQualityWith

func SeatQualityWith(c Candidate, shape SeatShape, pool []Candidate, seat Seat, prior Prior) float64

SeatQualityWith is SeatQuality with the pool's prior blended in for the seat given: a rating the prior holds for the candidate's canonical id, with a positive observation count, moves the seat's quality towards the rating's mean by N/(N+PriorWeightAt), and a seat the prior never rates — or rates with no observations — keeps the catalog quality unchanged.

Types

type Candidate

type Candidate struct {
	// ID names the candidate; the text before its first slash is the vendor.
	ID string
	// Open marks a candidate the open family may draw from.
	Open bool
	// Intelligence, Coding and Agentic are the three capability indexes
	// quality is read from.
	Intelligence float64
	Coding       float64
	Agentic      float64
	// PromptPrice, CompletionPrice and CacheReadPrice are the dollars per
	// 1M tokens of each kind.
	PromptPrice     float64
	CompletionPrice float64
	CacheReadPrice  float64
	// HasCacheRead reports whether CacheReadPrice is published at all.
	HasCacheRead bool
	// Context is the model's context window, in tokens.
	Context int
	// Images and Tools are what the model can take and call.
	Images bool
	Tools  bool
}

A Candidate is one model up for a seat. Prices are dollars per 1M tokens.

A candidate carrying none of the three indexes — zero values, which is how unpublished indexes read — or carrying a zero prompt price is not a candidate at all: it never occupies a seat, never clears a floor, and raises no ceiling. A candidate carrying some but not all of its indexes is scored on the ones it publishes, each missing one estimated from the donors among its fellow candidates, and the result says which of a pick's indexes those were. A model that publishes no cache-read price says so with HasCacheRead false and pays the prompt price on the cache share too.

type Cell

type Cell struct {
	Role  string
	Model string
	Mean  float64
	N     int
}

A Cell is one measurement a caller hands PriorFromCells: a role, a model id and the pool's mean quality with the observations behind it. It is the shape a measurement document reads into, so crewpick need not import the reader.

type Crew

type Crew struct {
	Worker     string
	High       string
	Mastermind string
	Bill       float64
	Quality    float64
	// Estimated marks, for the model in each seat, which of its indexes
	// were scored by estimate rather than measurement: the row is the seat
	// and the column the index. A crew picked on measured indexes alone is
	// all false.
	Estimated [3][3]bool
	// Measured marks, for the model in each seat, whether a pool's own
	// rating for that model entered the seat's quality: the row is the seat
	// and the value says the seat's model carried a rating with a positive
	// observation count. A crew picked with no prior, or whose seats the
	// prior never rated, is all false.
	Measured [3]bool
}

A Crew is one (worker, high, mastermind) pick. Bill is the expected dollars per 1M task tokens once each seat's share of the task is paid for at its seat cost; Quality is the mean of the three seat qualities on a 0-100 scale.

func AtKnob

func AtKnob(front []Crew, k float64) Crew

AtKnob reads one pick off the front at knob k in [0, 1]: the budget runs geometrically from the frugal bill to the max bill — lo*(hi/lo)^k — and the pick is the best-quality crew whose bill fits it, so k of 0 is frugal and k of 1 is max. An empty front picks nothing.

func Front

func Front(candidates []Candidate, shapes map[Seat]SeatShape, fam Family) []Crew

Front returns the pareto front of (bill, quality) over every crew the candidates can field in the family, sorted by bill: a crew on it is beaten by no other crew on both, and every crew off it is beaten by one on it.

shapes must carry an entry for every seat. Per seat the running is narrowed twice before crews are enumerated — family and the seat's needs first, then Floor against that seat's best in the family — and a crew is refused while its high seat comes from the worker's own vendor, so the two seats that see the same work never share one shop. Ties in (bill, quality) break to the crew scored on fewer estimated indexes, and ties past that by id order — worker, then high, then mastermind — which is what makes the front a property of the candidates rather than of their order.

func FrontWith

func FrontWith(candidates []Candidate, shapes map[Seat]SeatShape, fam Family, prior Prior) []Crew

FrontWith is Front with the pool's prior blended into each seat's quality: a rating the prior holds for a seat's model, with a positive observation count, moves that seat's quality towards the rating's mean by N/(N+PriorWeightAt), the seat's measured flag says so, and the floor and the crew's quality read the blended value. A nil prior leaves every seat on the catalog quality, exactly as Front does. The front is a property of the candidates and the prior, never of the order they arrive in.

func Presets

func Presets(front []Crew) (frugal, balanced, max Crew)

Presets reads the three named picks off a front: frugal is the cheapest crew, max the dearest, and balanced the knee between them — the crew farthest above the straight line from one end of the front to the other in (ln bill, quality), which is where money stops buying quality in earnest. A front of one crew is all three; an empty front is none.

type Family

type Family int

A Family is which candidates a crew may be drawn from.

const (
	// Open draws only from candidates marked Open.
	Open Family = iota
	// All draws from every candidate.
	All
)

type Index

type Index int

An Index is one of the three capability indexes quality is read from — intelligence, coding, agentic — in the order the seat weights and a candidate's reading carry them.

const (
	Intelligence Index = iota
	Coding
	Agentic
)

type Prior

type Prior map[Seat]map[string]Rating

A Prior is a pool's measured quality per seat, keyed by canonical model id: the rating a seat's quality is blended towards, by how many observations back it. A nil Prior is no prior at all, and every seat keeps the catalog quality.

func MergePriors

func MergePriors(a, b Prior) Prior

MergePriors folds two priors into one. Seat by seat, a rating both hold for a model is combined by observation count — the two means weighted by the counts behind them, the counts added — and a rating either holds alone is carried whole. A nil prior on either side is the other, and two nils are nil. The priors given are read and never changed, so the answer is never one of them.

func PriorFromCells

func PriorFromCells(cells []Cell, minInstalls int, canonical func(string) string) Prior

PriorFromCells reads a pool's measurements into a Prior: a cell names a seat by its role word, a model by an id that canonical resolves to the canonical id a candidate is looked up by, and carries a mean and a count. A cell whose role names no seat, or whose count is below minInstalls, is ignored, and so is a count of zero or less. A nil canonical leaves the id as it stands. When no cell survives the prior is nil, which is the same as no prior at all.

The floor here keeps counting rows rather than the installs a measurement document carries beside them: the reader applies the document's floor before these cells arrive — on installs where a cell spells them — and rows cannot be fewer than the installs that produced them, so a cell that floor kept meets this one too. The cells built outside a document carry rows alone, so an installs field on Cell would sit unset on every caller, and there is none.

type Rating

type Rating struct {
	Mean float64
	N    int
}

A Rating is a pool's measured quality for one model on one seat, on the same 0-100 scale quality is read on, with the number of observations behind it.

type Seat

type Seat int

A Seat is one of a crew's three roles: the worker carries the volume of the work, the high seat judges and repairs what the worker writes, and the mastermind spends a little context steering the whole task.

const (
	Worker Seat = iota
	High
	Mastermind
)

type SeatShape

type SeatShape struct {
	// Weights reads quality from the three indexes: intelligence, coding,
	// agentic.
	Weights [3]float64
	// InOut is the seat's input tokens per output token; the completion
	// price is spread over it.
	InOut float64
	// CacheShare is the share of the seat's input tokens served from cache.
	CacheShare float64
	// NeedTools and NeedImages are what a model must take and call to sit
	// the seat.
	NeedTools  bool
	NeedImages bool
	// MinContext is the window a model must have to sit the seat.
	MinContext int
	// Volume is the seat's share of a task's tokens, and the weight its
	// cost carries in a crew's bill.
	Volume float64
}

A SeatShape is how one seat uses whatever model sits it: the weights quality is read from, the token mix cost is computed on, what the model must be able to take and call, the window it must have, and the share of a task's tokens the seat burns.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL