judge

package
v0.4.1-rc.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 22, 2026 License: Apache-2.0 Imports: 10 Imported by: 0

Documentation

Overview

Package judge scores a finished run seat by seat, independently of the crew that produced it.

THE MECHANISM. A run's record names which model held each judged seat — worker, high, mastermind — and what the run delivered. Every seat the record carries is put to one model that was not in the crew: the caller supplies Ask, which sends a system and a user prompt and returns the answer. One question is asked per seat, in role order, and each answer is read as one JSON object carrying a score from 0 to 100 and a one-sentence reason. That is the same scale crewpick reads seat quality on and a pool records its role_quality metric in, so a run's seats land on the scale the picker already uses.

THE CALLER OWNS THE CALL. This package touches no disk and no network and holds no credentials: the model, the transport and the bill are the caller's, and the caller bills the judge's calls to its own seat. Pick chooses that model from a catalog, and the chosen id rides in the record (Record.Judge) so a seat already holding it is skipped rather than scoring its own work.

Index

Constants

View Source
const DefaultFloor = 60

DefaultFloor is the published coding index a judge must reach to be picked. A judge below the worker's own class rubber-stamps the work rather than scoring it: it has less of the capability the seat is scored on than the seat it is reading.

Variables

This section is empty.

Functions

func Candidates

func Candidates(models []catalog.Model, crew map[Role]string, floor float64) []string

Candidates returns every catalog row that can judge a crew, cheapest first.

A row qualifies exactly as the judge a crew would otherwise be handed does: its published coding index reaches floor, it carries "tools", no crew seat holds it, and its vendor differs from the worker's. Rows the provider published no price for are out of the running. TWO MORE ROWS ARE NEVER CANDIDATES: one whose prompt and completion prices sum to zero, and one whose id ends in ":free" whatever the case. A free row is a rate-limited row and not a price — the provider answers it with a 429 at the moment a judge most needs an answer — so it is asked for nothing, whatever its index says.

The list is ordered by cost and then by id, so the cheapest stands first and the answer is a property of the catalog rather than of its order. A crew with nothing left to draw from yields no candidates.

func Pick

func Pick(models []catalog.Model, crew map[Role]string, floor float64) (judgeID string, ok bool)

Pick returns the cheapest judge Candidates names, and ok is false when it names none. It is Candidates' head-or-nothing, so the properties its callers rely on hold here too: the cheaper of two rows wins, and a tie breaks to the lower id.

Types

type Ask

type Ask func(ctx context.Context, system, user string) (answer string, err error)

Ask puts one question to the judge model and returns its answer. The caller owns the model, the credentials and the transport, and bills the call to the judge's own seat; this package writes both prompts and reads the answer.

type Record

type Record struct {
	// Brief is what the run was asked to do.
	Brief string
	// Deliverable is what had to exist when it was over.
	Deliverable string
	// Report is the account the run landed with.
	Report string
	// Claim is the run's own summary of its answer, kept beside the report.
	Claim string
	// Ending is why a run that stopped early stopped where it did.
	Ending string
	// Files is every path the run wrote.
	Files []string
	// Changed is how many paths it changed.
	Changed int
	// Checks is the repeatable verification the run ran.
	Checks []string
	// Seats is which model held each judged seat. The chat door asks no planner
	// call, so a record may carry only the worker and the high seats.
	Seats map[Role]string
	// Judge is the model that will answer the seat questions, the id [Pick]
	// returned. A seat already holding it is skipped, so a model never scores
	// its own work; an empty Judge names no model and skips nothing.
	Judge string
	// CostUSD is what the run spent, in dollars.
	CostUSD float64
}

Record is one finished run as the judge reads it: the spec it ran under, its own account of itself, what it touched, and who held each seat. It is the judge's own copy, assembled by the caller from wherever the run is recorded, so this package imports neither the session store nor anything else that owns the run.

type Role

type Role string

Role is one of a crew's judged seats.

const (
	// RoleWorker is the seat that carried the work.
	RoleWorker Role = "worker"
	// RoleHigh is the seat that checked the work.
	RoleHigh Role = "high"
	// RoleMastermind is the seat that planned the work.
	RoleMastermind Role = "mastermind"
)

type Score

type Score struct {
	Role   Role
	Model  string
	Score  float64
	Reason string
}

Score is one seat's reading: the role, the model that held it, the score on the 0-100 scale and the judge's one-sentence reason.

func Judge

func Judge(ctx context.Context, ask Ask, rec Record) ([]Score, error)

Judge asks one question per seat present in rec.Seats, in role order, and answers with a Score for each seat that was read. A seat held by the judge's own model (Record.Judge) is skipped and named in the returned error; a seat whose call fails, or whose answer is not one JSON object carrying an integer score from 0 to 100, is an error for that seat alone and never stops the others. The scores obtained are returned beside a joined error naming every seat that failed, and a record with no seats returns no scores and no error.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL