codeaf-replay

command
v0.2.2-rc.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 18, 2026 License: Apache-2.0 Imports: 14 Imported by: 0

Documentation

Overview

codeaf-replay judges a chooser by REGRET ON TEN DAYS OF LOG rather than by a screenshot.

IT EXISTS BECAUSE WE WERE TUNING A BANDIT BY ANECDOTE. Every change to the machine chooser in the week to 2026-09-11 — #850, #853, #873 and two lanes still in flight — was argued from one person's screenshot and a ten-row grep, and every one of them moved a number nobody then measured. The model-call log already holds about a thousand finished requests a day: which machine was asked for, which one answered, how long its first token took, how fast it wrote, whether it refused, and what the call was for. That is a bandit replay dataset, and this is the instrument that reads it as one.

WHAT IT DOES, in one paragraph. It walks the log in time order. Every finished request becomes a question — "on this model, at this moment, for this role, which machine would you have demanded?" — put to each candidate policy, whose beliefs have been fed exactly the sightings and refusals that had arrived by then and not one that had not. The answer is priced against what that machine measurably did around that moment (world.go), the best machine available is priced the same way, and the difference is the regret. The table it prints is per role class, because a second of a watched answer and a second of an unattended errand are not the same second.

IT IS A DEVELOPER'S BINARY AND NEVER A VERB ON codeaf, for the reason cmd/codeaf-census and cmd/codeaf-changes are: the shipped binary is on a checked-in byte budget (SIZE-BUDGET) and a measuring tool must not spend the product's weight.

Usage:

codeaf-replay [flags]

  -log PATH         the model-call log (default: this machine's own)
  -sightings PATH   the lane observation journal (default: beside the belief file)
  -since DATE       ignore rows before this RFC3339 prefix
  -window D         how far either side of a request a machine's own answers
                    are read for what it was doing then (default 15m)
  -min N            the fewest machines a model needs before its requests are
                    scored at all (default 2 — nothing to choose between)
  -incidents N      how many of the worst moments to show every candidate at
  -out PATH         write the report to a file instead of stdout

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL