ouroboros

package
v0.3.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 7, 2026 License: Apache-2.0 Imports: 36 Imported by: 0

README

Ouroboros

Mine what your agents did for the next rule they need, and prove the rule before it ships.

What it is · What you can do · The loop · Quick start · Worked example · Write a miner · Status · Reference

What it is

CSF runs AI coding agents under rules: every agent works in a harness session, and every shell command and commit it makes passes a gate. Gates are only as good as the mistakes someone thought of in advance. Ouroboros finds the ones nobody did.

Ouroboros uses evaluated results to propose the next version of an agent's instructions or an allowed workflow; RRSI is its method.

That sentence is the ontology's definition, unchanged. This directory is its mining half. A miner reads a corpus (harness event logs, tickets, pull requests, files at a revision), turns it into facts, and derives verdicts with a Datalog rule. A miner is accepted only when a walk-forward backtest finds every labeled offense it should, and an accepted miner becomes the next gate. The gated runs then become the corpus for the next miner: the snake eats its tail.

You write two files, extract.ml and rules.dl, on a contract that is already built.

What you can do with it

  • Turn a complaint into a gate. A mistake the operator flags becomes labeled instances; a miner that finds every one of them, and nothing clean, closes the ticket.
  • Prove a rule before you enforce it. Thresholds are fitted on the earlier half of the record and judged on the later half, so a miner cannot pass by memorizing its examples.
  • See the evidence behind every finding. Each verdict carries its proof: the facts it used and the source line of each.
  • Measure how early a rule would have fired. The backtest reports lead time: how long before the operator noticed, the miner would have.
  • Close tickets only on proof. tools/close-ticket.sh closes a ticket only when the merged pull request's backtest shows no missed offense.

The loop

Figure 1 shows one turn of the loop.

The Ouroboros loop

Figure 1. The Ouroboros loop. A miner extracts facts from the corpus, a Datalog rule derives verdicts with proofs, and a walk-forward backtest decides: a miner with no missed offense becomes a typed finding that can close its ticket, and the next runs become the next corpus.

Each arrow is one typed artifact, and each has a check: extract.ml and rules.dl must pass miner_test, the backtest block is generated (never typed), and tools/close-ticket.sh closes a ticket only when the merged pull request's Backtest FN row is 0 and every check row exits 0.

Quick start

Copy the template and give it your miner's name (in extract.ml, set name, package and verdicts):

cp -r services/ouroboros/miners/_template services/ouroboros/miners/my_miner

Test, build and run the template as it ships. This is real output from this branch, with Bazel's progress lines, the findings' proof arrays and the backtest's first rows trimmed (…):

$ tools/bazel.sh test //services/ouroboros/miners/_template:miner_test --test_output=all
…
draft-pr-late: 4 tests passed
…
//services/ouroboros/miners/_template:miner_test                (cached) PASSED in 0.0s

$ tools/bazel.sh build //services/ouroboros/miners/_template:miner
INFO: Build completed successfully, 4 total actions

$ bazel-bin/services/ouroboros/miners/_template/miner.exe findings --knee 468 'services/ouroboros/miners/_template/fixtures/*/events.jsonl'
{"miner":"draft-pr-late","rule":"invisible","subject":[{"text":"6c1f2a10-3b7e-4d6a-9e51-0a1b2c3d4e31"}],"severity":"SEVERITY_S3","proof":[…]}
{"miner":"draft-pr-late","rule":"invisible","subject":[{"text":"6c1f2a10-3b7e-4d6a-9e51-0a1b2c3d4e60"}],"severity":"SEVERITY_S3","proof":[…]}

$ bazel-bin/services/ouroboros/miners/_template/miner.exe backtest services/ouroboros/miners/_template/fixtures/labels.tsv 'services/ouroboros/miners/_template/fixtures/*/events.jsonl'
…
| Backtest $\kappa^\star$ | 468 (argmax precision on $\Lambda_{\le t}$ s.t. FN = 0; 3 candidates from `score`) |
| Backtest TP | 1: 6c1f2a10-3b7e-4d6a-9e51-0a1b2c3d4e60 |
| Backtest FP | 0 |
| Backtest FN | 0 |

The miner's verbs are facts, findings [--knee K], backtest [--json] LABELS ITEM... mutate [--json] LABELS ITEM... and scope. Swap the fixture glob for '<state>/*/events.jsonl', where <state> is the harness state directory, to run over every recorded run.

Worked example: DRAFT-PR-LATE

The template flags a harness run whose work stayed invisible to the operator, with no draft pull request, for longer than a knee measured from the runs themselves (#102). Its fixtures are three real runs.

Facts (miner.exe facts on the fixtures; score is seconds from start to the first draft pull request, or to the last event when there is none):

Run gated opened score (s) Source lines
4e31 yes yes 3089 2, 3
4e59 yes yes 468 2, 3
4e60 yes no 1259 2, 3

Rule (rules.dl):

invisible(R) :- gated(R), score(R, S), knee(K), gt(S, K).

Labels (fixtures/labels.tsv), ordered by start and split at $t$ = 2026-10-02T18:04:45Z:

Run Label Half
4e31 + fit
4e59 − fit
4e60 + accept

Knee. The candidates are the measured scores, $K = {468, 1259, 3089}$, and the fit half chooses among them:

$$\kappa^\star = \arg\max_{\kappa \in K} \mathrm{precision}(\kappa; \Lambda_{\le t}) \ \text{s.t.}\ \mathrm{FN}(\kappa; \Lambda_{\le t}) = 0$$

$\kappa$ Fires in fit half Precision FN
468 4e31 1 0
1259 4e31 1 0
3089 none 0 1

Ties keep the smallest knee, which fires earliest, so $\kappa^\star = 468$.

Backtest (miners/_template/backtest.md, generated, and checked current by miner_test):

Measured Result
TP 1 (4e60)
FP 0
FN 0
Lead 304 s

Lead is $t_{\mathrm{flag}} - (t_{\mathrm{start}} + \kappa^\star) = 772 - 468 = 304$ s: the miner would have fired five minutes before the operator flagged 4e60. The finding carries its proof, gated(4e60), score(4e60, 1259) and knee(468), each with its source line.

Writing a miner

  1. Read the ticket and every comment: its named instances are your labels and its predicate is your rule.
  2. Copy miners/_template to miners/<name>; in extract.ml set name, package and verdicts.
  3. Write extract: one corpus item in, facts out, each with its source span. Emit a score fact if the rule needs a knee.
  4. Write rules.dl with its /* predicate */ block; thresholds read knee(K).
  5. Put each labeled instance in fixtures/<instance>/events.jsonl (only the lines a fact needs) and list it in fixtures/labels.tsv: instance, + or -, start, flagged time or -, source.
  6. Run tools/bazel.sh test //services/ouroboros/miners/<name>:miner_test until it passes; paste the blocks it prints into backtest.md and mutation.md. A surviving mutant is printed with the rules or labels that survived: strengthen the fixtures or the rule until the test kills it.
  7. Write the README from the template's, with the agent-drafted marker and the backtest block.
  8. Run python3 tools/check_operator_identifiers.py, commit with the CSF-Session and CSF-Model trailers, push, and open a draft pull request.
  9. Run bash tools/check-merge.sh; on exit 0 mark the pull request ready. Never merge.

Pitfalls

Each row is a mistake seen in tonight's proposals or in the template's own history.

Pitfall Counterexample (wrong) Example (right)
nested term gt(L, knee(K)) knee(K), gt(L, K)
written threshold gt(L, 2000) knee(K), gt(L, K)
unbound negation bad(X) :- ~seen(X). bad(X) :- run(X), ~seen(X).
relation no extractor emits latency(Q, L) extract.ml emits latency
label that is not an instance #222 6c1f2a10-…-4e60
knee fit on the judged labels fit and accept on $\Lambda$ fit $\Lambda_{\le t}$, accept $\Lambda_{>t}$
no positive after the split + labels only before $t$ a + in the later half
hand-set knee $\kappa = 900$: lead −128 s $\kappa^\star = 468$: lead 304 s
>= for "longer than" ge(S, K): $\kappa^\star = 1259$, lead −487 s gt(S, K)
test blind to a written knee gt(S, 468) passes every label check the knee-binds check: a knee below every score and one above fire different runs

The engine checks neither safety nor stratification, so Contract.check_rules does: an unbound negation raises Invalid_rules "negation before its variables are bound", and every miner's test runs it first.

The service

The Go package in this directory is the loop itself, mounted into csf serve (#120): miners run all the time over the harness corpus, a pre-check decides at no model cost which tickets a fixer can take, fixers run through the harness under a daily budget, the merge train merges what passes the gate, and the compounding rate is measured and shown on the ops view. It replaces the ad hoc scripts that launched fixers outside the session gates and receipts.

csf serve -database-config database.json -ouroboros-repository /path/to/checkout            # detector, pre-check, measure, merge train
csf serve -database-config database.json -ouroboros-repository /path/to/checkout -ouroboros-fixers   # and fixer launch
csf ouroboros precheck -repository /path/to/checkout [-ticket N]...                             # the pre-check by hand, one row per ticket
curl http://127.0.0.1:14120/api/ouroboros                                                        # the latest numbers as JSON

The ledger is six csf_ouroboros_* tables of CSF's schema (ipc/db/csfpg); every proposal, verdict, fixer session, merge result and series point is a row there. Fixers launch unattended only when sandboxed launch is available; until then -ouroboros-fixers is the switch, and without it the loop detects, pre-checks, measures and merges.

Triggers

Four cron triggers in America/Los_Angeles, the zone #330's measurement uses; every cadence is derived from a measurement, and every occurrence skips while the previous one runs.

Trigger Cadence Derived from
ouroboros.detector every 15 min, and on every change to a run's event log after a 10 s settle the template miner reads every log (125 runs, 240 MB) in 1.3 s; an executor logs several lines a second, so the settle bounds reruns to one per burst and the sweep is the catch-up after a restart
ouroboros.fixers every 10 min the mean fixer session is 21.8 min, so a fixer is reconciled about twice before it ends
ouroboros.merge_train every 30 min one run of tools/merge-pr.sh builds the checkers and takes about 10 min; the train takes the ready pull requests oldest first
ouroboros.measure hourly, and once at start the compounding number is a per-day series; an hour keeps the day's point current
The pre-check

Before any model is paid for a ticket, its latest proposal (the comment the overnight loop posted on every mining ticket) must label at least one instance that exists in the corpus and is of a kind the proposed extractor reads: a run directory for jsonl, an existing issue or pull request other than the ticket itself for issue and pull_request, a commit in the repository for file_at. A proposal that names none gets the verdict needs_labels and waits in the labeler's queue; one that does gets launch.

Replayed over the 107 tickets whose fixers ran on 2026-10-03 (csf ouroboros precheck -ticket ..., 0 model calls):

Outcome on 2026-10-03 Tickets Skipped by the pre-check Fixer spend the skip would have avoided
synthetic labels 85 84 $60.78 of $61.90
miner_test failed 11 11 $11.78 of $11.78
fabricated fixtures 5 5 $3.48 of $3.48
ready pull request 3 2 $1.08 of $2.21
error 3 3 $0.00 of $0.00
all 107 105 $77.12 of $79.37

The two that pass are #249 (its proposal names pull requests #226 and #245, and its fixer reached a ready pull request) and #122 (its proposal names commit a1c14af; its fixer still wrote synthetic labels, which no pre-check can foresee). The one ready pull request the check would have skipped, #236, came from a proposal that named no instance: the fixer found one on its own.

Fixers and the budget

A fixer is one harness session: the fixer agent on claude-opus-5-5 (the only model the harness admits, by the operator's ruling of 2026-10-05), a worktree of its own on miner/<name>-<assignment prefix>, the session gates on, and a brief that names the instances the pre-check found. The loop launches at most one per ticket, stops when the harness's admission reports no room, and stops for the day when the spend reaches the cap:

Threshold Value Derived from
daily cap $100 operator ruling, 2026-10-04
reservation per running fixer $0.74 the measured mean fixer session, $79.37 over 107 on 2026-10-03, until the ledger has 10 finished fixers to measure its own mean from
budget day calendar day in America/Los_Angeles the measurement's day

A fixer's cost is folded from its event log: the executor reports total_cost_usd cumulatively within one process and from zero when the process is reopened, so the cost is the sum of the maximum of each non-decreasing run. Once a fixer's one turn is done the loop closes the session and reads the outcome from its pull request: ready, draft or no_pull_request.

The merge train

The train runs tools/merge-pr.sh over every ready, non-LANG pull request, oldest first, once per head; a refused head gets a comment with the refusal lines and is retried only after its author pushes a new one. NO-SELF-MERGE (#249): the merger is csf-serve/<pid>, never a session, and a pull request whose commits carry the merger's own identity in a CSF-Session: trailer is refused without running the path.

The compounding number

The objective of #120 is the struggle rate's compounding rate, shown with its interval. Amendment 1 (operator rulings, 2026-10-04) fixes the primary series as #330's weekly one and adds an internal check. What the loop measures, exactly:

Measured Definition
struggle episode the deterministic part of #330's definition (rrsi-mine traces): a hit is a tool result with is_error, an executor permission_denied, a session gate denial, or a tool call equal to one of the previous 4 (rrsi-mine's retry at exact equality in place of 0.85 similarity); hits at most 6 events apart are one episode; every episode counts, since no labeler marks noise, canonical vocabulary or harness-fixability here, so this rate is an upper bound on #330's
tool calls every tool_use block of every assistant record, subagents included: the per-session denominator #330 reads from the harness run record
struggle rate, primary episodes per 1,000 tool calls per week (weeks start Monday, days in America/Los_Angeles), with the 95% Garwood Poisson interval (Wilson-Hilferty), as measure.py computes it; the per-day rate is kept beside it
compounding rate, primary the weekly growth factor: the exponent of a weighted log-linear fit of ln(episodes/calls) on the week index over weeks with at least one episode and 200 tool calls (weights: episodes), with its 95% interval; measure.py's trend; not estimable under three such weeks, and the same fit per day is the early read until then
baseline to beat ×1.156 per week (×1.135 to ×1.177) at 49.9 per 1,000 in the week of 2026-09-28 (#330, 2026-10-02 comment); exponential improvement is the factor below 1 and staying there; the view shows the factor, its interval and this baseline
structure per intervention, internal merged enforced structure gained on main per week over the operator interventions of that week; structure is session gates (Gate constants of the session gate), accepted miners (directories with rules.dl) and ontology terms (term declarations), counted from git at each week's end (recorded intents live in csfpg and are not counted yet); interventions are the operator-authored messages into sessions, every prompt after a run's first turn, marked proxy until the AFFECT miner labels corrections
compounding exponent, internal the ordinary least-squares slope of log(structure per intervention) on the week index with its 95% interval; not estimable under three weeks; first point: the week of 2026-09-28
$ per ready pull request fixer spend over fixers whose pull request is ready
fixer yield ready pull requests over finished fixers
findings per day findings over the days since the first

Series: ouroboros.struggle_rate_per_1k_tool_calls_weekly (one point per week) and ouroboros.struggle_rate_per_1k_tool_calls (per day), ouroboros.compounding_per_week and ouroboros.compounding_per_day, ouroboros.structure_per_intervention (per week) and ouroboros.structure_exponent, ouroboros.findings, ouroboros.fixer_spend_usd, ouroboros.cost_per_ready_pull_request_usd, ouroboros.fixer_yield. The snapshot is written to <state>/ouroboros.json, which the ops view follows into its panel, and served at /api/ouroboros.

Generic and tenant gates

A miner's scope is decided when it is accepted and carried on every finding (scope in extract.ml, Scope in records.proto, miner.exe scope): a generic gate holds for any harness user (a draft pull request opened late, a sleep loop, a self-matching pgrep); a tenant gate encodes one operator's intents, rulings or vocabulary and never leaves the tenant. The ledger keeps the scope on every finding and the snapshot counts the accepted miners by scope. Today's accepted miners: 1 generic (_template, DRAFT-PR-LATE), 0 tenant. The tenant record, its opt-in field (default off) and the export of generic findings into a shared miner set (after the operator-identifier gate, facts and spans only, instances re-keyed) are planned: CSF has no tenant record yet to carry the field.

The labeler's queue

csf_ouroboros_proposals rows with the verdict needs_labels are the queue the labeler (#318) reads; the record is LabelRequest in contract/records.proto: ticket, proposal, miner, corpus kinds, the instances that named nothing, the reason and when. The labeler never edits the row: it posts a new proposal comment with real instances, and the loop pre-checks that as a new proposal.

Labeler

Labels are the bottleneck. On 2026-10-03, 85 of 107 fixer sessions ($61.90 of $79.37) ended on a ticket that named no real instance, so each invented one. The labeler (labeler/, #318) proposes real ones at no model spend: a local model on the host's otherwise idle GPU, reached through the Ollama provider behind the brain contract.

csf label -repo OWNER/NAME -model-container ollama-1 -model-endpoint http://HOST:11434 \
  -checkout "$PWD" -ticket 105 -ticket 71            # or -tickets-file landed.tsv

One ticket is one call, in four steps:

  1. Parse. The request is a LabelRequest: the loop's queue row (ticket, proposal, miner, corpus, the instances the proposal named, the pre-check's reason) plus the labeler's own reading of the ticket (repository, predicate, facts). A row that carries only the ticket is completed here: the labeler reads the ticket through gh, takes the predicate from its title and body, and the miner, corpus kinds and fact lines from its latest proposal comment, keeping every value the row already carries.
  2. Prefilter, deterministically. The fact lines' quoted literals, snake_case names and flags, and the predicate's backticked spans, are the search terms. Every line of the corpus (<state>/*/events.jsonl through the file capability; issues, pull requests and git log through the process capability) that carries one is a candidate, scored by inverse document frequency, so a term on every line counts for nothing and a rare one for much, and interleaved across instances so one verbose run cannot fill a batch. No threshold is picked.
  3. Label. The model reads a batch of candidates, each with its instance, source line and an excerpt centred on the rarest matched term, and answers a schema-constrained JSON object: for each candidate + or -, a verbatim quote and a reason. The prompt asks for the concrete evidence the predicate describes; shared vocabulary is -.
  4. Accept, with no human in the loop. A positive stands only when it reproduces: through the ticket's miner when one is registered (-miner name=executable; its facts must name the instance), otherwise by its quote being verbatim on the line it names: every fragment of the quote, split at the excerpt's cut mark, on the line in order, read as written or with the JSON line's string escapes decoded. A quote lifted from the ticket's own words fails. Everything else is a ProposedLabel with ACCEPTANCE_REJECTED and why, or ACCEPTANCE_UNCHECKED for a negative.

Every label goes to <state>/labeler.jsonl and the run record LabelerRun to <state>/labeler.json, which the ops view's panel shows live. Both records are typed in contract/records.proto, the shape the LOOP slice and the labeler share.

GPU sharing. Before each batch the labeler lists the running containers and yields while any other than the model server was created with the GPU (Docker's --runtime nvidia, or a device request for the nvidia driver or the gpu capability); it looks again every batch's worth of model time. Every request carries a keep-alive, so the model unloads soon after a run ends: twice the median gap between batch starts, floored at twice the measured load.

Derivations. Every threshold is measured, and the derivation stands beside it in labeler.go:

Threshold Value Measured
context window 16384 tokens 8192 overran: 7674 prompt + 824 answer tokens in one batch
prompt tokens per candidate running max, starts at 265 222 (event lines) to 265 (issue bodies)
answer tokens per label running max, starts at 57 39 to 57
batches per ticket 2 34 to 37 s per batch of 44 to 51 candidates
yield wait 35 s one batch of model time
keep-alive 2 × median gap, floor 2 × load load 4.7 s cold, 1.4 s warm; gaps 37 s

Model choice, by measurement. The host holds one local model, qwen3:8b (the 27B GGUF the exited llama.cpp container names is not on disk), so the choice was between its two configurations, measured on tickets #105, #71 and #199 on 2026-10-04 with the same corpus and prompt:

Configuration Tickets with an instance Positives reproduced Precision Wall time
thinking off 3 of 3 27 of 31 0.87 207 s
thinking on 2 of 3 9 of 10 0.90 249 s

Thinking labels fewer lines positive, so it loses whole tickets for three points of precision and a fifth more time; the labeler runs with it off.

The proof run. On 2026-10-04 the labeler ran over the 85 tickets whose fixer session had invented an instance (landed.tsv, gate column "synthetic labels"), on this host's RTX 5070, at $0:

Measured Value
tickets labeled 85 (2 ended in a model-server error: a 10-minute timeout after a 6225-token answer, and one partial answer)
tickets with at least one reproducible real instance 48
labels proposed 3370
positives proposed / reproduced / rejected 417 / 376 / 41
precision under the acceptance rule 0.90
labels per hour 3620
GPU utilization, mean of a sample after each batch 90% (6.7 GB in use)
wall time 56 min, 136 batches, 1 window overflow
accepted spans by corpus 308 event-log lines, 48 commits, 13 issues, 7 pull requests

Every 21st proposed positive, 20 in all, was read against its source line. All 19 accepted quotes are verbatim on their lines and the one rejected quote is not, so the rule judged 20 of 20 correctly. Only 4 of the 20 lines are an instance of their predicate (a session refusing with the ticket's exact mkdir error, a cold-cache usage record, a list of terms minted after the fact, a brief stating an operator's intent and deadline); the other 15 reproduce a line that carries the ticket's vocabulary: a brief, a tool result, a commit footer. The rule guarantees that a positive is real text at a real span, not that the span is an offense; that judgment waits for the ticket's extractor, which is the rule's first clause. The same rule read strictly (no fragments, no decoded escapes) over the first 11 tickets kept 54 of 88 positives (0.61); the final rule kept 58 of 68 on the same tickets (0.85).

Status

The contract, the template miner, its gates and the loop service are built; the PBO guard and dreaming are planned.

Measured Before this service Now Command
miners that build and run 0 1 ls services/ouroboros/miners
miners with walk-forward FN = 0 0 1 tools/bazel.sh test //services/ouroboros/...
contract and template tests 0 2 pass same
miner proposals from tickets 0 110 first mining run
proposals that can reach TP > 0 on the template extractor — 0 census
tickets of 2026-10-03 the pre-check skips before a model call 0 of 107 105 of 107 csf ouroboros precheck -repository . -ticket ...
loop specs 0 54 pass go test -race ./services/ouroboros/
synthetic-label tickets with a reproducible real instance 0 of 85 48 of 85 the proof run
model spend per labeled ticket $0.73 (fixer sessions) $0 same
First mining run

On 2026-10-03 the orchestrator ran a first pass over the mining tickets. These are proposals, not miners: the orchestrator's miner loop asked one stateless safe-mode Haiku call per ticket batch to write a predicate, facts, rules and labels for each open mining ticket, and posted each as a ticket comment. No backtest has run on any of them.

Measured Value
tickets mined 110
model calls 17
largest batch 11 tickets
artifacts written 110
ticket comments posted 110
refused by the identifier gate 0
cost $1.4033
cost per ticket $0.0128
call seconds per ticket 18.4
input tokens (uncached + cache creation) 170 + 170,878 = 171,048
cache read tokens 0
output tokens 212,278
load1 at call 33.73–278.46
backtested 0

Source: the loop's state/costs.tsv (one row per call: family, tickets, artifacts, input, cache read, cache creation, output tokens, USD, wall seconds, load1), state/posted.tsv (one row per comment) and no state/refused.tsv. Reproduce the sums from the loop's directory:

python3 -c "R=[l.split('\t') for l in open('state/costs.tsv') if l.strip()]; s=lambda i: sum(float(r[i]) for r in R); print(len(R), s(1), max(int(r[1]) for r in R), s(3), s(4), s(5), s(6), round(s(7), 4), round(s(8) / s(1), 1))"

Call seconds per ticket sum each call's wall time; calls ran four at a time, so elapsed time was shorter. Every call wrote its whole prompt into the cache and none read it back. The census on #41 measures these proposals against the contract: 0 of 110 can reach TP > 0 with the template's extractor, because each needs its own.

Reference

The contract

Every record is typed in contract/contract.mli, with its wire form in contract/records.proto; Corpus reads event logs and transcripts in place, gh tickets and pull requests, and files at a revision.

Record Fields
Fact relation, args, span
Finding miner, rule, subject, severity, proof
Backtest labels, split, knee, families, tp, fp, fn, lead
Mutation miner, at, killed, survived, excluded, score, floor, accepted, surviving, seconds

An argument is Text or Number, and a span is a file and a line. Severity runs S0 to S3 (#105): S0 is a gate escape, S3 an offense found only by mining. families counts the $\kappa$ candidates tried, so multiple testing is visible on every backtest.

The gates this service owns
Gate Passes when
miner_test rules safe, labels fire as labeled, FN = 0, TP > 0, the knee binds, backtest.md current, every mutant killed, mutation.md current
tools/close-ticket.sh <ticket> <pr> Backtest FN row is 0 and every check row exits 0

tools/close-ticket.sh --dry-run 211 23 exits 1 with the pull request body has no | Backtest FN | row; without --dry-run it comments that row and leaves the ticket open. Walk-forward acceptance follows Bergmeir and Benítez 2012; the acceptance first written for #390 chose the knee on the record it then judged, and #86 replaced it.

Mutation score

A miner's test is itself tested (#317): Mutate (contract/mutate.ml) changes the rules or the labels one way at a time, reruns the miner's checks on each mutant, and counts.

Mutant kind Change
drop atom one body literal removed
swap comparison gt for ge, lt for le, and back
write knee knee(K) removed, a fixture score written for K
negate atom a positive atom negated
drop rule one clause removed
swap head two verdict relations exchange their rules
flip label one label's sign flipped
move split the two labels either side of the split exchange starts

A mutant the checks fail is killed; one they pass has survived; the score is killed over killed plus survived. Two kinds leave the denominator, each with its reason in mutation.md: a mutant check_rules refuses (an unbound variable after a drop or a negation), and an equivalent mutant, one that derives the same verdicts as the original under no knee and under every knee candidate.

The gate is Mutate.rejections: the score must reach the floor, and no surviving mutant may change a labeled instance's verdict; the survivor is printed with its rules or labels. The floor is derived, not picked: it is the template miner's measured score, recorded as Mutate.template and checked by the template's own test, so it moves only when the template's measurement does.

Measured on the template Value
mutants 16
excluded 6: 1 equivalent (gated(R) dropped: every fixture run is gated), 5 refused
killed 10 of 10
floor 10/10 = 1.00
wall time 4 ms (miner.exe mutate --json, seconds)

Mutation analysis found one blind spot in the template's test as it first shipped: a knee written into the rules, gt(S, 468), passed every check, because the fitted knee was 468 too. The knee-binds check closes it; mutation.md records that it is the only check that kills that mutant. Mutants change rule text and labels only, so they run inside the test's own Bazel sandbox; the extractor runs no further.

The $h_{\mathrm{fleet}}$ floor

$h_{\mathrm{fleet}}$ is the fleet's prompt-cache hit ratio (#235) over every harness result event's usage:

$$h_{\mathrm{fleet}} = \frac{\sum \mathrm{cache_read}}{\sum (\mathrm{input} + \mathrm{cache_creation} + \mathrm{cache_read})}$$

The 0.9867 floor is a stated number, so it carries its reality check (#241):

Stated Measured Runs Runs below Used by
0.9867 0.9870 73 40 none

Measured on 2026-10-03 with:

jq -n '[inputs | select(.event_type == "result") | .event.usage // {}
  | {r: (.cache_read_input_tokens // 0),
     t: ((.input_tokens // 0) + (.cache_creation_input_tokens // 0) + (.cache_read_input_tokens // 0))}]
  | (map(.r) | add) / (map(.t) | add)' <state>/*/events.jsonl

The fleet clears the floor (0.9870 ≥ 0.9867), yet 40 of the 73 runs fall below it on their own ratio $h_r$, so a fleet-wide pass hides a majority of offending runs; #235 ranks those offenders for a miner to drive down. Tonight's 17 proposal calls read 0 cached tokens, so their ratio is 0.

PBO guard (planned)

Walk-forward answers in-sample selection (Bailey, Borwein, López de Prado and Zhu 2014) but not multiple testing. The probability of backtest overfitting (Bailey et al. 2016) splits the record into $S$ blocks; each of the $\binom{S}{S/2}$ halvings $c$ picks the best candidate in sample, with relative rank $\omega_c$ out of sample:

$$\mathrm{PBO} = \Pr_c\left[\lambda_c \le 0\right], \qquad \lambda_c = \ln\frac{\omega_c}{1 - \omega_c}$$

Record Labels After $t$ families PBO
template fixtures 3 1 3 undefined
every recorded run 3 1 60 undefined

With $S = 2$ each half holds one or two labels, too few to rank candidates; the guard waits for a label set with a positive in each of $S \ge 4$ blocks.

Dreaming (planned)

Dreaming is the loop pointed at the language and the queue rather than at one gate (#140): a dreamer classifies operator intents and proposes terms, gates and ranked work, and an actor executes turns and never edits the language. Its objective is minimum description length (Rissanen 1978), choosing the language $L$ that minimizes $|L| + \sum_{i} |\mathrm{encode}(i \mid L)|$. It runs at two speeds, after complementary learning systems (McClelland, McNaughton and O'Reilly 1995): the nightly dream ratifies, and online dreaming (#141) mints provisional session gates that expire unless ratified.

  • Example (build it, #141): swap the gates at a tool-call boundary; the hook binary reads its rules on every call, so a rule published between two calls binds from the next one.
  • Counterexample (dropped, #141): swap code into the Go process during a GC pause; Go has no JIT, and a plugin can never be unloaded.

Neither speed has code here yet; a miner's findings are the dream's input once it lands.

Documentation

Overview

Package ouroboros is the mining loop: the service that runs every accepted miner over the harness corpus all the time, decides for free whether a mining ticket can be fixed, launches one fixer session per ticket through the harness under a daily budget, merges ready pull requests through the merge path, and measures the compounding number the loop exists for. It replaces the ad hoc scripts that launched fixers outside the session gates.

It is a service in the harness host app: it owns no listener and no process. The binary grants it the ledger over csfpg, the harness's session capability, the process capability, the ticket capability over gh, the clock, the file and watch capabilities over the corpus and the repository, and mounts its four triggers into the cron service. Its own goroutine is the corpus watch, started on the scope the runtime hands it.

Index

Constants

View Source
const (
	// TriggerDetector sweeps every miner over the corpus items that grew
	// since they were last mined; the watch runs the same sweep the moment a
	// log changes, so this is the catch-up after a restart or a missed
	// notification.
	TriggerDetector = "ouroboros.detector"
	// TriggerFixers reconciles the running fixer sessions, pre-checks every
	// open mining ticket and launches fixers under the daily budget.
	TriggerFixers = "ouroboros.fixers"
	// TriggerMergeTrain runs the merge path over every ready pull request,
	// oldest first; one run of the path builds the checkers and takes about
	// ten minutes, so the train runs every thirty.
	TriggerMergeTrain = "ouroboros.merge_train"
	// TriggerMeasure computes the struggle rate, its compounding rate and the
	// loop's own yields, records them as series and projects them for the
	// ops view.
	TriggerMeasure = "ouroboros.measure"
	// ScheduleLocation is the time zone the triggers, the budget day and the
	// measured days are declared in: the one #330's measurement uses.
	ScheduleLocation = "America/Los_Angeles"

	// SnapshotFile is where the loop projects its numbers, under the state
	// directory, for the ops view that follows that directory.
	SnapshotFile = "ouroboros.json"
	// DefaultDailyBudgetUSDMicros is the operator's hard cap on fixer spend:
	// $100 per day, in integer microdollars like every amount in csfpg.
	DefaultDailyBudgetUSDMicros = 100_000_000
	// DefaultFixerModel is the model fixers run on: every real session runs
	// claude-opus-5-5 (operator, 2026-10-05), which the harness's model
	// policy enforces at submit.
	DefaultFixerModel = "claude-opus-5-5"
	// MinersDirectory holds one directory per miner under the repository.
	MinersDirectory = "services/ouroboros/miners"
	// MinerExecutables holds each miner's built executable, as Bazel publishes
	// it under the repository.
	MinerExecutables = "bazel-bin/services/ouroboros/miners"

	// APIPath serves the latest snapshot as JSON on the host's router.
	APIPath = "/api/ouroboros"
)

The triggers the loop declares, as csfcron names them, and their cadence. Each interval is derived from a measurement recorded in the README: the detector sweep is the catch-up behind the watch, the fixer and merge cadences follow the mean fixer session (21.8 min), and the measure follows the day the compounding number is computed over.

View Source
const (

	// ScopeGeneric and ScopeTenant are the proto names of a miner's scope
	// (#120 amendment C): generic holds for any harness user and may be
	// shared from an opted-in tenant; tenant never leaves the tenant.
	ScopeGeneric = "SCOPE_GENERIC"
	ScopeTenant  = "SCOPE_TENANT"
)

A miner's files under its directory and the verbs of its executable, as the contract's main defines them.

View Source
const (
	SeriesStruggleRate       = "ouroboros.struggle_rate_per_1k_tool_calls"
	SeriesStruggleRateWeekly = "ouroboros.struggle_rate_per_1k_tool_calls_weekly"
	SeriesCompounding        = "ouroboros.compounding_per_week"
	SeriesCompoundingDaily   = "ouroboros.compounding_per_day"
	SeriesFindings           = "ouroboros.findings"
	SeriesFixerSpend         = "ouroboros.fixer_spend_usd"
	SeriesCostPerReadyPull   = "ouroboros.cost_per_ready_pull_request_usd"
	SeriesFixerYield         = "ouroboros.fixer_yield"

	// The baseline to beat (#330, 2026-10-02 comment): the weekly growth
	// factor of the struggle rate with its interval, and the rate in the
	// week of 2026-09-28. Exponential improvement is this factor below 1,
	// and staying there.
	BaselineWeeklyFactor = 1.156
	BaselineWeeklyLow    = 1.135
	BaselineWeeklyHigh   = 1.177
	BaselineRatePerK     = 49.9
	BaselineWeek         = "2026-09-28"

	// SignalToolError, SignalPermissionDenial and SignalRetry name the three
	// struggle signals read from a harness event log, as rrsi-mine names them.
	SignalToolError        = "tool_error"
	SignalPermissionDenial = "permission_denial"
	SignalRetry            = "retry"
)

The series the measure records, and the constants of the measurement. The primary series is candace-server #330's: struggle episodes per 1,000 agent tool calls, reported weekly with the 95% Garwood Poisson interval, and the weekly growth factor from a weighted log-linear fit over the active weeks. The struggle definition is the deterministic part of #330's (rrsi-mine traces, the nine signals merged into episodes): the three signals a harness event log carries without a model are kept, with the same episode gap, and every episode counts, since no labeler marks noise or harness-fixability here. Days are calendar days in ScheduleLocation; weeks start on Monday. The per-day fit is kept as the early read while the corpus holds fewer than three active weeks.

View Source
const (
	SeriesStructure         = "ouroboros.structure_per_intervention"
	SeriesStructureExponent = "ouroboros.structure_exponent"
	// StructureProxy marks the denominator as the proxy it is until the
	// AFFECT miner labels corrections.
	StructureProxy = "proxy: operator-authored messages stand in for interventions until the AFFECT miner labels corrections"
)

The internal check (#120 amendment B): structure gained per operator intervention, per week. The numerator is merged enforced structure, counted from git on main at each week's end: session gates (the Gate constants of the session gate), accepted miners (directories with a rules file) and ontology terms (term declarations). Recorded intents live in csfpg, not in git, so they are not counted here. The denominator is operator interventions: until the AFFECT miner labels corrections, the operator-authored messages into sessions (every prompt after a run's first turn), so the series is marked proxy. The slope of log(y) on the week index is the compounding exponent.

View Source
const (
	CorpusEvents      = "jsonl"
	CorpusIssue       = "issue"
	CorpusPullRequest = "pull_request"
	CorpusFile        = "file_at"

	InstanceRun         = "run"
	InstanceIssue       = "issue"
	InstancePullRequest = "pull_request"
	InstanceCommit      = "commit"
)

The corpus kinds a proposal declares, as the miner contract's Corpus readers name them, and the instance kind each one reads.

View Source
const (
	// FixerAgentID names the fixer agent in every recipe and run record.
	FixerAgentID = "ouroboros-fixer"
)

The fixer recipe: one agent definition, one brief per ticket, one branch per miner, and the tools the brief needs. The brief is the one the ad hoc loop handed its sessions, made a harness recipe.

View Source
const (

	// SessionTrailer is the commit trailer every harness session leaves on
	// its commits; the merge train reads the author from it.
	SessionTrailer = "CSF-Session:"
)

The gh invocations and the labels the mining subset is read by.

Variables

View Source
var (
	// ErrInvalidOption reports a nil option or a value the service cannot use.
	ErrInvalidOption = errors.New("ouroboros: invalid option")
	// ErrMissingCapability reports a loop built without a required capability.
	ErrMissingCapability = errors.New("ouroboros: a required capability is missing")
	// ErrAlreadyStarted reports a second Start of one loop.
	ErrAlreadyStarted = errors.New("ouroboros: the loop is already started")
	// ErrBudgetSpent reports a launch refused because the day's budget is
	// spent or reserved by the fixers already running.
	ErrBudgetSpent = errors.New("ouroboros: the daily fixer budget is spent")
	// ErrFixersDisabled reports a launch refused because the loop was built
	// without the fixer switch.
	ErrFixersDisabled = errors.New("ouroboros: fixers are not enabled")
	// ErrSelfMerge reports a pull request the train refused to merge because
	// its commits carry the merger's own session: NO-SELF-MERGE (#249).
	ErrSelfMerge = errors.New("ouroboros: the merger is the author session")
)
View Source
var ErrDatabaseRequired = errors.New("ouroboros: a database is required")

ErrDatabaseRequired reports a ledger built without the database capability.

View Source
var ErrNoRepositorySlug = errors.New("ouroboros: the origin URL names no GitHub repository")

ErrNoRepositorySlug reports an origin URL no GitHub owner/name can be read from.

View Source
var ErrNotPullRequestURL = errors.New("ouroboros: not a GitHub pull request URL")

ErrNotPullRequestURL reports a URL that names no GitHub pull request.

Functions

func CorpusKinds

func CorpusKinds(declared string) []string

CorpusKinds reads the corpus kinds a proposal declares, such as "pull_request, jsonl".

func MinerName

func MinerName(proposed string) string

MinerName is the Bazel package name a proposal's miner gets.

func PoissonRate

func PoissonRate(events int64, exposure int64) (rate *float64, low *float64, high *float64)

PoissonRate is events per 1,000 units of exposure with its 95% Garwood interval; no rate when there is no exposure.

func PullRequestOf

func PullRequestOf(url string) (string, int64, error)

PullRequestOf reads the repository slug and the number of the pull request a GitHub URL names.

func Refusal

func Refusal(report string) string

Refusal is the refusal reason a merge report carries: its first lines naming a refusal, regression, drift or failure, bounded.

func RepositorySlug

func RepositorySlug(origin string) (string, error)

RepositorySlug reads owner/name from a GitHub remote URL in either the https or the ssh spelling.

func SessionTrailers

func SessionTrailers(text string) []string

SessionTrailers reads every distinct CSF-Session trailer value in text, in first-seen order.

func WeekOf

func WeekOf(day string) string

WeekOf is the Monday of the week holding a calendar day.

Types

type Baseline

type Baseline struct {
	Factor   float64 `json:"factor"`
	Low      float64 `json:"low"`
	High     float64 `json:"high"`
	RatePerK float64 `json:"rate_per_1k"`
	Week     string  `json:"week"`
}

Baseline is #330's baseline, carried in every snapshot so the view shows what the factor is measured against.

type Comment

type Comment struct {
	ID        string
	Body      string
	CreatedAt time.Time
}

Comment is one comment on a ticket.

type DailyBudget

type DailyBudget struct {
	Day                      string
	CapUSDMicros             int64
	SpentUSDMicros           int64
	ReservedUSDMicros        int64
	ExpectedSessionUSDMicros int64
	MeasuredSessions         int64
}

DailyBudget is the day's fixer budget as the ledger reads it now: the cap, what the day's fixers spent, what the running ones are expected to add, and that expectation per session with the number of finished sessions it was measured from (zero when it is the measured baseline).

func (DailyBudget) LeftUSDMicros

func (budget DailyBudget) LeftUSDMicros() int64

LeftUSDMicros is the cap less the spend and the reservation, never below zero.

type Exponent

type Exponent struct {
	Slope *float64 `json:"slope"`
	Low   *float64 `json:"low"`
	High  *float64 `json:"high"`
	Weeks int      `json:"weeks"`
	Proxy string   `json:"proxy"`
}

Exponent is the compounding exponent of the internal check: the slope of log(structure per intervention) on the week index, with its 95% interval.

func StructureExponent

func StructureExponent(weeks []StructureWeek) Exponent

StructureExponent fits log(y) on the week index by ordinary least squares over the weeks with a positive y, and returns the slope with its 95% interval; it needs structureMinWeeks such weeks.

type Finding

type Finding struct {
	Miner    string
	Rule     string
	Subject  string
	Item     string
	Severity string
	Scope    string
	Record   json.RawMessage
	FoundAt  time.Time
}

Finding is one verdict row a miner derived over one item: the subject's instance, the rule, the severity and the miner's own record whole.

func ParseFindings

func ParseFindings(miner string, items []string, output []byte, at time.Time) ([]Finding, error)

ParseFindings reads the JSON lines a miner printed. A finding's subject is its first argument; it is attributed to the item whose run directory the subject names, or to the first item when the subject names none.

type Fixer

type Fixer struct {
	AssignmentID   string
	Ticket         int64
	ProposalID     string
	Miner          string
	Model          string
	StartedAt      time.Time
	FinishedAt     time.Time
	Outcome        Outcome
	PullRequestURL string
	CostUSDMicros  int64
	InputTokens    int64
	OutputTokens   int64
	Seconds        int64
}

Fixer is one fixer session the loop launched, with what it cost.

type FixerSummary

type FixerSummary struct {
	Sessions      int      `json:"sessions"`
	Running       int      `json:"running"`
	Finished      int      `json:"finished"`
	Ready         int      `json:"ready"`
	SpendUSD      float64  `json:"spend_usd"`
	SpendTodayUSD float64  `json:"spend_today_usd"`
	BudgetUSD     float64  `json:"budget_usd"`
	CostPerReady  *float64 `json:"cost_per_ready_pull_request_usd"`
	Yield         *float64 `json:"yield"`
	Launch        bool     `json:"launch_enabled"`
}

FixerSummary is the loop's own yield, from the ledger.

type GitHubTickets

type GitHubTickets struct {
	// contains filtered or unexported fields
}

GitHubTickets is the ticket capability over gh, run through the process capability against one repository.

func NewGitHubTickets

func NewGitHubTickets(launcher proc.ILauncher, repository string) (*GitHubTickets, error)

NewGitHubTickets grants the repository's tickets through gh. repository is the owner/name slug.

func (*GitHubTickets) Comment

func (tickets *GitHubTickets) Comment(ctx context.Context, number int64, body string) error

Comment posts one comment on a pull request, read from standard input so its body never crosses an argument vector.

func (*GitHubTickets) ListMiningTickets

func (tickets *GitHubTickets) ListMiningTickets(ctx context.Context) ([]Ticket, error)

ListMiningTickets lists the open severity-labeled, non-lang issues.

func (*GitHubTickets) ListReadyPullRequests

func (tickets *GitHubTickets) ListReadyPullRequests(ctx context.Context) ([]PullRequest, error)

ListReadyPullRequests lists the open pull requests that are not drafts, oldest first.

func (*GitHubTickets) PullRequest

func (tickets *GitHubTickets) PullRequest(ctx context.Context, number int64) (PullRequest, error)

PullRequest reads one pull request.

func (*GitHubTickets) PullRequestSessions

func (tickets *GitHubTickets) PullRequestSessions(ctx context.Context, number int64) ([]string, error)

PullRequestSessions reads the CSF-Session trailers of a pull request's commit messages.

func (*GitHubTickets) Reference

func (tickets *GitHubTickets) Reference(ctx context.Context, number int64) (ReferenceKind, error)

Reference asks the issues API what a number is; a number gh cannot find is nothing.

func (*GitHubTickets) Repository

func (tickets *GitHubTickets) Repository() string

Repository is the owner/name slug.

func (*GitHubTickets) TicketComments

func (tickets *GitHubTickets) TicketComments(ctx context.Context, number int64) ([]Comment, error)

TicketComments lists a ticket's comments, oldest first.

type HiddenRuns

type HiddenRuns func(ctx context.Context) (map[string]bool, error)

HiddenRuns lists the run directories of the corpus no miner and no measure may read: the evaluation suite's tickets, held out (#416).

type IStore

type IStore interface {
	Item(ctx context.Context, miner string, item string) (Item, bool, error)
	RecordItem(ctx context.Context, item Item) error
	ItemsMinedSince(ctx context.Context, since time.Time) ([]Item, error)
	RecordFindings(ctx context.Context, findings []Finding) (int64, error)
	FindingsSince(ctx context.Context, since time.Time) ([]Finding, error)
	RecentFindings(ctx context.Context, limit int) ([]Finding, error)
	Proposal(ctx context.Context, ticket int64, proposal string) (ProposalRecord, bool, error)
	RecordProposal(ctx context.Context, record ProposalRecord) error
	ProposalsNeedingLabels(ctx context.Context) ([]ProposalRecord, error)
	RecordFixer(ctx context.Context, fixer Fixer) error
	UpdateFixer(ctx context.Context, fixer Fixer) error
	RunningFixers(ctx context.Context) ([]Fixer, error)
	FixersStartedSince(ctx context.Context, since time.Time) ([]Fixer, error)
	FixersForTicket(ctx context.Context, ticket int64) ([]Fixer, error)
	RecentFixers(ctx context.Context, limit int) ([]Fixer, error)
	Merge(ctx context.Context, pullRequest int64, headSHA string) (Merge, bool, error)
	RecordMerge(ctx context.Context, merge Merge) error
	RecentMerges(ctx context.Context, limit int) ([]Merge, error)
	RecordSeries(ctx context.Context, point Point) error
	Series(ctx context.Context, series string) ([]Point, error)
}

IStore is the ledger: every proposal, pre-check verdict, fixer session, merge result and series point the loop records, in the out-of-process store the binary grants.

type ITickets

type ITickets interface {
	// Repository is the owner/name slug the tickets live in.
	Repository() string
	// ListMiningTickets lists the open tickets carrying a severity label
	// and no lang label, the mining subset, lowest number first.
	ListMiningTickets(ctx context.Context) ([]Ticket, error)
	// TicketComments lists a ticket's comments, oldest first.
	TicketComments(ctx context.Context, number int64) ([]Comment, error)
	// Reference reports whether a number is an issue, a pull request or
	// nothing in the repository.
	Reference(ctx context.Context, number int64) (ReferenceKind, error)
	// ListReadyPullRequests lists the open pull requests that are not drafts,
	// oldest first.
	ListReadyPullRequests(ctx context.Context) ([]PullRequest, error)
	// PullRequest reads one pull request.
	PullRequest(ctx context.Context, number int64) (PullRequest, error)
	// PullRequestSessions lists the CSF-Session trailers of a pull request's
	// commits: the sessions that authored it.
	PullRequestSessions(ctx context.Context, number int64) ([]string, error)
	// Comment posts one comment on a pull request.
	Comment(ctx context.Context, number int64, body string) error
}

ITickets is the ticket capability: the repository's issues and pull requests, reached through gh in the binary and a double in specs.

type Item

type Item struct {
	Miner    string
	Item     string
	ByteSize int64
	MinedAt  time.Time
}

Item is one corpus file a miner has read, at the size it read it.

type Ledger

type Ledger struct {
	// contains filtered or unexported fields
}

Ledger is the IStore over CSF's PostgreSQL schema, reached through the csfpg capability: the csf_ouroboros_* tables and the queries csfpg generates for them. It borrows the capability and never closes it.

func NewLedger

func NewLedger(database csfpg.IDB) (*Ledger, error)

NewLedger returns the ledger over a pool the binary opened through ipc/db/csfpg, or over pgmem's IDB in a spec.

func (*Ledger) FindingsSince

func (ledger *Ledger) FindingsSince(ctx context.Context, since time.Time) ([]Finding, error)

FindingsSince lists the findings recorded at or after since, oldest first; a zero since lists every finding.

func (*Ledger) FixersForTicket

func (ledger *Ledger) FixersForTicket(ctx context.Context, ticket int64) ([]Fixer, error)

FixersForTicket lists every fixer launched for one ticket, oldest first.

func (*Ledger) FixersStartedSince

func (ledger *Ledger) FixersStartedSince(ctx context.Context, since time.Time) ([]Fixer, error)

FixersStartedSince lists the fixers started at or after since, oldest first; a zero since lists every fixer.

func (*Ledger) Item

func (ledger *Ledger) Item(ctx context.Context, miner string, item string) (Item, bool, error)

Item reads where one miner stands on one corpus item.

func (*Ledger) ItemsMinedSince

func (ledger *Ledger) ItemsMinedSince(ctx context.Context, since time.Time) ([]Item, error)

ItemsMinedSince lists every item a miner last read at or after since, oldest first.

func (*Ledger) Merge

func (ledger *Ledger) Merge(ctx context.Context, pullRequest int64, headSHA string) (Merge, bool, error)

Merge reads the merge path's result on one pull request head.

func (*Ledger) Proposal

func (ledger *Ledger) Proposal(ctx context.Context, ticket int64, proposal string) (ProposalRecord, bool, error)

Proposal reads the verdict recorded on one proposal of one ticket.

func (*Ledger) ProposalsNeedingLabels

func (ledger *Ledger) ProposalsNeedingLabels(ctx context.Context) ([]ProposalRecord, error)

ProposalsNeedingLabels is the labeler's queue, oldest first.

func (*Ledger) RecentFindings

func (ledger *Ledger) RecentFindings(ctx context.Context, limit int) ([]Finding, error)

RecentFindings lists the newest findings, newest first.

func (*Ledger) RecentFixers

func (ledger *Ledger) RecentFixers(ctx context.Context, limit int) ([]Fixer, error)

RecentFixers lists the newest fixers, newest first.

func (*Ledger) RecentMerges

func (ledger *Ledger) RecentMerges(ctx context.Context, limit int) ([]Merge, error)

RecentMerges lists the newest merge results, newest first.

func (*Ledger) RecordFindings

func (ledger *Ledger) RecordFindings(ctx context.Context, findings []Finding) (int64, error)

RecordFindings records the findings not yet recorded and reports how many were new.

func (*Ledger) RecordFixer

func (ledger *Ledger) RecordFixer(ctx context.Context, fixer Fixer) error

RecordFixer records a fixer session the loop launched, as running.

func (*Ledger) RecordItem

func (ledger *Ledger) RecordItem(ctx context.Context, item Item) error

RecordItem records that a miner read an item at a size.

func (*Ledger) RecordMerge

func (ledger *Ledger) RecordMerge(ctx context.Context, merge Merge) error

RecordMerge records one run of the merge path.

func (*Ledger) RecordProposal

func (ledger *Ledger) RecordProposal(ctx context.Context, record ProposalRecord) error

RecordProposal records or replaces the verdict on one proposal.

func (*Ledger) RecordSeries

func (ledger *Ledger) RecordSeries(ctx context.Context, point Point) error

RecordSeries records or replaces one day of one series.

func (*Ledger) RunningFixers

func (ledger *Ledger) RunningFixers(ctx context.Context) ([]Fixer, error)

RunningFixers lists the fixers not yet finished, oldest first.

func (*Ledger) Series

func (ledger *Ledger) Series(ctx context.Context, series string) ([]Point, error)

Series lists one series, oldest day first.

func (*Ledger) UpdateFixer

func (ledger *Ledger) UpdateFixer(ctx context.Context, fixer Fixer) error

UpdateFixer records what a fixer session cost and where it stands.

type Loop

type Loop struct {
	// contains filtered or unexported fields
}

Loop is the mining loop. Its state between occurrences is the ledger; the only in-process datum is the latest snapshot, published whole.

func NewLoop

func NewLoop(options ...LoopOption) (*Loop, error)

NewLoop validates the whole option set and returns a stopped loop: it reads nothing and starts no goroutine until mounted.

func (*Loop) DailyBudget

func (loop *Loop) DailyBudget(ctx context.Context) (DailyBudget, error)

DailyBudget reads the day's budget from the ledger. The expectation per running fixer is the ledger's own mean once it has ledgerMeanMinimum finished fixers, the measured baseline before.

func (*Loop) Detect

func (loop *Loop) Detect(ctx context.Context, only []string) error

Detect runs every built miner over the corpus items that grew since the miner last read them: every item when only is nil, the named run directories otherwise. It is deterministic and calls no model; the new findings go to the ledger and each item's size is recorded as read.

func (*Loop) Fixers

func (loop *Loop) Fixers(ctx context.Context) error

Fixers is one occurrence of the fixer trigger: reconcile the running fixers with the harness, then pre-check every open mining ticket whose latest proposal has no verdict yet and launch a fixer for each one that passes, one per ticket, while the day's budget and the harness's admission allow. A ticket whose proposal names no real instance is queued for the labeler at no model cost.

func (*Loop) Measure

func (loop *Loop) Measure(ctx context.Context) (*Snapshot, error)

Measure computes the loop's numbers over the corpus and the ledger, records them as series and publishes the snapshot.

func (*Loop) MergeTrain

func (loop *Loop) MergeTrain(ctx context.Context) error

MergeTrain runs the merge path over every ready pull request, oldest first, exactly once per head: a refused head is commented on and retried only after its author pushes a new one. The merger is never an author session: a pull request whose commits carry the merger's identity is refused without running the path.

func (*Loop) Miners

func (loop *Loop) Miners(ctx context.Context) ([]Miner, error)

Miners lists the miners under the repository: every directory with a rules file, built or not, with the scope a built one reports.

func (*Loop) Register

func (loop *Loop) Register(router gin.IRouter)

Register mounts the snapshot route on the caller's router.

func (*Loop) Start

func (loop *Loop) Start(scope *runtime.Scope) error

Start starts the corpus watch on scope, so every change to a run's event log mines that log once the change has settled, and runs the detector and the measure once, so the numbers are there as soon as the host is.

func (*Loop) Triggers

func (loop *Loop) Triggers() []cronservice.Option

Triggers declares the loop's four triggers for the cron service, in ScheduleLocation. Every one skips an occurrence while the previous one still runs, so a long merge never overlaps the next.

type LoopOption

type LoopOption func(loop *Loop) error

LoopOption configures a Loop.

func WithClock

func WithClock(source clock.IClock) LoopOption

WithClock replaces the host's clock, which stamps every ledger row and decides the budget day.

func WithCorpus

func WithCorpus(directory string, files iofs.IFiles, watcher iofs.IWatcher) LoopOption

WithCorpus grants the harness state directory the miners read: its absolute path, read access and change notification. Required.

func WithDailyBudget

func WithDailyBudget(usdMicros int64) LoopOption

WithDailyBudget sets the hard cap on fixer spend per day, in USD microdollars. The default is DefaultDailyBudgetUSDMicros.

func WithFixerLaunch

func WithFixerLaunch(enabled bool) LoopOption

WithFixerLaunch is the always-on switch: fixers are launched only when it is on. Until the harness reports that sandboxed launch is available, the binary turns it on only from an explicit flag.

func WithFixerModel

func WithFixerModel(model string) LoopOption

WithFixerModel sets the model fixers run on.

func WithHidden

func WithHidden(hidden HiddenRuns) LoopOption

WithHidden grants the list of held-out run directories; every corpus reader of the loop skips them. Without it nothing is hidden.

func WithLauncher

func WithLauncher(launcher proc.ILauncher) LoopOption

WithLauncher grants the process capability: miners, git and the merge path run through it. Required.

func WithLedger

func WithLedger(store IStore) LoopOption

WithLedger grants the ledger. Required.

func WithLogger

func WithLogger(logger *slog.Logger) LoopOption

WithLogger receives the loop's own records.

func WithMergeObserver

func WithMergeObserver(observer MergeObserver) LoopOption

WithMergeObserver tells observer about every pull request the merge train merges: the slice dispatcher releases what depended on it.

func WithMergerIdentity

func WithMergerIdentity(identity string) LoopOption

WithMergerIdentity names the principal the merge train merges as, which NO-SELF-MERGE compares with the sessions a pull request's commits carry. Required.

func WithRepository

func WithRepository(directory string, files iofs.IFiles) LoopOption

WithRepository grants the checkout the miners are read from, fixer worktrees are created from and the merge path runs in. Required.

func WithSessions

func WithSessions(sessions csf.IAgentSessions) LoopOption

WithSessions grants the harness's session capability, through which every fixer is launched and observed. Required.

func WithState

func WithState(directory string, files iofs.IFiles) LoopOption

WithState grants the state directory of the harness host the loop runs in: where fixer sessions run and where the snapshot is written. Required.

func WithTickets

func WithTickets(tickets ITickets) LoopOption

WithTickets grants the ticket capability over the repository's issues and pull requests. Required.

type Merge

type Merge struct {
	PullRequest    int64
	HeadSHA        string
	AuthorSessions []string
	Merger         string
	Merged         bool
	ExitCode       int
	Reason         string
	RecordedAt     time.Time
}

Merge is one run of the merge path over one pull request head.

type MergeObserver

type MergeObserver func(ctx context.Context, pull PullRequest)

MergeObserver is told about every pull request the merge train merged, once the merge is recorded.

type MergeSummary

type MergeSummary struct {
	Merged  int `json:"merged"`
	Refused int `json:"refused"`
}

MergeSummary is the merge train's record.

type Miner

type Miner struct {
	Name       string
	Directory  string
	Executable string
	Built      bool
	// Scope is what the miner's findings encode, as its executable reports
	// it; empty for a miner that is not built.
	Scope string
}

Miner is one accepted miner: a directory under the repository's miners directory with its rules, and its built executable.

type MinerStatus

type MinerStatus struct {
	Name     string `json:"name"`
	Built    bool   `json:"built"`
	Scope    string `json:"scope"`
	Knee     int64  `json:"knee"`
	HasKnee  bool   `json:"has_knee"`
	Findings int    `json:"findings"`
}

MinerStatus is what the snapshot shows about one miner.

type Outcome

type Outcome string

Outcome is where a fixer session stands.

const (
	OutcomeRunning       Outcome = "running"
	OutcomeReady         Outcome = "ready"
	OutcomeDraft         Outcome = "draft"
	OutcomeNoPullRequest Outcome = "no_pull_request"
	OutcomeFailed        Outcome = "failed"
	OutcomeCanceled      Outcome = "canceled"
)

The outcomes, as the ledger stores them.

func (Outcome) Finished

func (outcome Outcome) Finished() bool

Finished reports whether the outcome is terminal.

type Period

type Period int

Period is the unit a compounding fit indexes time by.

const (
	PerDay Period = iota
	PerWeek
)

The periods the fit runs over.

type Point

type Point struct {
	Series      string
	Day         string
	Value       *float64
	Low         *float64
	High        *float64
	Numerator   int64
	Denominator int64
	ComputedAt  time.Time
}

Point is one day of one series: a value with its interval, or no value when the day had no exposure, and the counts it was computed from.

type Precheck

type Precheck struct {
	Verdict Verdict
	Reason  string
	Found   []Reference
}

Precheck is the pre-check's result on one proposal.

type Prechecker

type Prechecker struct {
	// contains filtered or unexported fields
}

Prechecker is the free pre-check: it decides, without a model, whether a proposal's miner has anything to backtest on. It reads the corpus, asks the ticket capability what a number is and git whether a commit exists.

func NewPrechecker

func NewPrechecker(corpus iofs.IFiles, tickets ITickets, launcher proc.ILauncher, repository string) (*Prechecker, error)

NewPrechecker grants the pre-check the corpus, the tickets, the process capability and the repository checkout commits are resolved in.

func (*Prechecker) Check

func (checker *Prechecker) Check(ctx context.Context, ticket int64, proposal Proposal) (Precheck, error)

Check decides whether the proposal's miner has anything to backtest on: at least one labeled instance must exist in the corpus and be of a kind the miner's extractor reads. A run must have a run directory, an issue or pull request must exist in the repository and not be the ticket itself, and a commit must be in the repository.

type Proposal

type Proposal struct {
	ID        string
	Miner     string
	Corpus    []string
	Instances []string
}

Proposal is one mining proposal as a ticket carries it: the miner it names, the corpus kinds its extractor reads and the instances it labels.

func LatestProposal

func LatestProposal(comments []Comment) (Proposal, bool)

LatestProposal is the newest proposal among a ticket's comments.

func ParseProposal

func ParseProposal(comment Comment) (Proposal, bool)

ParseProposal reads a proposal from a comment body: the miner and corpus rows of its field table and every row of its labeled-instance table. A comment without the labeled-instance table is not a proposal.

type ProposalRecord

type ProposalRecord struct {
	Ticket    int64
	ID        string
	Miner     string
	Corpus    string
	Instances []string
	Verdict   Verdict
	Reason    string
	CheckedAt time.Time
}

ProposalRecord is the pre-check's verdict on one mining proposal of one ticket. A verdict of needs_labels is the labeler's queue entry.

type PullRequest

type PullRequest struct {
	Number    int64
	Title     string
	URL       string
	HeadSHA   string
	Draft     bool
	State     string
	CreatedAt time.Time
}

PullRequest is one pull request as the merge train and the fixer ledger read it.

func (PullRequest) Merged

func (pull PullRequest) Merged() bool

Merged reports whether the pull request has merged.

type Rate

type Rate struct {
	Day       string   `json:"day"`
	ToolCalls int64    `json:"tool_calls"`
	Struggles int64    `json:"struggles"`
	PerK      *float64 `json:"per_1k"`
	Low       *float64 `json:"low"`
	High      *float64 `json:"high"`
}

Rate is one period's struggle rate: episodes per 1,000 tool calls with its interval, or no rate when the period had no tool call. Day is the calendar day, or the Monday of the week.

func DailyRates

func DailyRates(readings []Reading) ([]Rate, map[string]int64)

DailyRates folds the readings of several event logs into one rate per day, oldest first, and the operator interventions per week.

func Weekly

func Weekly(days []Rate) []Rate

Weekly folds daily rates into one rate per week, oldest first.

type Reading

type Reading struct {
	Calls         map[string]int64
	Episodes      map[string]int64
	Interventions map[string]int64
}

Reading is what one event log contributes to the measure, each count by the calendar day it fell on.

func Struggles

func Struggles(lines []byte, location *time.Location) Reading

Struggles reads one event log into its tool calls, struggle episodes and operator interventions, each dated in location.

func StrugglesWithin

func StrugglesWithin(lines []byte, location *time.Location, budget int) Reading

StrugglesWithin is Struggles over the log's first budget tool calls, and everything up to the next one: a fixed budget, the unit the evaluation suite reads a replay at. A budget of 0 reads the whole log.

type Reference

type Reference struct {
	Kind  string
	Value string
}

Reference is one thing a labeled instance may name in the corpus.

func References

func References(instance string) []Reference

References reads what a labeled instance may name: harness runs by their assignment identifier, issues and pull requests by a #number, and commits by a hex revision. It is pure; existence is the pre-check's question.

type ReferenceKind

type ReferenceKind string

ReferenceKind is what a number names in the repository's issue space.

const (
	ReferenceIssue       ReferenceKind = "issue"
	ReferencePullRequest ReferenceKind = "pull_request"
	ReferenceNone        ReferenceKind = "none"
)

The kinds a number resolves to.

type Snapshot

type Snapshot struct {
	ComputedAt        time.Time       `json:"computed_at"`
	Day               string          `json:"day"`
	Week              Rate            `json:"week"`
	Weeks             []Rate          `json:"weeks"`
	Compounding       Trend           `json:"compounding"`
	Baseline          Baseline        `json:"baseline"`
	Struggle          Rate            `json:"struggle_rate"`
	Days              []Rate          `json:"days"`
	CompoundingDaily  Trend           `json:"compounding_daily"`
	Structure         []StructureWeek `json:"structure"`
	StructureExponent Exponent        `json:"structure_exponent"`
	Findings          int             `json:"findings"`
	FindingsToday     int             `json:"findings_today"`
	FindingsPerDay    *float64        `json:"findings_per_day"`
	Fixers            FixerSummary    `json:"fixers"`
	Merges            MergeSummary    `json:"merges"`
	Queue             int             `json:"needs_labels"`
	Miners            []MinerStatus   `json:"miners"`
	GenericMiners     int             `json:"generic_miners"`
	TenantMiners      int             `json:"tenant_miners"`
}

Snapshot is the loop's numbers at one instant: what the ops view shows and the API serves.

type StructureCounts

type StructureCounts struct {
	Gates  int `json:"gates"`
	Miners int `json:"miners"`
	Terms  int `json:"terms"`
}

StructureCounts is the enforced structure on main at one commit.

func (StructureCounts) Total

func (counts StructureCounts) Total() int

Total is the structure as one number.

type StructureWeek

type StructureWeek struct {
	Week            string          `json:"week"`
	Commit          string          `json:"commit"`
	Counts          StructureCounts `json:"counts"`
	Gained          int             `json:"gained"`
	Interventions   int64           `json:"interventions"`
	PerIntervention *float64        `json:"per_intervention"`
}

StructureWeek is one week of the internal check.

type Ticket

type Ticket struct {
	Number   int64
	Title    string
	Labels   []string
	Severity string
}

Ticket is one open mining ticket: an issue carrying a severity label.

func MiningTicket

func MiningTicket(number int64, title string, labels []string) (Ticket, bool)

MiningTicket reads an issue as a mining ticket: it carries a severity label and no lang label, which marks the operator-only language tickets.

type Trend

type Trend struct {
	Factor *float64 `json:"factor"`
	Low    *float64 `json:"low"`
	High   *float64 `json:"high"`
	Weeks  int      `json:"periods"`
}

Trend is a compounding rate: the multiplicative change of the struggle rate per period with its 95% interval, over the periods that entered the fit.

func Compounding

func Compounding(periods []Rate, period Period) Trend

Compounding fits ln(struggles/calls) on the period index over the periods with at least one struggle and activeCalls tool calls, weighting each by its struggles, and returns the per-period factor with its 95% interval: measure.py's trend, over days or weeks.

type Usage

type Usage struct {
	CostUSDMicros int64
	InputTokens   int64
	OutputTokens  int64
	Seconds       int64
	Results       int
}

Usage is what a fixer session cost, folded from its event log.

func FoldUsage

func FoldUsage(lines []byte) Usage

FoldUsage reads a session's cost from its event log. The executor reports total_cost_usd cumulatively within one process and from zero when the process is reopened, so the cost is the sum of the maximum of each non-decreasing run; tokens are summed over every result; seconds span the first record to the last.

type Verdict

type Verdict string

Verdict is the pre-check's answer on one proposal.

const (
	// VerdictLaunch: the ticket names an instance that exists in the corpus,
	// so a fixer may be launched.
	VerdictLaunch Verdict = "launch"
	// VerdictNeedsLabels: no named instance exists, so no model is paid and
	// the proposal waits in the labeler's queue.
	VerdictNeedsLabels Verdict = "needs_labels"
)

The verdicts, as the ledger stores them.

Directories

Path Synopsis
Package codes holds CSF's failure codes and codes complaints with them.
Package codes holds CSF's failure codes and codes complaints with them.
Package evaluate scores csf builds on a held-out, rotated suite of past tickets before they go live (EVAL-SUITE, #416).
Package evaluate scores csf builds on a held-out, rotated suite of past tickets before they go live (EVAL-SUITE, #416).
Package labeler is the ouroboros labeler (csf_staging#318): for a mining ticket that names no real instance, it finds candidate spans in the corpus the ticket's miner would read, asks a local model to label them, and keeps only the labels that reproduce.
Package labeler is the ouroboros labeler (csf_staging#318): for a mining ticket that names no real instance, it finds candidate spans in the corpus the ticket's miner would read, asks a local model to label them, and keeps only the labels that reproduce.
Package mocks is a generated GoMock package.
Package mocks is a generated GoMock package.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL