Documentation
¶
Overview ¶
Package codes holds CSF's failure codes and codes complaints with them.
A failure code is one recurring way the agent system fails. Each is declared once, in csf/compiler/language/architecture.csf, with the edge between the two components where it starts and the fault side that owns its repair; the language generator projects the declarations to Catalogue. MAST's 14 modes are the seed (Cemri et al., arXiv:2503.13657); a code CSF induces from complaints the seed does not fit names the mode it refines, as AdaMAST induces its codes from traces (arXiv:2607.16387). The edge and side follow the interaction-centric taxonomy (arXiv:2607.28802).
A complaint is an operator correction: an operator-authored message that pkg/affect reads as flagging the agent's mistake. ComplaintCoder asks a local model behind the brain contract to code each one, and keeps a code only when the model's quote from the complaint is in it verbatim; a complaint no code fits comes back with the model's proposal for a new code. Tally counts the coded complaints per code, edge and side. Nothing here reaches a session: the codes are structure, read by the views and by whoever decides which gate or miner to build next.
Index ¶
- Constants
- Variables
- func Clusters(coded []Coded) map[string]string
- func Kappa(left map[string]string, right map[string]string) float64
- func Matching(left map[string]string, right map[string]string) (share float64, shared int)
- func ReadReference(reference io.Reader) (classes map[string]string, sides map[string]string, err error)
- func Select(messages []Complaint, reference map[string]string) ([]Complaint, ReaderRecall)
- func SidesOf(coded []Coded, catalogue []Code) map[string]string
- type ClusterAgreement
- type Code
- type CodeCount
- type Coded
- type Complaint
- type ComplaintCoder
- type ComplaintCoderOption
- type Count
- type Edge
- type Induced
- type Proposal
- type ProposalCount
- type ReaderRecall
- type Record
- type Side
- type Summary
Constants ¶
const ( RejectedNoQuote = "quote not in the complaint" RejectedUnknownCode = "code not in the catalogue" RejectedUnanswered = "the model gave no answer for the complaint" )
The reasons an answer is not kept.
const CodeNone = "none"
CodeNone is the code a model answers when no catalogue code fits; the answer then carries its proposal for a new code.
const DefaultBatch = 8
DefaultBatch is how many complaints one model request codes. The window is not the bound: qwen3:8b answers fewer complaints than a batch holds as the catalogue grows. Measured 2026-10-05 on the 28 complaints of the night before, 16k window, no thinking: with the 24-code catalogue a batch of 24 left 8 of the 16 corrections unanswered and a batch of 8 left 4; with the 14-code seed a batch of 8 left 1 of 28 complaints rejected.
const RecordFile = "failure_codes.jsonl"
RecordFile is the coded complaints' file under the harness state directory, one Record per line, appended by every coding run.
Variables ¶
var ( // ErrNoBrain reports a coder built without the model. ErrNoBrain = errors.New("failure codes: a brain is required") // ErrInvalidOption reports a nil option or a value the coder cannot use. ErrInvalidOption = errors.New("failure codes: invalid option") )
var Catalogue = []Code{ { ID: "mast_disobey_task", Name: "Disobey task specification", Definition: "The agent does not do what the assignment specifies, or does it under constraints the assignment ruled out.", Edge: Edge{From: "assignment", To: "agent"}, Side: SideModel, MAST: "FM-1.1", }, { ID: "mast_disobey_role", Name: "Disobey role specification", Definition: "The agent acts outside the role its recipe gives it, doing another station's work.", Edge: Edge{From: "harness", To: "agent"}, Side: SideModel, MAST: "FM-1.2", }, { ID: "mast_step_repetition", Name: "Step repetition", Definition: "The agent repeats a step it already completed.", Edge: Edge{From: "agent", To: "harness"}, Side: SideModel, MAST: "FM-1.3", }, { ID: "mast_history_loss", Name: "Loss of conversation history", Definition: "The agent loses earlier context and acts as if it had not been said.", Edge: Edge{From: "harness", To: "context"}, Side: SideHarness, MAST: "FM-1.4", }, { ID: "mast_termination_unaware", Name: "Unaware of termination conditions", Definition: "The agent does not know when the work is done and keeps going or stops at the wrong point.", Edge: Edge{From: "agent", To: "assignment"}, Side: SideModel, MAST: "FM-1.5", }, { ID: "mast_conversation_reset", Name: "Conversation reset", Definition: "The conversation restarts and progress made in it is lost.", Edge: Edge{From: "session", To: "agent"}, Side: SideHarness, MAST: "FM-2.1", }, { ID: "mast_no_clarification", Name: "Fail to ask for clarification", Definition: "The agent proceeds on an ambiguous instruction instead of asking.", Edge: Edge{From: "agent", To: "human"}, Side: SideModel, MAST: "FM-2.2", }, { ID: "mast_task_derailment", Name: "Task derailment", Definition: "The agent drifts from the assignment onto other work.", Edge: Edge{From: "agent", To: "assignment"}, Side: SideModel, MAST: "FM-2.3", }, { ID: "mast_information_withholding", Name: "Information withholding", Definition: "The agent has information the operator or another agent needs and does not share it.", Edge: Edge{From: "agent", To: "human"}, Side: SideModel, MAST: "FM-2.4", }, { ID: "mast_ignored_input", Name: "Ignored other agent's input", Definition: "The agent disregards what another agent or the operator told it.", Edge: Edge{From: "relay", To: "agent"}, Side: SideModel, MAST: "FM-2.5", }, { ID: "mast_reasoning_action_mismatch", Name: "Reasoning-action mismatch", Definition: "What the agent does differs from what its own reasoning said it would do.", Edge: Edge{From: "model", To: "agent"}, Side: SideModel, MAST: "FM-2.6", }, { ID: "mast_premature_termination", Name: "Premature termination", Definition: "The agent stops before the assignment's outcome is reached.", Edge: Edge{From: "agent", To: "assignment"}, Side: SideModel, MAST: "FM-3.1", }, { ID: "mast_no_verification", Name: "No or incomplete verification", Definition: "The work is reported done without checking it, or with checks that miss what matters.", Edge: Edge{From: "agent", To: "session_gate"}, Side: SideHarness, MAST: "FM-3.2", }, { ID: "mast_incorrect_verification", Name: "Incorrect verification", Definition: "A check passes work that is wrong.", Edge: Edge{From: "session_gate", To: "agent"}, Side: SideHarness, MAST: "FM-3.3", }, { ID: "induced_communication_inefficiency", Name: "Communication inefficiency", Definition: "Inefficient communication between the agent and human due to low information density.", Edge: Edge{From: "agent", To: "human"}, Side: SideHarness, }, { ID: "induced_configuration_error", Name: "Configuration error", Definition: "Errors in configuration leading to model switching issues.", Edge: Edge{From: "harness", To: "model"}, Side: SideHarness, }, { ID: "induced_design_reinvention", Name: "Design reinvention", Definition: "Unnecessary complexity and reinvention instead of reusing existing components.", Edge: Edge{From: "harness", To: "host"}, Side: SideHarnessEnvironment, }, { ID: "induced_visibility_loss", Name: "Visibility loss", Definition: "Missing visibility into system metrics and dashboards.", Edge: Edge{From: "harness", To: "infra"}, Side: SideEnvironment, }, { ID: "induced_unexpected_failure", Name: "Unexpected failure", Definition: "Unexpected failures in session execution or processing.", Edge: Edge{From: "session", To: "harness"}, Side: SideHarness, }, { ID: "induced_dependency_failure", Name: "Dependency failure", Definition: "Issues with external dependencies leading to system instability.", Edge: Edge{From: "harness", To: "vendor"}, Side: SideProvider, }, { ID: "induced_tooling_script_failure", Name: "Tooling script failure", Definition: "Failures in tooling and script execution due to host configuration.", Edge: Edge{From: "harness", To: "host"}, Side: SideEnvironment, }, { ID: "induced_search_reranking_failure", Name: "Search reranking failure", Definition: "Issues with search and reranking mechanisms affecting information retrieval.", Edge: Edge{From: "model", To: "harness"}, Side: SideModel, }, { ID: "induced_inconsistency_alignment", Name: "Inconsistency alignment", Definition: "Inconsistencies in ontology alignment leading to system instability.", Edge: Edge{From: "harness", To: "model"}, Side: SideModel, }, { ID: "induced_mining_failure", Name: "Mining failure", Definition: "Failures in mining processes leading to missing primitives.", Edge: Edge{From: "miner", To: "harness"}, Side: SideHarness, }, }
Catalogue is every failure code architecture.csf declares, in its order.
var Components = map[string]string{
"human": "the operator",
"agent": "the agent doing the work in a session",
"model": "the large model behind the agent",
"harness": "the CSF harness that runs sessions, its verbs and its recipes",
"session_gate": "a structural check the harness installs into every session",
"assignment": "the work an agent was given: ticket, recipe and instructions",
"context": "what the agent knows in a session: conversation, memory, rulings",
"session": "one agent conversation and its lifecycle",
"relay": "messages between agents and sessions",
"turn_executor": "the CLI that runs the agent's turns (Claude Code, Copilot)",
"host": "the machine and its processes, files and caches",
"infra": "deployment and infrastructure outside the harness",
"worktree": "the session's git checkout and branch",
"miner": "an ouroboros rule that detects a failure in records",
"receipt": "the links a run returns: branch, pull request, trace",
"csfpg": "CSF's PostgreSQL state",
"vendor": "an external provider or service CSF depends on",
}
Components is every ontology term a code's edge may name, with the meaning a coder is told: the endpoints the catalogue already uses and the parts of the system a complaint can be about.
var Sides = []Side{SideHarness, SideHarnessEnvironment, SideEnvironment, SideProvider, SideModel}
Sides is every fault side, in the order a tally reports them.
Functions ¶
func Clusters ¶
Clusters is each kept complaint's cluster, keyed by its ref: its catalogue code, or its proposal's name; a rejected answer is in no cluster.
func Kappa ¶
Kappa is Cohen's kappa between two labelings over the keys both hold: the agreement beyond what the two labelings' own frequencies give by chance. It is 0 with no shared key and 1 when both labelings use one label alike.
func Matching ¶
Matching is the share of the keys both labelings hold on which they give the same label, and how many keys that is. Kappa corrects it for chance, but kappa is 0 whenever one labeling uses a single label, so a reference that puts every complaint on one side is read by this share.
func ReadReference ¶
func ReadReference(reference io.Reader) (classes map[string]string, sides map[string]string, err error)
ReadReference reads a hand clustering of complaints, the reference a coding is measured against: one complaint per line, its ref, class and fault side separated by tabs. Blank lines and lines starting with # are skipped; a side may be left empty.
func Select ¶
func Select(messages []Complaint, reference map[string]string) ([]Complaint, ReaderRecall)
Select is the messages to code: every correction, and every message the reference labels, so a coding is measured on the reference's own messages even where the correction reader misses them; the reference's labels never reach the model. It returns the reader's recall on the reference.
Types ¶
type ClusterAgreement ¶
type ClusterAgreement struct {
Complaints int `json:"complaints"`
Clusters int `json:"clusters"`
Classes int `json:"classes"`
Purity float64 `json:"purity"`
InversePurity float64 `json:"inverse_purity"`
F float64 `json:"f"`
}
ClusterAgreement compares one clustering of complaints with a reference clustering of the same complaints: purity is the share of complaints in their cluster's majority reference class, inverse purity the share in their class's majority cluster, and F their harmonic mean. A clustering that puts every complaint alone scores purity 1, and one that puts all in one cluster scores inverse purity 1, so only F rewards matching both.
type Code ¶
type Code struct {
ID string
Name string
Definition string
Edge Edge
Side Side
// MAST is the MAST mode the code refines, FM-<category>.<mode>; empty for
// an induced code that refines none.
MAST string
// Miner is the directory of the miner whose rule detects the code; empty
// until one is accepted.
Miner string
}
Code is one failure code as architecture.csf declares it.
type CodeCount ¶
type CodeCount struct {
Code string `json:"code"`
Name string `json:"name"`
Edge Edge `json:"edge"`
Side Side `json:"side"`
Route string `json:"route"`
Count int `json:"count"`
}
CodeCount is one catalogue code's complaints.
type Coded ¶
type Coded struct {
Complaint Complaint
// Code is the catalogue code kept; empty when the answer was rejected or
// proposes a new code.
Code string
// Quote is the model's verbatim evidence from the complaint.
Quote string
// Proposal is the new code proposed when no catalogue code fits.
Proposal *Proposal
// Rejected is why the answer was not kept; empty when it was.
Rejected string
}
Coded is one complaint and the code it was given.
type Complaint ¶
type Complaint struct {
// Source names the corpus item, such as a transcript's session.
Source string
// Ref identifies the complaint within its source.
Ref string
At time.Time
Text string
// Context is the agent's reply the complaint answers; empty when none
// was recorded.
Context string
// Correction is pkg/affect reading the message as an operator
// correction; a message read only because a reference labels it is not.
Correction bool
}
Complaint is one operator correction to code: where it came from, what the operator wrote and what the agent had said just before.
func ReadTranscript ¶
func ReadTranscript(source string, transcript io.Reader, since time.Time, until time.Time) ([]Complaint, error)
ReadTranscript reads one Claude Code session transcript, named source, for the operator's own messages (origin human) sent in [since, until), each with the agent's last text before it as its context and, as Correction, whether pkg/affect reads it as an operator correction: a complaint. A zero until is open-ended.
type ComplaintCoder ¶
type ComplaintCoder struct {
// contains filtered or unexported fields
}
ComplaintCoder codes complaints with the catalogue through a local model. It holds no state between calls.
func NewComplaintCoder ¶
func NewComplaintCoder(options ...ComplaintCoderOption) (*ComplaintCoder, error)
NewComplaintCoder builds the coder; the brain is required.
func (*ComplaintCoder) Code ¶
Code codes every complaint, in order, one batch per model request. A failed request ends the run with the complaints coded so far.
type ComplaintCoderOption ¶
type ComplaintCoderOption func(coder *ComplaintCoder) error
ComplaintCoderOption configures a ComplaintCoder.
func WithBatch ¶
func WithBatch(size int) ComplaintCoderOption
WithBatch codes size complaints per model request instead of DefaultBatch.
func WithCatalogue ¶
func WithCatalogue(catalogue []Code) ComplaintCoderOption
WithCatalogue codes with catalogue instead of Catalogue.
func WithKeepAlive ¶
func WithKeepAlive(duration time.Duration) ComplaintCoderOption
WithKeepAlive asks the model server to keep the model loaded for duration after each answer; the server's default applies otherwise.
type Induced ¶
type Induced struct {
Proposal Proposal `json:"proposal"`
Members []string `json:"members"`
Quotes []string `json:"quotes"`
}
Induced is one code the model induced from a set of complaints, with the complaints it covers: those holding one of its quotes verbatim. A code with no member is kept, so a run shows what the model proposed even when none of its quotes reproduces.
type Proposal ¶
type Proposal struct {
Name string `json:"name"`
Definition string `json:"definition"`
Edge Edge `json:"edge"`
Side Side `json:"side"`
}
Proposal is a model's proposal for a new code, for a complaint no catalogue code fits.
type ProposalCount ¶
type ProposalCount struct {
Name string `json:"name"`
Proposal Proposal `json:"proposal"`
Count int `json:"count"`
}
ProposalCount is the complaints whose proposals share one name, with the first proposal made under it.
type ReaderRecall ¶
ReaderRecall is how many of a reference's labeled messages the corpus holds, and how many of those pkg/affect reads as corrections.
type Record ¶
type Record struct {
Source string `json:"source"`
Ref string `json:"ref"`
At time.Time `json:"at"`
// Code is the catalogue code, empty for a proposal or a rejection.
Code string `json:"code"`
Edge Edge `json:"edge"`
Side Side `json:"side"`
Quote string `json:"quote"`
Proposal *Proposal `json:"proposal,omitempty"`
Rejected string `json:"rejected,omitempty"`
// Model names the provider and model that coded it.
Model string `json:"model"`
}
Record is one coded complaint as the state directory keeps it: where the complaint is, the code it was given with its edge and side, and the model's evidence. The complaint's own text stays in its source.
type Side ¶
type Side string
Side is a failure code's fault side: who owns its repair.
type Summary ¶
type Summary struct {
Complaints int `json:"complaints"`
Coded int `json:"coded"`
Proposed int `json:"proposed"`
Rejected int `json:"rejected"`
Codes []CodeCount `json:"codes"`
Proposals []ProposalCount `json:"proposals"`
// Edges and Sides count every kept answer: a catalogue code's edge and
// side, or a proposal's.
Edges []Count `json:"edges"`
Sides []Count `json:"sides"`
}
Summary is a coded corpus counted per code, edge and side; codes and proposals are ordered by count, most first, so the first code is the one whose gate pays most.