Documentation
¶
Overview ¶
Package terms reads the terms of a text and tells which of them a vocabulary has not seen. It is the extraction behind the harness's unvetted operator terms: the words of an operator's message that appear in none of the operator's earlier messages. It is a pure library: it starts no goroutines and crosses no boundary.
A term is a word of the text after everything that is not vocabulary is set aside: fenced and inline code, words carrying digits or identifier punctuation (paths, flags, hashes, dotted and underscored names), camelCase identifiers, words shorter than MinimumLength, and common English (the stoplist). Words are lowercased, runs of three or more of one letter are collapsed to one, and inflections are folded by Stem, so "cgroups" and "cgroup" are one term. The stoplist is matched on the word and on its stem.
Derivation ¶
Every parameter was fit on the operator's history: 1,903 operator messages over 157 sessions, 2026-08-01 to 2026-10-03, a private corpus that is not in this repository. The corpus was walked in time order; the first 1,522 messages built the vocabulary and the last 381 were the holdout whose novel terms were hand-labelled as a term worth checking, a borderline domain use of an English word, a typo, or an ordinary English word. The rules and what each measured on the holdout:
- Minimum length 3 (2 and 4 were tried): 2 admits letter fragments of identifiers; 4 drops ros, gin, mcp, ssh, bpf and the like.
- Identifier words dropped whole before splitting: UUIDs and hashes were a tenth of the novel words when their letter runs were read as words.
- Hyphenated words split: whole, the recipe slugs and ad hoc compounds were a third of the novel words (673 against 455 terms).
- Stoplist cutoff: the whole frequency list. Flagged terms fell monotonically with the cutoff (320 at none, 287 at 3,000, 258 at 5,000, 225 at 10,000) and the words the last step absorbed were common English (census, formula, hypothesis, tonight), none a term.
- Stemming: Porter on the text and the stoplist. Against a light plural-and-tense fold it took the flagged share of messages from 33% to 28% and the false-positive rate from 69% to 62%, losing four domain uses of English words (compaction, contrastive, coupling, responder).
- Rejected: folding a novel word onto a vetted one at edit distance one absorbed real terms (cgroup into group, dockerd into docker, premise into precise: 14 of its 50 absorptions). Folding transpositions alone absorbed cgroup into an earlier typo of it. Stripping prefixes (un-, re-, over-, sub-) saved 3 points and lost subtree, dequeue, unexported, oversize and overfile.
On the holdout the chosen rules, as this package implements them, flag 107 of 381 operator messages (28.1%), 191 distinct terms, a median of 1 and at most 7 per flagged message; by the hand labels 62 are terms, 9 borderline, 42 typos and 78 ordinary English, a false-positive rate of 120/191 = 62.8% (95% Wilson interval 55.8% to 69.4%). The typos are the operator's and are left to the reply: a research check that says "a typo of literally" costs a line. The ordinary words are English the frequency list does not reach; a larger list would absorb them and the borderline class with them. Measured 2026-10-04; the corpus, the labels and the replay program stay outside the repository because they are the operator's words.
Index ¶
Constants ¶
const MinimumLength = 3
MinimumLength is the shortest word read as a term, in letters.
Variables ¶
This section is empty.
Functions ¶
func Same ¶
Same reports whether two terms share a stem: one spelled "Cgroups" names the same term as one spelled "cgroup".
func Stem ¶
Stem is the Porter stem of word: M. F. Porter, "An algorithm for suffix stripping", Program 14(3):130-137, 1980, implemented step for step as the paper states it. word is lowercase ASCII letters; anything else is returned unchanged. A stem is not a word ("analogies" and "analogy" both stem to "analogi"); it is the key two inflections of one word share.
Types ¶
type Vocabulary ¶
type Vocabulary struct {
// contains filtered or unexported fields
}
Vocabulary is the set of terms seen so far, keyed by stem.
func (*Vocabulary) Add ¶
func (vocabulary *Vocabulary) Add(terms ...Term)
Add records terms as seen.
func (*Vocabulary) Has ¶
func (vocabulary *Vocabulary) Has(term Term) bool
Has reports whether a term with the same stem was added.
func (*Vocabulary) Len ¶
func (vocabulary *Vocabulary) Len() int
Len is the number of distinct stems added.
func (*Vocabulary) Novel ¶
func (vocabulary *Vocabulary) Novel(text string) []Term
Novel returns the terms of text the vocabulary has not seen, in order of first appearance, one per stem. It does not add them.