Documentation
¶
Index ¶
- Constants
- func FindTokenSignatureMatches(mediaType slugs.MediaType, queryName string, candidates []database.MediaTitle) []string
- func FuzzyLengthWindow(queryLength, slack int) (low, high int)
- func GenerateTokenSignature(mediaType slugs.MediaType, gameName string) string
- func SameTitleNumbers(a, b string) bool
- func TokenCoverageRatio(mediaType slugs.MediaType, query, candidateName string) float64
- func TokenPrefixMatch(mediaType slugs.MediaType, query, candidateName string) bool
- type FuzzyMatch
- func ApplyDamerauLevenshteinTieBreaker(query string, matches []FuzzyMatch, topN int) []FuzzyMatch
- func FilterByTokenCoverage(mediaType slugs.MediaType, query string, matches []FuzzyMatch, ...) []FuzzyMatch
- func FindFuzzyMatches(query string, candidates []string, maxDistance int, minSimilarity float32) []FuzzyMatch
Constants ¶
const ( MinSlugLengthForFuzzy = 5 FuzzyMatchMaxLengthDiff = 2 FuzzyMatchMinSimilarity = 0.85 )
Shared thresholds keep title discovery and launch resolution aligned.
Variables ¶
This section is empty.
Functions ¶
func FindTokenSignatureMatches ¶
func FindTokenSignatureMatches( mediaType slugs.MediaType, queryName string, candidates []database.MediaTitle, ) []string
FindTokenSignatureMatches finds candidates where the token signature exactly matches the query signature. This enables word-order independent matching: "Crystal Space Quest" matches "Quest Space Crystal".
The query and candidates must have word boundaries (e.g., "Super Mario World"), not slugs. The mediaType ensures consistent parsing between query and indexed candidates. Returns the slugs of matched titles.
func FuzzyLengthWindow ¶ added in v2.18.0
FuzzyLengthWindow is how far a candidate slug's length may sit from the query's before it is discarded without being scored.
slack widens it upwards only, for a query that may have lost an abbreviation expansion to a typo: slug normalisation turns "Super Mario Bros." into "supermariobrothers" (18) while the typo "Super Mario Bross" stays "supermariobross" (15), so the title the user meant sits three characters away and a flat window of two threw it out unscored. Expansion only ever lengthens, so the lower bound does not move.
The widening is deliberately not unconditional. Applying it to every query measured 26ms to 531ms per lookup on the MiSTer test device against 24,000 titles, because the length bounds prune whole candidate blocks before any of them is read.
func GenerateTokenSignature ¶
GenerateTokenSignature creates a normalized, sorted token signature for word-order independent matching. Uses the same tokenization pipeline as slugification to ensure consistency.
IMPORTANT: Requires input with word boundaries (e.g., "Super Mario World"), not slugs. Slugs have already lost word boundary information and will produce incorrect signatures.
The mediaType parameter ensures media-type-specific parsing is applied (e.g., ParseGame for games), matching the indexing pipeline used in GenerateSlugWithMetadata.
Example:
GenerateTokenSignature(slugs.MediaTypeGame, "Super Mario World") → "mario_super_world" GenerateTokenSignature(slugs.MediaTypeGame, "Mario World Super") → "mario_super_world"
func SameTitleNumbers ¶ added in v2.19.0
SameTitleNumbers reports whether two already-slugified titles carry the same embedded numbers, comparing runs of ASCII digits in order. It never parses a run as an integer — titles are untrusted input, and a digit run of arbitrary length must not risk an overflow or allocation blowup.
Leading zeros are ignored ("touhou06" agrees with "touhou6"), and a lone "1" is treated as no number: a series' first game is often unnumbered ("finalfantasy" / "finalfantasy1" name the same game). Any other difference — including one side having a number the other lacks entirely — means the titles disagree. This is what stops "streetfighter2" from being treated as a typo of "streetfighter": Jaro-Winkler alone can't tell a typo from a sequel number, only SameTitleNumbers can.
func TokenCoverageRatio ¶ added in v2.19.0
TokenCoverageRatio reports the fraction of query's *required* word tokens that have a close match among candidateName's tokens, matching each candidate token to at most one query token via a maximum bipartite matching - not a greedy, query-order-dependent assignment, which can miss a valid pairing that exists: e.g. query tokens ["cattle", "castel"] against candidate tokens ["castle", "battle"] has a perfect matching (cattle-battle, castel-castle), but greedily assigning "cattle" to its single best match first ("castle", the closer of its two eligible candidates) leaves "castel" with no eligible candidate left, even though swapping the query tokens' order would have found it - the query's incidental word order must not change whether coverage is complete. 1.0 means every required query token was accounted for; candidateName may still carry extra tokens the query never mentioned without being penalized here (a bare-prefix relationship, scored separately by TokenPrefixMatch). A query token that is a lone "1" is never required: SameTitleNumbers treats a lone "1" the same way, since a series' first game is often unnumbered.
A token "matches" if it's identical, close by Jaro-Winkler similarity, or - since normal slugification only expands a correctly-spelled abbreviation, leaving a typo of one exactly as typed - a typo of a known abbreviation whose expansion matches instead ("bross" against "brothers", the same tolerance AbbreviationExpansionSlack already gives the whole-string case).
This exists because a whole-string Jaro-Winkler score, however heavily patched, can't reliably tell "these are the same title with a typo" from "these happen to share a lot of characters": "streetfighter2turbo" scores 0.927 against "streetfighterzero2" by whole-string similarity alone, sharing a "streetfighter" prefix and a "2" that SameTitleNumbers can't separate, while the words "turbo" and "zero" are simply unrelated - which per-token comparison sees directly (0.75 coverage, one required token unmatched) instead of having to infer it from character-level side effects.
It is not a universal fix. Two title pairs can have the identical *shape* of disagreement with opposite ground truth, and no string-shape metric resolves that from the text alone - this function does not try to.
query and candidateName must be original titles with word boundaries (e.g. "Street Fighter II"), not slugs - SlugifyWithTokens needs the spaces. A word-order match from GenerateTokenSignature should not be checked this way: it already requires every token to match, by definition.
func TokenPrefixMatch ¶ added in v2.19.0
TokenPrefixMatch reports whether query's word tokens are a strict, non-empty prefix of candidateName's word tokens: same order, every query token matched, and the candidate has at least one more. It compares tokens rather than slug bytes so a compound word like "Firefly" cannot be prefix-matched by "Fire" — slugification drops the space that would otherwise distinguish them, but tokenization (done before spaces are dropped) still has it.
query and candidateName must be original titles with word boundaries (e.g. "Street Fighter II"), not slugs — SlugifyWithTokens needs the spaces.
Types ¶
type FuzzyMatch ¶
FuzzyMatch represents a slug that matches the query with a similarity score.
func ApplyDamerauLevenshteinTieBreaker ¶
func ApplyDamerauLevenshteinTieBreaker(query string, matches []FuzzyMatch, topN int) []FuzzyMatch
ApplyDamerauLevenshteinTieBreaker refines fuzzy matches using Damerau-Levenshtein distance to handle transposition errors (e.g., "crono tigger" → "Chrono Trigger").
It takes the top N candidates from Jaro-Winkler and re-ranks them by edit distance. This two-stage approach is more accurate than either algorithm alone while remaining fast.
func FilterByTokenCoverage ¶ added in v2.19.0
func FilterByTokenCoverage( mediaType slugs.MediaType, query string, matches []FuzzyMatch, namesBySlug map[string]string, ) []FuzzyMatch
FilterByTokenCoverage drops matches whose candidate (looked up in namesBySlug by FuzzyMatch.Slug) does not fully cover query's word tokens per TokenCoverageRatio. Applied after Jaro-Winkler and its tie-breaker have already narrowed the candidate set to a handful, since computing token coverage for every length-eligible candidate would cost as much as the similarity scoring it exists to double-check. A candidate missing from namesBySlug scores zero coverage and is dropped - every caller builds the map from the same candidate list FindFuzzyMatches scored, so this should not happen in practice, but a match this function cannot evaluate is not one it can vouch for either.
func FindFuzzyMatches ¶
func FindFuzzyMatches(query string, candidates []string, maxDistance int, minSimilarity float32) []FuzzyMatch
FindFuzzyMatches returns slugs that fuzzy match the query using Jaro-Winkler similarity. Jaro-Winkler is optimized for short strings and heavily weights matching prefixes, making it ideal for game titles where users typically get the start correct. It also naturally handles British/American spelling variations (e.g., "colour" vs "color"). Results are filtered by maxDistance and minSimilarity, sorted by similarity (best first).