bench

package
v0.7.9 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jul 20, 2026 License: AGPL-3.0 Imports: 21 Imported by: 0

README

memini benchmark harness

A retrieval benchmark: ingest a dataset of memories, then for each question measure how well a system retrieves the gold supporting memories.

The Go eval harnesses in this directory (*_test.go) live behind the bench build tag, so they stay out of the default go test ./... (they need a live embedder and take minutes). Run them with -tags bench, as shown below; a plain go test ./bench/ reports "no test files" by design.

mise run bench                 # offline sample, local embedder
go run ./cmd/bench -k 5        # same, explicit K

Against a real embeddings model and a real dataset:

export MEMINI_EMBED_BASE_URL=http://localhost:8081/v1
export MEMINI_EMBED_MODEL=bge-m3 MEMINI_EMBED_DIMS=1024
# Optional: instruction-tuned asymmetric embedders (Qwen3-Embedding, bge) score
# higher when queries carry a retrieval instruction; documents stay bare.
# Measured on Qwen3-Embedding-8B: +6.0pp R@5 on the LongMemEval vector leg
# (91.2% -> 97.2%), +1.0pp MRR on the fused ranking on both datasets.
export MEMINI_EMBED_QUERY_PREFIX=$'Instruct: Given a user query, retrieve relevant memories that answer it\nQuery:'
go run ./cmd/bench -suite longmemeval -data ./longmemeval_s.json -k 5
go run ./cmd/bench -suite locomo      -data ./locomo.json        -k 5

# Isolate the recency-aware re-ranker against pure RRF on the same candidates,
# using each question's date as "now" (needs a timestamped dataset):
go run ./cmd/bench -suite longmemeval -data ./longmemeval_s.json -rerank -k 5

Full results

Everything this harness measures, in one table — sourced from the committed results/ JSON, all on the same all-MiniLM-L6-v2 (384-d) endpoint. Cells are recall_any@5 / @10 / MRR (%); p50 is in-process recall latency (rerank rows show the added cost). The detailed per-dataset sections below explain the methodology, sweeps, and caveats behind each column.

Strategy LongMemEval · session LoCoMo · turn-level LoCoMo · session-level p50
vector 92.6 / 95.4 / 80.7 41.3 / 51.8 / 28.1 64.1 / 79.8 / 45.2 <1 ms
keyword (Porter BM25) 97.6 / 99.0 / 92.2 58.7 / 67.1 / 44.8 92.6 / 96.8 / 79.4 ~3 ms
hybrid (default, production path) 98.4 / 99.2 / 93.0 59.7 / 69.9 / 42.4 90.9 / 96.6 / 74.3 ~5 ms
+ cross-encoder (MEMINI_RERANK=<url>) 98.4 / 99.2 / 93.1 70.9 / 75.0 / 59.8 90.9 / 96.6 / 74.3 +20–230 ms
+ LLM rerank (MEMINI_RERANK=llm) 98.4 / 99.2 / 93.0 74.4 / 76.5 / 67.4 +350–420 ms

Questions per dataset: LongMemEval 500 (session granularity), LoCoMo turn-level 1,982 (gold = exact evidence turns), LoCoMo session-level 1,981 (gold = sessions holding those turns). Rerank backends: Qwen3-Reranker-0.6B (cross-encoder) and Qwen3.5-9B (LLM). Reproduce with the per-suite commands in the sections below (-suite longmemeval, locomo, locomo-sessions; add -rerank-url/-llm-rerank for the rerank rows).

Reading it: hybrid never trails either single leg on the saturated session sets (it ties keyword on LoCoMo-session, where keyword's exact-token match is already near-ceiling). On turn-level LoCoMo base recall has real headroom, so the rerank tier earns its keep — the cross-encoder lands +11pp R@5 / +17pp MRR over hybrid at a fraction of the LLM's latency, and the LLM adds a few more points (+15pp / +25pp) if you already run a chat model. Where recall is already at ceiling (both session sets), reranking is a measured no-op.

Results: memini vs other memory systems

All memini numbers below are measured by this harness against a live all-MiniLM-L6-v2 (384-d) endpoint — the same embedding model agentmemory benchmarks with. Competitor numbers are cited from their own publications — we cannot re-run their systems here, and they use different embedding models, readers, and judges. Treat cross-system rows as directional, not a controlled head-to-head. (This mirrors how agentmemory documents its comparison.)

LongMemEval-S — retrieval recall_any@K

Full 500-question LongMemEval-S (~48 sessions/question), same metric agentmemory reports: does any gold session appear in the top-K retrieved? No LLM in the loop — pure retrieval. The run is the full 500 questions with the identical embedding model agentmemory benchmarks with (all-MiniLM-L6-v2, 384-d) for a true apples-to-apples comparison.

Hybrid recall over-fetches a deep candidate pool per leg (max(k*5, 50)) before fusing, so a memory just outside the top-k of both legs can still win — the production Recall path does the same. Fusion is convex score fusion (alpha 0.5, the baked default): each leg's scores are min-max normalized to [0,1] and combined 0.5·vector + 0.5·keyword, keeping score magnitude so a memory a leg ranks far above its runners-up dominates one that is merely middling in both. A negative alpha falls back to Reciprocal Rank Fusion; deep pools then need a steep decay (rrfK=5, not the classic 60), since a flat decay lets both-leg mediocrity outscore single-leg excellence (2/(60+20) > 1/(60+0)). Score fusion gets the same effect from score magnitude directly, and beat RRF on 3 of 4 model×dataset cells (and on MRR in all 4).

System Embedding model R@5 R@10 Source
memini — hybrid (score) all-MiniLM-L6-v2 (384-d) 98.4% 99.4% measured
memini — keyword (Porter BM25) 97.6% 99.0% measured
memini — vector all-MiniLM-L6-v2 91.8% 96.6% measured
agentmemory — BM25 + Vector all-MiniLM-L6-v2 95.2% 98.6% published
agentmemory — BM25 only 86.2% 94.6% published
MemPalace (vector only) larger model ~96.6% self-reported

On the same model/dataset/metric (full 500 questions), memini hybrid beats agentmemory at R@5 (98.4% vs 95.2%), R@10 (99.4% vs 98.6%), and MRR (92.3% vs 88.2%). memini's keyword leg is +11.4pp over agentmemory's BM25-only (97.6% vs 86.2%) thanks to Porter stemming, and hybrid fusion now beats either leg alone. Relative to fetching only k per leg with the classic rrfK=60, the deep-pool + score fusion is worth +2.0pp R@5 / +1.0pp R@10.

LoCoMo — retrieval recall_any@K

LoCoMo retrieval at dialogue-turn granularity (1,982 questions over 10 long conversations, gold = exact evidence turns among ~590 turns/conversation) — a much harder target than LongMemEval's session granularity, and the regime where flat-decay RRF over deep pools degrades badly.

System (all-MiniLM-L6-v2) R@5 R@10
memini — hybrid (score) 59.8% 69.8%
memini — keyword (Porter BM25) 58.7% 67.1%
memini — vector 41.5% 52.1%

No published turn-level retrieval baselines exist to compare against (mem0 / Letta report LLM-judged QA accuracy, below). This is the one cell where the default score fusion is edged by RRF (60.1% / 71.0%): when the vector leg is near-noise (MiniLM scores only 41.5% here), giving it an equal-weight normalized vote hurts, whereas RRF's rank-only vote is more robust. Score fusion still wins this cell on MRR and wins outright on every cell with a stronger embedder — so it is the baked default; the benchmark harness can still select RRF (a negative fusion alpha via its flag) for weak-vector experiments. (Ablation: rrfK=60 over the same deep pools scored just 52.8% R@5, below the keyword leg alone — both score fusion and rrfK=5 fix that.)

Pool-depth robustness (-pool-factor / -pool-floor)

Min-max normalization could in principle be fragile to pool depth (the score at the bottom of the pool sets each leg's zero point), so score fusion was swept at per-leg depths 30 / 50 / 80 on both datasets and both embedders (hybrid R@5 / R@10 / MRR):

cell depth 30 depth 50 (default) depth 80
LME · MiniLM 97.8 / 99.4 / 92.0 98.4 / 99.4 / 92.3 98.6 / 99.4 / 92.6
LME · Qwen3+prefix 98.8 / 99.4 / 94.5 98.8 / 99.6 / 94.6 98.8 / 99.6 / 94.6
LoCoMo · MiniLM 60.0 / 70.1 / 42.1 59.8 / 69.8 / 42.6 59.3 / 69.6 / 42.7
LoCoMo · Qwen3+prefix 70.1 / 77.9 / 52.1 70.1 / 78.5 / 52.4 70.1 / 78.7 / 52.5

Quality moves at most ±0.6pp R@5 across a 2.7× depth range — no tail collapse — with the two datasets drifting in opposite directions (deeper pools help session-granularity LongMemEval slightly and hurt turn-granularity LoCoMo slightly), so the default max(k*5, 50) sits at the crossover.

Recency-aware re-ranking (-rerank)

memini re-ranks the fused candidates by a composite of relevance, recency, and importance. The recency weight is deliberately light (0.05): a sweep on LongMemEval-S (knowledge-update + temporal-reasoning, q.Now = question date, sessions timestamped from haystack_dates) shows recency is a net win only as a tie-breaker, and actively harmful when over-weighted.

recency weight R@1 (both cats) knowledge-update R@1 temporal R@1 MRR
0 (pure RRF) 82.9% 91.0% 78.2% 90.1%
0.05 (default) 83.4% 91.0% 78.9% 90.5%
0.15 83.9% 89.7% 80.5% 90.7%
0.25 83.4% 87.2% 81.2% 90.4%

At 0.05 the re-ranker is +0.5pp R@1 / +0.4pp MRR over pure RRF with no knowledge-update cost, and recall@5 is identical across all weights (the re-rank only reorders within the top results). The steep RRF decay made the composite far more robust to the recency weight than the flat rrfK=60 decay was (where 0.15+ buried correct-but-older memories); the default stays at the conservative 0.05 since the gains beyond it are within noise.

Temporal targeting (temporal0.40)

Recency weighting trades off against itself: raising it helps temporal-reasoning (78.2→81.2% R@1) but hurts knowledge-update (91.0→87.2%), whose answers aren't necessarily recent. Temporal targeting avoids that: when a query names a relative time ("three weeks ago"), it computes target = now − offset and boosts candidates dated near that point, not near now. It only fires on temporal queries, so other categories are unaffected.

Strategy all R@1 knowledge-update R@1 temporal-reasoning R@1 MRR
recency 0.05 (prior default) 83.4% 91.0% 78.9% 90.5%
recency 0.25 83.4% 87.2% 81.2% 90.4%
temporal 0.40 85.3% 91.0% 82.0% 91.5%

Temporal targeting is +1.9pp R@1 overall over the recency default and beats even the heaviest recency weight on temporal-reasoning without the knowledge-update regression — so it ships on in production (temporal boost 0.40, the baked default). The no-LLM regex extractor only catches templated phrasing; an LLM anchor extractor (plugging into the same search.AnchorExtractor interface) can resolve looser references and is the intended with-LLM tier.

Held-out split (-holdout)

To avoid overfitting tuning decisions to the full benchmark, -holdout splits LongMemEval deterministically by load order: every 10th question is held (50/500), the rest are tune (450/500). Sweep parameters on -holdout tune, then report the final number on -holdout held (unseen). Default all runs the full set. Results files are suffixed (longmemeval-held.json) so splits don't overwrite each other.

Measured (memini-hybrid, all-MiniLM-L6-v2 — the parameters were swept on tune, not held):

Split Questions R@5 R@10 MRR
all (full) 500 98.4 99.2 93.0
tune 450 98.2 99.1 93.0
held 50 100.0 100.0 93.5

The held split does not regress against tune, so the tuning choices generalize (no tuned-to-test inflation). Per-category R@5:

Category tune (450) held (50)
knowledge-update 100.0 100.0
multi-session 99.2 100.0
single-session-assistant 100.0 100.0
single-session-user 96.8 100.0
temporal-reasoning 98.3 100.0
single-session-preference 88.9 100.0

Read the per-category numbers off tune (450 questions); it shows the real headroom is single-session-preference (88.9% R@5). On held each category is only 2–13 questions, so its across-the-board 100% is small-sample, not a separate claim of perfection.

Session-doc construction (-session-doc)

LongMemEval sessions are embedded as one document per session; -session-doc controls what text that document contains, to measure the vector leg's sensitivity to document shape:

  • full (default) — "role: content" for every turn.
  • user-only — only the user turns, no role prefixes. Assistant turns dilute the embedding for user-question recall; this is the shape MemPalace reports 96.6% R@5 vector-only with on the same MiniLM model.
  • datedfull prefixed with the session date, giving temporal questions a textual anchor embeddings would otherwise ignore.

Compare the vector row's recall_any@5 across modes (cached embeddings make the sweep cheap); the keyword and hybrid rows shift too but the vector leg is the target.

memini hybrid per-category (all-MiniLM, recall_any@10): multi-session 100%, knowledge-update 100%, single-session-user 98.6%, single-session-assistant 98.2%, temporal-reasoning 97.0%, single-session-preference 96.7%.

Rerank tier — cross-encoder vs LLM (-rerank-url / -llm-rerank)

The read-side rerank reorders the top of the production candidate order. The bench drives either backend through the same comparison (one reranker call per question — use -limit):

# cross-encoder (fast; e.g. Qwen3-Reranker-0.6B via llama-server --rerank):
go run ./cmd/bench -suite locomo -data ./locomo.json -rerank-url http://localhost:8002/v1 -rerank-model qwen3-reranker-0.6b -limit 100 -k 5,10
# LLM reranker (slow; MEMINI_LLM_*):
go run ./cmd/bench -suite locomo -data ./locomo.json -llm-rerank -limit 100 -k 5,10

Measured on all-MiniLM-L6-v2 (cross-encoder = Qwen3-Reranker-0.6B, LLM = Qwen3.5-9B), recall_any@5 / @10 / MRR:

Config LongMemEval (session) LoCoMo turn-level added p50
hybrid (base) 98.4 / 99.2 / 93.0 59.7 / 69.9 / 42.4
+ cross-encoder 98.4 / 99.2 / 93.1 70.9 / 75.0 / 59.8 ~20–230 ms
+ LLM rerank 98.4 / 99.2 / 93.0 74.4 / 76.5 / 67.4 ~350–420 ms

Reranking is a no-op at recall ceiling (session-level) and a big win where recall has headroom (turn-level: +11pp R@5 / +17pp MRR for the cross-encoder, +15pp / +25pp for the LLM). The cross-encoder captures most of the LLM's lift at a fraction of the latency with no chat model — the recommended production rerank (MEMINI_RERANK=<url>); the LLM tier (MEMINI_RERANK=llm) buys the last points if you already run one.

LoCoMo — end-to-end QA accuracy (LLM-judge)

The metric mem0/Letta publish: retrieve → generate an answer → an LLM judges it against the gold answer. memini's number uses a fast instruct reader+judge (Llama-3.3-70B-Instruct); the competitor numbers use their own readers/judges, so this is directional.

System LoCoMo QA accuracy Source
memini (hybrid retrieval + instruct reader) full run pending measured
Letta / MemGPT 83.2% published
Mem0 68.5% published

Sources: agentmemory COMPARISON.md/LONGMEMEVAL.md; LongMemEval (arXiv 2410.10813); LoCoMo (snap-stanford.github.io/LoCoMo); mem0.ai; letta.com.

LongMemEval-S — end-to-end QA accuracy (LLM-judge)

Full 500-question LongMemEval-S with the production answer path: hybrid retrieval (bge-m3, 1024-d) → chain-of-thought answer prompt → LLM judge. The answer prompt uses chronological context formatting, <mem_thinking> CoT reasoning tags, "most-recent-wins" for knowledge updates, temporal grounding (question date injected), and 1-hop linked-memory expansion. Competitor numbers use GPT-4o as answerer+judge; memini uses Qwen3.6-35b-fast — so this is directional, not a controlled head-to-head.

System Answerer/Judge QA accuracy Source
mem0 GPT-4o 94.4% mem0ai/memory-benchmarks
Hindsight (TEMPR + CARA) Gemini-3 91.4% Hindsight paper
memini (CoT prompt + agentic loop) Qwen3.6-35b-fast 82.8% measured
Zep (Graphiti) GPT-4o 71.0% published
Full-context (entire history) GPT-4o 60.2% LongMemEval paper

Per-category breakdown (memini, Qwen3.6-35b-fast, k=10, bge-m3):

Category Accuracy n vs original prompt
single-session-assistant 100.0% 56 0
single-session-user 93.8% 64 +1.6pp
knowledge-update 77.8% 72 +8.4pp
temporal-reasoning 86.6% 127 +52.0pp
multi-session 72.7% 121 +33.0pp
single-session-preference 56.7% 30 +50.0pp
abstention (all types) ~92% 36 ~0
Overall 82.8% 500 +25.2pp

The answer-path improvements (chronological context, CoT reasoning, temporal grounding, most-recent-wins, IncludeLinked, agentic tool loop) lifted overall accuracy from 57.6% to 82.8% on the same model and embedder — a +25.2pp gain from prompt engineering and tool-loop enablement alone.

Reproduce:

export MEMINI_EMBED_BASE_URL=https://your-embedder/v1
export MEMINI_EMBED_MODEL=bge-m3 MEMINI_EMBED_DIMS=1024
export MEMINI_LLM_BASE_URL=https://your-llm/v1
export MEMINI_LLM_MODEL=qwen3.6-35b-fast
go run ./cmd/qa -suite longmemeval -data bench/data/longmemeval_s.json \
  -ingest upsert -k 10 -workers 4

Note on comparability: mem0 uses GPT-4o (trillion-parameter) for both answering and judging. memini uses Qwen3.6-35b-fast (35B). The gap to mem0 (80.6% vs 94.4%) is primarily answerer LLM quality — mem0's own data shows a 12pp spread between GPT-5 (91.0%) and GPT-OSS-120B (89.8%) as extraction models, and the answerer matters even more. With a GPT-4o-class answerer, memini's retrieval (R@5=98.4%) and answer prompt are competitive.

Metrics

  • Recall@K — fraction of questions whose gold memory appears in the top K.
  • MRR — mean reciprocal rank of the first gold hit.
  • p50/p95 — recall latency; ingest — total ingest time.

Output is a Markdown table (stdout) plus JSON under bench/results/.

What it compares today

Three memini retrieval strategies over the same ingested store, to show the value of hybrid fusion:

System Retrieval
memini-hybrid vector + keyword, score fusion (production path)
memini-vector dense vector only
memini-keyword BM25 keyword only

memini-hybrid should never score below either single strategy.

Datasets

  • sample — committed at bench/data/sample.json, runs fully offline.
  • Normalized schema (-suite file) — {name, items:[{id,content}], questions:[{query,gold:[id]}]}.
  • LongMemEval / LoCoMo — loaders map the published JSON shapes to the normalized schema (each session/turn becomes an item; answer/evidence ids become gold). Download the datasets and pass -data.

Recall@K on LongMemEval/LoCoMo is easy to overfit — treat scores as directional.

Recall-quality scoreboard (quality_test.go)

The LongMemEval/LoCoMo suites ingest a single tier and mostly measure recall; the failures memini has actually shipped (spray) were precision failures — recall injecting irrelevant memories. TestRecallQualityScoreboard is the tier-mixed baseline both axes are judged against: durable-fact recall, episodic detail recall, and injection precision on one labeled corpus (near-duplicate durable/episodic pairs, same-template distractor durables, episodic-only topics, and confusable topics that lexically neighbour fact topics). Ranking changes are not done until this scoreboard moves in the right direction on ≥2 embedders.

go test -tags bench ./bench/ -run TestRecallQualityScoreboard -v
# add the cross-encoder rows:
MEMINI_RERANK_URL=http://127.0.0.1:8002/v1 MEMINI_RERANK_MODEL=qwen3-reranker-0.6b \
  go test -tags bench ./bench/ -run TestRecallQualityScoreboard -v

Pre-fix baseline (2026-07, additive quality composite, reserve gate ratio 0.6; reserve=0 reference in parentheses). spray = injected durables on queries with no relevant durable; inj = injected durables over all 40 queries:

Embedder · mode dur R@5 / MRR@5 epi R@5 / MRR P@3 / P@5 spray inj
qwen3-0.6b · comp 100% / .66 (70/.60) 100% / 1.0 0.725 / 0.670 0 50 (43)
qwen3-0.6b · rerank 100% / .20 (70/.14) 100% / 1.0 0.733 / 0.670 0 50 (43)
MiniLM-L6 · comp 100% / .20 (0/0) 100% / 1.0 0.742 / 0.695 0 51 (35)
MiniLM-L6 · rerank 100% / .20 (0/0) 100% / 1.0 0.775 / 0.695 0 51 (35)
nomic-v1.5 · comp 100% / .20 (0/0) 100% / 1.0 0.758 / 0.695 0 54 (40)
nomic-v1.5 · rerank 100% / .20 (0/0) 100% / 1.0 0.767 / 0.695 0 54 (40)
bge-small · comp 70% / .14 (0/0) 100% / 1.0 0.750 / 0.700 0 45 (19)
bge-small · rerank 70% / .14 (0/0) 100% / 1.0 0.775 / 0.700 0 45 (19)

What the baseline said: spray was clean, but all injection happened on low-signal detail questions — score anatomy showed durables at fused relevance 0.00 ranked #2–6 because the additive 0.2·quality bonus dwarfed the noise tail — and durable MRR sat at ~.20 on the small embedders because the reserve placed recovered facts at the window bottom. The cross-encoder additionally demotes terse "Decision: …" facts to the bottom of the window (qwen3 MRR .66→.20) while nudging P@3 up.

After the durable-ranking fix

Three coupled changes (2026-07): the composite's quality term is relevance-modulated (rel·(w_r + w_q·q̂) — a zero-relevance durable has nothing to amplify, so tier salience reorders comparable candidates instead of floating off-topic facts into weak windows; single-tier corpora like LongMemEval/LoCoMo have uniform quality and are provably order-invariant); the reserve gate ratio is re-expressed as 0.5 in the new score space (the same effective relevance bar the settled 0.6 imposed under the old composite, which carried a flat +0.2 durable floor); and the gate gains an absolute top-anchor leg (promotion also needs ≥0.4× the window's top hit — the leg that holds when the evictee is noise). Promoted durables now surface directly below the top hit instead of at the window bottom; the top hit is never displaced, so an episodic gold answer cannot be shadowed.

Embedder · mode dur R@5 / MRR@5 epi R@5 / MRR P@3 / P@5 spray inj
qwen3-0.6b · comp 100% / .75 (60/.55) 100% / 1.0 0.733 / 0.695 0 5 (5)
qwen3-0.6b · rerank 100% / .20 (60/.12) 100% / 1.0 0.758 / 0.695 0 5 (5)
MiniLM-L6 · comp 100% / .50 (0/0) 100% / 1.0 0.758 / 0.720 0 9 (9)
MiniLM-L6 · rerank 100% / .20 (0/0) 100% / 1.0 0.767 / 0.720 0 9 (9)
nomic-v1.5 · comp 100% / .50 (0/0) 100% / 1.0 0.775 / 0.735 0 14 (13)
nomic-v1.5 · rerank 100% / .20 (0/0) 100% / 1.0 0.783 / 0.735 0 14 (13)
bge-small · comp 70% / .35 (0/0) 100% / 1.0 0.783 / 0.750 0 4 (5)
bge-small · rerank 70% / .14 (0/0) 100% / 1.0 0.792 / 0.750 0 4 (5)

Reading it against the baseline:

  • Injections −72 to −91% (50/51/54/45 → 5/9/14/4), and production now adds at most 1 over the reserve=0 reference — the reserve leak is closed. What remains is base relevance (e.g. a BM25 match on a shared token), not tier spam.
  • Durable MRR on the small embedders .20 → .50 (promoted facts sit at rank 2, capped by the never-displace-the-top rule); qwen3 .66 → .75. Under the cross-encoder MRR stays ~.20: the reranker reorders the window and demotes terse facts — the known remaining defect, and it lives in the rerank tier, not the composite.
  • Durable R@5 matches the pre-fix baseline exactly (100/100/100/70). bge-small's three missing facts are blocked by the evictee leg; ratio 0.4 recovers all three for +1 injection (measured) — a candidate loosening, deliberately not taken to avoid tuning on the eval corpus.
  • Episodic recall and spray are untouched (100% / 1.0, 0).
The production window (k=3) and the rerank demotion

Every shipped turn-injection integration (hermes, openclaw, opencode, pi, openwebui) requests limit = recall_limit (default 3) and injects the entire response; the MCP tool defaults to k=10, delivered whole. Nobody truncates below k, so the cross-encoder demoting a promoted fact within the window reorders delivered memories rather than dropping any — a rerank pin/blend was scoped and dropped on that evidence. The prod k=3 scoreboard row measures the real consumer window: membership is identical to k=5 (durable R@3 100/100/100/70, both modes), worst-case rerank MRR is .333 (rank 3 of 3 delivered), and precision is slightly better (inj 4/9/11/4).

A gate-simplification probe (drop the evictee leg, keep only the 0.4 top anchor) was identical on this corpus and recovered bge-small to 100%, but leaked one spray injection on the tier-mix corpus (an off-topic durable at 0.38× the top of a flat chatter window) — the evictee-relative leg is load-bearing exactly there, so the gate keeps both legs.

Multi-hop diagnostics (multihop_test.go, entitydiag_test.go)

LoCoMo category-1 questions need all of ≥2 gold memories ("Where did Caroline move from 4 years ago?" = the moved 4 years ago turn + the home country, Sweden turn). TestMultiHopRetrievalCeiling measures whether that is a retrieval problem and what a second recall would buy (each 2nd-hop regime is the union of two k=10 result sets):

regime (277 questions, full-gold %) qwen3-0.6b MiniLM-L6
single recall 16.2% 15.9%
+ realistic 2nd hop (top-1 content) 22.7% 19.5%
+ oracle 2nd hop (known-gold content) 26.4% 22.0%
Entity edges don't bridge it (measured, phase stopped)

The planned fix was an associative edge: extract entities at write time (extract.Entities, capitalized spans, no LLM), index entity → memory ids, and let recall pull memories sharing entities with the top hits, relevance- gated like the durable reserve. TestEntityBridgeDiagnostic priced every step of that mechanism before any of it was wired, on two embedders:

  • Extraction is precise: on hand-labeled samples, P 1.00 / R 1.00 over 30 LoCoMo turns, P 1.00 / R 0.75 on scoreboard-corpus templates (the misses are lowercase technical names like "pgvector" — out of scope for a no-POS heuristic). Pinned by TestEntitiesPrecisionRecall.
  • But the entity graph has no bridging power: under a hub-safe document- frequency cap (≤10% of the conversation) only 1.3–2.6% of failed questions are fully entity-bridgeable. Without the cap it looks like 13%, but every extra link is a speaker-name hub (caroline = 119/419 memories) — expansion through a hub is spray by construction.
  • A lowercase concept vocabulary (stemmed noun-chunk proxy for the "noun phrase" tier) links 42–43% of failures in principle — but simulating the actual mechanism (promote ≤3 linked pool candidates into the window under the reserve's two-leg gate) yields no gain or a net loss: 46–49 vs 48 covered on qwen3, 39–41 vs 44 on MiniLM, with ~95% of promotions being noise. The contradiction is structural: a true multi-hop bridge has low query relevance by definition, so any relevance gate strong enough to block spray also blocks the bridges; candidates that clear the gate were already at the window's edge.
  • An entity-anchored 2nd hop (augment the query with the top hits' DF-capped terms — the union mechanism above, minus the full-content echo) reaches 19.1% / 22.4% (entities / concepts) on qwen3 and 17.3% / 20.6% on MiniLM — at or below the realistic content hop, never near the oracle.

Conclusion: on this benchmark the multi-hop gap is not an associative-edge problem a no-LLM entity index can close — 59% of missing golds are not even in the 50-deep fused pool, and what connects gold sets is either a hub person or generic lowercase concepts. Entity-aware recall was not wired into the service; extract.Entities stays as a validated building block (a later phase wants entity+attribute+value contradiction triggers, a precision-first use it fits). Re-run with:

go test -tags bench ./bench/ -run 'TestMultiHopRetrievalCeiling|TestEntityBridgeDiagnostic' -v
MEMINI_EMBED_MODEL=text-embedding-all-minilm-l6-v2-embedding MEMINI_EMBED_DIMS=384 \
  go test -tags bench ./bench/ -run 'TestMultiHopRetrievalCeiling|TestEntityBridgeDiagnostic' -v

Contradiction / update handling (contradiction_test.go)

On the no-LLM target a fact update loses to the fact it corrects, and both sit in the delivered window together. corroborateNearestAsync only grows confidence on a restatement; its contradiction mirror did not exist (the only supersede-on-contradiction was LLM-gated), and DurableScore has no recency term — so an entrenched old fact outranks a fresh contradicting write indefinitely. This phase adds a precision-first, LLM-free contradiction detector (internal/contradict, in the tradition of de Marneffe et al., ACL 2008: only surface-detectable value/polarity changes are in scope) and, on a confirmed contradiction, invalidates the stale fact the way Zep/Graphiti's temporal graph does — stamp valid_to, keep the row and its history, never delete.

Detector precision (36 authored quads — base / restatement / update / distinct, pure text, no embedder). The costly error is flagging a restatement as an update (it would downrank a live fact and lose its corroboration):

config restatement→update update recall (value / polarity) distinct→update
Default 0 / 36 29/36 (val 17/22, pol 12/14) 1 / 36

Similarity-gate routing (per embedder). The write path already pays a top-1 same-tier vector search; the detector runs on that neighbour, gated by the existing writeDedupScore (0.625). Restatement misfires are 0 at every floor on both embedders; at 0.625 distinct misfires are 0 too, and 0.625 sits below corroborateMinScore (0.70) as it must (an update diverges from its base more than a restatement):

floor qwen3-0.6b update recall qwen3 rest / dist misfire MiniLM-L6 update recall MiniLM rest / dist
0.550 29/36 (81%) 0 / 1 27/36 (75%) 0 / 1
0.625 26/36 (72%) 0 / 0 20/36 (56%) 0 / 0

Wild false positives on 500 real LongMemEval nearest-neighbour pairs (all distinct by construction, qwen3): 0.00% flagged as update at every floor (4 pairs even clear 0.625; none fire).

The harm, and why confidence alone can't fix it (TestContradictionStaleVsFresh, 12 topics; the old fact entrenched through the real corroboration path — mean effective confidence ~0.42, AccessCount ~4 — then contradicted by a fresh durable write, at the production config):

action on the stale fact stale-above-fresh (qwen3, k=3) stale-above-fresh (MiniLM, k=3)
baseline (none) 10/12 9/12
shrink confidence (×0.25 / usage-aware) 4–5/12 5–6/12
stamp valid_to (supersede / invalidate) 0/12 0/12

Confidence shrink only halves the harm: composite rank is 80% relevance / 20% quality (and fresh episodic chatter sets the quality normaliser), so zeroing the stale fact's confidence barely moves its score. Only removing it from the current-state window flips the order — hence the valid_to invalidation (reversible via Restore, still reachable via AsOf), which recall now honours in the default (non-AsOf) path.

Shipped mechanism. contradictNearestAsync (mirror of corroborateNearestAsync, internal/service/service.go) fires on fresh durable writes: top-1 durable neighbour ≥ 0.625, 24h cooldown, extract-anchored detector confirms a value/polarity change, then store.MarkContradicted stamps valid_to, shrinks confidence usage-aware (0.9·seed/usage, so the fresh write outranks it for any AccessCount), and records contradicted_by / contradicted_prev_confidence for audit and reversal. On by default; MEMINI_CONTRADICT_DOWNRANK=false is the kill-switch. Metric: memini_contradict_results_total.

No scoreboard regressionTestRecallQualityScoreboard, both embedders, composite + reranked, at k=5 and prod k=3: durable R@5/MRR, episodic 100%/1.000, spray 0, and injection counts are all identical to the pre-change baseline (the corpus carries no valid_to'd facts, so the new recall filter is inert there). Re-run with:

MEMINI_SWEEP_EMBEDDERS="http://127.0.0.1:8001/v1|text-embedding-qwen3-embedding-0.6b|1024,http://127.0.0.1:8001/v1|text-embedding-all-minilm-l6-v2-embedding|384" \
  go test -tags bench ./bench/ -run 'TestContradiction|TestRecallQualityScoreboard' -v
# add the wild-FP probe (needs the LongMemEval file):
MEMINI_LME_DATA=$PWD/bench/data/longmemeval_s_cleaned.json \
  go test -tags bench ./bench/ -run TestContradictionWildFalsePositives -v

External baselines

bench.System is the extension point. To compare against mem0, Zep/Graphiti, Letta, Cognee, agentmemory, or supermemory, implement System (Name / Ingest / Recall) over each service's API and add it to the run list in cmd/bench. These require the respective services/keys and are intentionally not vendored here.

Documentation

Overview

Package bench is a retrieval benchmark harness: it ingests a dataset of memories and scores each question's gold retrieval (Recall@K, MRR) and latency. Runs offline on the committed sample with a deterministic local embedder, or against a real endpoint and a converted LongMemEval/LoCoMo set.

Index

Constants

This section is empty.

Variables

View Source
var CodingAgentCategories = map[string]bool{
	"decision":        true,
	"convention":      true,
	"rationale":       true,
	"current-state":   true,
	"synthesis":       true,
	"temporal-update": true,
	"abstention":      true,
}

CodingAgentCategories is the fixed question-category vocabulary for the coding-agent-memory suite. The gold audit rejects any other value.

Functions

func AnswerAndJudge added in v0.5.10

func AnswerAndJudge(
	ctx context.Context, svc *service.Service, judge llm.Completer, q Question, k int, level service.ReasoningLevel,
) (correct bool, answer string, err error)

AnswerAndJudge runs the production answer path (recall + service.Answer's reader prompt; the agentic tool loop when a reasoning level is set) and grades the reply against the reference. It returns the verdict and the raw answer text so a paired comparison can list the discordant cases for inspection.

func IngestQAUpsert added in v0.5.10

func IngestQAUpsert(ctx context.Context, st store.Store, e embed.Embedder, items []Item) error

IngestQAUpsert loads items directly into the store (retrieval-only baseline): semantic tier, dated at the item time so temporal targeting can aim.

func IngestQAWrite added in v0.5.10

func IngestQAWrite(
	ctx context.Context, st store.Store, e embed.Embedder, items []Item,
	distiller llm.Distiller, consolidator llm.Consolidator,
) error

IngestQAWrite feeds items through service.Remember sequentially in dataset order (write-path corroborate/contradict is order-sensitive), clocked at each item's time so temporal targeting and valid_to invalidation see the real chronology. TTL is forced to never-expire (question dates can fall long after a session, and the bench measures answer quality, not retention). Item.Session rides along as session_id metadata so the session-echo guard and distill batching see realistic keys. distiller, when non-nil, wires LLM distill-on-write (superseding the heuristic extractor, as in production) — one completion per capture. consolidator, when non-nil, wires LLM consolidation (dedup/contradiction resolution against existing memories) at the production gate (0.3) in async mode — writes don't block on the LLM.

func JudgeSystemFor added in v0.5.10

func JudgeSystemFor(category string) string

JudgeSystemFor returns the per-category judge rubric. The coding-agent suite's "temporal-update" reuses the knowledge-update rubric (the answer must reflect the latest value) and "abstention" the decline rubric; LongMemEval's knowledge-update / temporal-reasoning / *_abs categories keep their existing mappings.

func LoadCodingAgent added in v0.5.10

func LoadCodingAgent(path string) (*Dataset, *CodingAgentMeta, error)

LoadCodingAgent reads the coding-agent dataset. Unlike the LongMemEval/LoCoMo loaders, every item MUST carry a valid RFC3339 time and every question a valid now — the suite is temporally ordered, so a missing or malformed timestamp is an error, not a silent zero. Items are returned sorted by (time, id) so write-mode ingest replays the real chronology deterministically. The returned meta carries the audit-only fields (kind, superseded_by).

func Markdown

func Markdown(results []Result) string

Markdown renders results that share a K as a comparison table, best first.

func McNemarExact added in v0.5.10

func McNemarExact(b, c int) float64

McNemarExact returns the two-sided exact-binomial p-value for a paired comparison with discordant counts b and c: b questions where arm A is correct and arm B is wrong, c where B is correct and A is wrong (concordant pairs carry no information about which arm is better). Under H0 each discordant pair is a fair coin, so the count is Binomial(n=b+c, p=0.5); the two-sided p-value is the total probability of a split at least as lopsided as observed. This exact form is used instead of the chi-square approximation because the pilot's discordant counts are small (n≈45 questions), where the approximation is unreliable.

func NamespaceOf added in v0.5.10

func NamespaceOf(group string) string

NamespaceOf exposes nsOf to the external bench_test package (the synthesis spike needs the store namespace a question group maps to).

func RerankGateMarkdown added in v0.4.13

func RerankGateMarkdown(rows []RerankGateResult, k int) string

RerankGateMarkdown renders the sweep, lowest threshold first.

func RerankMarkdown

func RerankMarkdown(results []RerankResult, k int) string

RerankMarkdown renders the RRF-vs-composite comparison, grouped by category.

func VecGateMarkdown added in v0.4.13

func VecGateMarkdown(rows []VecGateResult, k int) string

VecGateMarkdown renders the sweep, lowest threshold first.

Types

type ChatStats added in v0.5.10

type ChatStats struct {
	Completes  int64 `json:"completes"`
	ToolRounds int64 `json:"tool_rounds"`
	InTokens   int64 `json:"in_tokens_est"`
	OutTokens  int64 `json:"out_tokens_est"`
	LatencyMS  int64 `json:"latency_ms"`
}

ChatStats is a snapshot of a CountingChat's counters. LatencyMS is the wall clock spent inside the wrapped LLM calls (not recall or judging).

func (ChatStats) Add added in v0.5.10

func (s ChatStats) Add(o ChatStats) ChatStats

Add returns the field-wise sum s+o, for re-aggregating per-question deltas.

func (ChatStats) Sub added in v0.5.10

func (s ChatStats) Sub(o ChatStats) ChatStats

Sub returns the counter deltas s-o: the cost of the work done between two snapshots (one benchmark question, typically).

type CodingAgentMeta added in v0.5.10

type CodingAgentMeta struct {
	Kind         map[string]string
	SupersededBy map[string]string
}

CodingAgentMeta holds the per-item fields the harness audits but the retrieval types (Item) do not carry: item kind and the supersession pointer. Keyed by item ID.

type CountingChat added in v0.5.10

type CountingChat struct {
	// contains filtered or unexported fields
}

CountingChat wraps an llm.Client with per-direction token counters so an answer arm can report what the reader (and, in the agentic loop, each tool round) spent — the answer-path analogue of CountingDistiller. It implements both llm.Completer (single-shot reader + judge) and llm.ToolChat (the agentic loop): service.Answer type-asserts the answerer to ToolChat, so the wrapper must be passed as the WithAnswerer value or the loop silently degrades to single-shot. Token counts use the same ~4 bytes/token estimate as the retrieval metrics and cover the prompt/response text only.

func NewCountingChat added in v0.5.10

func NewCountingChat(c llm.Client) *CountingChat

NewCountingChat wraps c; if c implements llm.ToolChat the agentic loop is supported, otherwise ChatTools reports an error the loop treats as a fallback.

func (*CountingChat) ChatTools added in v0.5.10

func (c *CountingChat) ChatTools(
	ctx context.Context, system string, turns []llm.ChatTurn, tools []llm.Tool, choice llm.ToolChoice,
) (llm.ChatResult, error)

ChatTools forwards one agentic round, counting the round and payload tokens.

func (*CountingChat) Complete added in v0.5.10

func (c *CountingChat) Complete(ctx context.Context, system, user string) (string, error)

Complete forwards to the wrapped completer, counting the call and payload tokens in each direction.

func (*CountingChat) Stats added in v0.5.10

func (c *CountingChat) Stats() ChatStats

Stats returns the current counter values.

type CountingDistiller added in v0.5.9

type CountingDistiller struct {
	// contains filtered or unexported fields
}

CountingDistiller wraps a Distiller with cost/compression counters so a write-mode run can report what distill-on-write spent at ingest. Token counts use the same ~4 bytes/token estimate as the retrieval metrics and cover the JSON payloads only; the fixed distill prompt template is per-call overhead on top.

func NewCountingDistiller added in v0.5.9

func NewCountingDistiller(inner llm.Distiller) *CountingDistiller

NewCountingDistiller wraps inner with usage counters.

func (*CountingDistiller) Distill added in v0.5.9

func (c *CountingDistiller) Distill(ctx context.Context, in llm.DistillInput) ([]llm.Fact, error)

Distill forwards to the wrapped distiller, counting episodes consumed, facts produced, and estimated payload tokens in each direction.

func (*CountingDistiller) Stats added in v0.5.9

func (c *CountingDistiller) Stats() DistillStats

Stats returns the current counter values.

type Dataset

type Dataset struct {
	Name      string     `json:"name"`
	Items     []Item     `json:"items"`
	Questions []Question `json:"questions"`
}

Dataset is a normalized retrieval benchmark.

func LoadFile

func LoadFile(path string) (*Dataset, error)

LoadFile reads a dataset in memini's normalized JSON schema.

func LoadLoCoMo

func LoadLoCoMo(path string) (*Dataset, error)

LoadLoCoMo converts the published LoCoMo file into the normalized Dataset. Each conversation is its own group/namespace (dialogue ids repeat across conversations); each dialogue turn is an item, and each QA's evidence ids are its gold set. Questions without evidence (e.g. adversarial) are skipped.

func LoadLoCoMoSessions added in v0.0.4

func LoadLoCoMoSessions(path string) (*Dataset, error)

LoadLoCoMoSessions loads LoCoMo at SESSION granularity: each conversation session becomes one document (its turns concatenated), and a question's gold set is the session(s) holding its evidence turns. This matches how session-level memory systems (e.g. MemPalace) score LoCoMo, enabling an apples-to-apples comparison; LoadLoCoMo scores the harder turn granularity.

func LoadLongMemEval

func LoadLongMemEval(path string, mode DocMode) (*Dataset, error)

func Poison added in v0.0.11

func Poison(ds *Dataset, perGroup int, filler string) *Dataset

Poison returns a copy of ds with perGroup debris items added to every group that has questions — simulating a low-quality bulk import (e.g. a mem0 export of restatements) collapsed into the namespace. The debris shares one content template so a dedup pass clusters and collapses it, modelling the realistic "exports are full of near-duplicates" case. Use it to measure the Recall@K delta a poisoned store suffers, and that dedup/curation recover it.

func Sample

func Sample() (*Dataset, error)

Sample returns the committed offline sample dataset.

func SplitHoldout added in v0.5.9

func SplitHoldout(ds *Dataset, mode string) (*Dataset, error)

SplitHoldout filters longmemeval questions into a deterministic tune/held split: every 10th question by load order is "held" (50 of 500), the rest are "tune" (450). "all" (or empty) returns the dataset unchanged. Items are pruned to the surviving questions' groups, and the name is suffixed so results files don't collide across splits.

type DistillStats added in v0.5.9

type DistillStats struct {
	Calls     int64 `json:"calls"`
	Errors    int64 `json:"errors"`
	Episodes  int64 `json:"episodes"`
	Facts     int64 `json:"facts"`
	InTokens  int64 `json:"in_tokens_est"`
	OutTokens int64 `json:"out_tokens_est"`
}

DistillStats is a snapshot of a CountingDistiller's counters.

type DocMode added in v0.0.4

type DocMode string

LoadLongMemEval converts a LongMemEval file: each haystack session becomes an item, each question's answer_session_ids becomes its gold set. DocMode selects how a LongMemEval haystack session is rendered into one embedded item, for the vector-leg document-construction experiment.

const (
	// DocFull renders "role: content\n" for every turn (the production shape).
	DocFull DocMode = "full"
	// DocUserOnly renders only user turns, with no role prefixes (MemPalace's
	// raw mode: assistant turns dilute the vector leg on user-question recall).
	DocUserOnly DocMode = "user-only"
	// DocDated prefixes the full session with its date, so temporal questions
	// have a textual anchor the embedder can see.
	DocDated DocMode = "dated"
)

type IngestMode added in v0.5.9

type IngestMode string

IngestMode selects how the corpus enters the store.

const (
	// IngestUpsert writes items directly via store.Upsert (the historical
	// default): pure retrieval measurement, write-path features inert.
	IngestUpsert IngestMode = "upsert"
	// IngestWrite routes items through service.Remember, exercising the shipped
	// write path: tier classification, gates, fingerprint/write dedup,
	// corroboration, and contradiction invalidation all participate.
	IngestWrite IngestMode = "write"
)

type Item

type Item struct {
	ID      string    `json:"id"`
	Content string    `json:"content"`
	Group   string    `json:"group,omitempty"`
	Time    time.Time `json:"-"`
	// Session is a dev-session key (coding-agent suite): passed as session_id
	// metadata on write-mode ingest. Source is a provenance pointer to the
	// primary source the item was mined from (git hash, note file, plan). Both
	// are zero for the LongMemEval/LoCoMo loaders.
	Session string `json:"session,omitempty"`
	Source  string `json:"source,omitempty"`
}

Item is one memory to ingest; Group scopes it to a namespace, empty falls back to a shared default. Time, when set, is the memory's source timestamp (used to ground recency in the recency-aware re-ranking comparison).

type Question

type Question struct {
	Query    string    `json:"query"`
	Gold     []string  `json:"gold"`
	Group    string    `json:"group,omitempty"`
	Answer   string    `json:"answer,omitempty"`
	Category string    `json:"category,omitempty"`
	Now      time.Time `json:"-"`
	// GoldAll is the full evidence set for synthesis questions (answering
	// requires combining every id); recall Gold credits any-hit, GoldAll drives
	// coverage@k. Empty falls back to Gold. Provenance points at the primary
	// source the gold answer was verified against. Both are used only by the
	// coding-agent suite.
	GoldAll    []string `json:"gold_all,omitempty"`
	Provenance string   `json:"provenance,omitempty"`
}

Question is a query plus the gold memory IDs it should retrieve. Group must match its items; Answer/Category are populated for QA evaluation where available. Now, when set, is the query's reference time (e.g. the question date) — the "now" against which recency is measured.

type RecallHit added in v0.5.9

type RecallHit struct {
	IDs     []string
	Content string
}

RecallHit is one retrieved memory. IDs usually holds a single dataset item ID; write-mode ingest can merge several items into one stored memory (fingerprint/write dedup), in which case the hit carries every item ID that landed on it. Rows the write path derived itself (extract-on-write) map to no item and keep their memory ID, which never matches gold — write-mode recall is conservative by construction.

type RerankGateResult added in v0.4.13

type RerankGateResult struct {
	Threshold        float64 `json:"threshold"`
	PosRecallAtK     float64 `json:"pos_recall_at_k"`
	NegInjectionRate float64 `json:"neg_injection_rate"`
}

RerankGateResult is one cross-encoder relevance-score threshold's effect under a per-query gate: if a query's best rerank score (over its recall pool) is below the threshold, nothing relevant exists and recall returns empty. Positive = own namespace (recall must survive); negative = a foreign namespace (injection must collapse). Cross-encoders emit calibrated absolute relevance, unlike bi-encoder cosine — this measures whether that separation is real.

func RerankGateSweep added in v0.4.13

func RerankGateSweep(
	ctx context.Context, st store.Store, e embed.Embedder, ce *rerank.CrossEncoder,
	ds *Dataset, k, pool int, thresholds []float64, queryPrefix string,
) ([]RerankGateResult, error)

RerankGateSweep ingests once, then for every question reranks its recall pool (hybrid fusion + composite, top `pool`) against the query in its own namespace (positive) and in a foreign namespace (negative), recording the top rerank score and whether the gold lands in the reranked top-k. It reports the top rerank-score distribution and, per threshold, positive recall@k vs negative injection. Negatives pair each question with the next question's namespace.

type RerankResult

type RerankResult struct {
	System    string
	Category  string
	Questions int
	RecallAt1 float64
	RecallAtK float64
	MRR       float64
}

RerankResult is one ranking strategy's score over a question set.

func LLMRerankCompare added in v0.0.4

func LLMRerankCompare(
	ctx context.Context, st store.Store, e embed.Embedder, rr rerank.Reranker,
	ds *Dataset, k, fetch int, queryPrefix string,
) ([]RerankResult, error)

LLMRerankCompare measures the with-LLM read-side rerank lift on pure retrieval. For each question it builds the production candidate order (hybrid score fusion -> composite re-rank), then re-orders the top `fetch` with an LLM reranker, and scores recall@1/@k and MRR for both. The LLM tier is slow (one chat call per question), so drive it over a subset with cmd/bench -limit.

func RerankCompare

func RerankCompare(
	ctx context.Context, st store.Store, e embed.Embedder, ds *Dataset, cats []string, k int, queryPrefix string,
) ([]RerankResult, error)

RerankCompare isolates the effect of recency-aware re-ranking: it ingests ds (items carry source timestamps), then for each selected question scores the SAME fused candidate set two ways — pure RRF order vs the composite re-ranker using the question's reference time. Reports recall@1, recall@K, and MRR per category and overall, for both strategies. cats empty means all categories.

type Result

type Result struct {
	System    string  `json:"system"`
	Dataset   string  `json:"dataset"`
	K         int     `json:"k"`
	Questions int     `json:"questions"`
	RecallAtK float64 `json:"recall_at_k"`
	MRR       float64 `json:"mrr"`
	P50Millis float64 `json:"p50_ms"`
	P95Millis float64 `json:"p95_ms"`
	IngestMs  float64 `json:"ingest_ms"`
	// TokensInjectedMean is the mean estimated token count of the top-K
	// retrieved contents per question — what recall would inject into a
	// consumer's context. TokenEfficiency divides it by the corpus's total
	// estimated tokens: the cost axis of answering from memory vs full context.
	TokensInjectedMean float64            `json:"tokens_injected_mean"`
	TokenEfficiency    float64            `json:"token_efficiency"`
	PerCategory        map[string]float64 `json:"per_category,omitempty"`
}

Result is one system's score on a dataset at a given K.

func Run

func Run(ctx context.Context, sys System, ds *Dataset, ks []int) ([]Result, error)

Run ingests the dataset into a system once, then scores recall_any@K and MRR for every K in ks from a single retrieval pass (retrieving max(ks) per question). Returns one Result per K.

type System

type System interface {
	Name() string
	Ingest(ctx context.Context, items []Item) error
	Recall(ctx context.Context, group, query string, k int) ([]RecallHit, error)
}

System is a memory system under test.

func MeminiSystems

func MeminiSystems(
	st store.Store, e embed.Embedder, concurrency int, queryPrefix string, fusionAlpha float64,
	poolFactor, poolFloor int, mode IngestMode, distiller llm.Distiller,
) []System

MeminiSystems returns the hybrid, vector-only, and keyword-only retrieval strategies sharing one ingested store. queryPrefix, when non-empty, is prepended to query embeddings (hybrid and vector legs), matching MEMINI_EMBED_QUERY_PREFIX in production. fusionAlpha < 0 uses RRF; >= 0 uses convex-combination score fusion with that vector weight. poolFactor/poolFloor override hybrid recall's per-leg pool sizing (non-positive keeps defaults). mode selects direct upserts (historical default) or the production write path. distiller, non-nil with write mode, enables LLM distill-on-write (nil keeps the heuristic extractor).

func MeminiSystemsOpts added in v0.5.10

func MeminiSystemsOpts(st store.Store, e embed.Embedder, o SystemOpts) []System

MeminiSystemsOpts is MeminiSystems with the full option set, including dated ingest (SystemOpts.Dated) for temporally-ordered corpora.

type SystemOpts added in v0.5.10

type SystemOpts struct {
	Concurrency int
	QueryPrefix string
	FusionAlpha float64 // < 0 uses RRF; >= 0 uses convex score fusion
	PoolFactor  int
	PoolFloor   int
	Mode        IngestMode
	Distiller   llm.Distiller
	// Dated honors Item.Time instead of the fixed benchClock: upsert rows are
	// stamped and dated at Item.Time; write-mode ingest advances a per-item clock
	// (with ValidFrom, never-TTL, and session_id metadata), so contradiction and
	// temporal recall see the real chronology. RecallNow is the clock recall runs
	// under once ingest completes (zero = benchClock). Ignored when Dated is false.
	Dated     bool
	RecallNow time.Time
	// Chunk turns chunked embedding on (MEMINI_CHUNK_EMBED), so a run measures
	// the union of the document and chunk vector legs rather than the document
	// leg alone. ChunkScoreWeight scales the chunk leg (MEMINI_CHUNK_SCORE_WEIGHT,
	// 0 means 1): max-pooling has a length bias, and this is the knob for it, so
	// it is the one thing a benchmark is actually needed to settle.
	Chunk            bool
	ChunkCfg         chunk.Config
	ChunkScoreWeight float64
	// QueryRewrite enables LLM query expansion on the hybrid system's Recall
	// path — the A/B lever for measuring read-path LLM value. Needs Answerer.
	QueryRewrite bool
	// Answerer, when non-nil, enables LLM-backed recall features (query
	// expansion). Must implement llm.Completer.
	Answerer llm.Completer
}

SystemOpts configures MeminiSystemsOpts. The zero value reproduces MeminiSystems' historical defaults (fixed benchClock, undated ingest).

type VecGateResult added in v0.4.13

type VecGateResult struct {
	Threshold        float64 `json:"threshold"`
	PosRecallAtK     float64 `json:"pos_recall_at_k"`
	NegInjectionRate float64 `json:"neg_injection_rate"`
}

VecGateResult is one absolute-vector-score threshold's effect under a per-query semantic-relevance gate: if a query's best raw vector score (1/(1+L2)) is below the threshold, nothing relevant exists and recall returns empty. Positive = each query against its own namespace (recall must survive); negative = the same query against a foreign namespace (injection must collapse). The right default is the knee: highest threshold where PosRecallAtK is ~unchanged but NegInjectionRate has dropped.

func VecGateSweep added in v0.4.13

func VecGateSweep(
	ctx context.Context, st store.Store, e embed.Embedder, ds *Dataset,
	k int, thresholds []float64, concurrency int, queryPrefix string, fusionAlpha float64,
) ([]VecGateResult, error)

VecGateSweep ingests once, then for every question measures the top raw vector score in its own namespace (positive) and in a foreign namespace (negative), plus whether the real fused recall already retrieves the gold. It reports, per threshold, the per-query gate's effect: positive recall@k (lost only when the own-namespace top vector score falls below the gate) and negative injection rate (a foreign query passes the gate when its top vector score clears it). Negatives pair each question with the next question's namespace; group ids are unique per question, so the paired namespace never holds the answer.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL