Documentation
¶
Overview ¶
Command qa measures end-to-end QA accuracy (LLM-judged) over a benchmark conversation corpus: ingest into memini (direct upserts or the production write path), answer each question with the shipped service.Answer, and grade against the reference with per-category judge rubrics. This is the answer-quality companion to cmd/bench's retrieval scores.
Click to show internal directories.
Click to hide internal directories.