RETRIEVAL
RELIABILITY,
MEASURED.
A retrieval evaluation harness for RAG over legal contracts. It measures four levers — chunking, hybrid retrieval, reranking and metadata filtering — against lawyer-annotated ground truth, then labels every failing query with the lever that caused it.
The point is not the number that comes out. It is that the number is falsifiable: every chunk carries (start, end) offsets back into its source document, so “did we retrieve the answer?” is decided by span arithmetic rather than by string matching or vibes. Nothing here uses an LLM as a judge.
THE BOUNDARY IS
ALREADY A DECISION.
ACROSS CHUNK BOUNDARIES BY FIXED-WIDTH CHUNKING
BY CLAUSE-AWARE CHUNKING
No ranking function can undo that: a clause split across three chunks cannot be returned whole. Neither figure requires running a retriever — both are properties of the chunker, measurable before a single query executes.
| strategy | gold spans intact in one chunk | relevant chunks per query | ceiling on precision@5 |
|---|---|---|---|
| naive (fixed 1000-char) | 43.0% | 2.70 | 0.504 |
| clause-aware | 99.1% | 1.58 | 0.312 |
The clause detector is the engineering. The obvious header regex finds zero headers in 7 of these 20 contracts: headers are indented, numbers are followed by a period, and where PDF-to-text collapsed line breaks the sections run mid-line. Inner whitespace has to be [ \t] and never \s, or every page-number footer becomes a spurious section boundary. The tuned pattern finds 1001 headers across the 20 contracts, leaving 2 with no recoverable structure, which fall back to sentence-boundary splitting rather than failing silently.
COMPLETE GROUNDING
IS THE COLUMN THAT MATTERS.
It is the share of questions where every gold clause the answer depends on was retrieved — the binary of whether a correct, fully-cited answer was possible at all. Everything to its left is a diagnostic for why it reads what it reads.
It is deliberately harsher than span coverage, and the gap between them is the reason it exists. Pooled character coverage cannot tell one fully-retrieved clause from two half-retrieved ones. On clause, coverage reads 0.735 while complete grounding is 0.660: those points of apparent success are questions answerable only from part of what they needed. In contract review that is a wrong answer carrying a confident citation.
| config | chunks | precision@5 | recall@10 | MRR | span coverage@5 | span recall@5 | complete grounding@5 |
|---|---|---|---|---|---|---|---|
| naive | 1103 | 0.276 | 0.751 | 0.607 | 0.705 | 0.701 | 0.580 |
| clause | 972 | 0.208 | 0.800 | 0.641 | 0.735 | 0.723 | 0.660 |
| clause + rerank | 972 | 0.204 | 0.827 | 0.617 | 0.709 | 0.695 | 0.560 |
| clause + bm25 | 972 | 0.196 | 0.820 | 0.616 | 0.659 | 0.662 | 0.580 |
| clause + hybrid | 972 | 0.240 | 0.857 | 0.700 | 0.823 | 0.818 | 0.700 |
| clause + hybrid + rerank | 972 | 0.208 | 0.827 | 0.605 | 0.717 | 0.705 | 0.580 |
| clause, unfiltered | 972 | 0.064 | 0.203 | 0.166 | 0.213 | 0.207 | 0.140 |
WHAT ACTUALLY
MOVED THE NUMBER.
Fusing dense and BM25 with reciprocal rank fusion beats dense alone on every column: grounding 0.660 → 0.700, coverage 0.735 → 0.823. Neither retriever wins on its own — BM25 alone scores 0.580 against dense’s 0.660 — and fusing two retrievers that each lose separately is the whole argument for fusion rather than choosing between them. The mechanism is checkable rather than asserted: fusion promoted 11 relevant chunks into the top 5 across 9 queries that dense alone ranked below 5. Net effect on grounding: 4 questions fixed, 2 broken.
The cross-encoder lifts recall@10 from 0.800 to 0.827 by promoting relevant chunks out of ranks 11–30, and drops complete grounding from 0.660 to 0.560. ms-marco-MiniLM is trained on web passage ranking, and legal clause language is far outside that distribution. Stacked on hybrid retrieval it gives back most of hybrid’s gain too, 0.700 → 0.580. It stays on this page because it was measured.
Drop the doc_id payload filter and grounding collapses from 0.660 to 0.140. 26 of 50 queries are grounded only because the filter is there, and without it 89.6% of retrieved top-5 chunks come from the wrong contract entirely. That row is a diagnostic, not a candidate configuration: CUAD’s questions are templated identically across all 510 contracts, so an unscoped search is unanswerable by construction. It is included because it is the only honest way to put a number on what filtering buys — inventing effective dates and entity names so the lever had something to filter on would produce a more impressive table and no evidence.
Grounding 0.580 → 0.660 against the fixed-width baseline, and 15 of the 21 baseline failures point straight at it. Section 01 is the mechanism; this is what it costs downstream.
AGGREGATES HIDE
THE REASON.
Every failing query is given exactly one label, assigned by rule from chunk offsets and rank positions, first match wins. Each label carries the chunk ids, offsets and rank positions that produced it, so any of them can be re-derived by hand from the run’s own output files. The dominance of chunk-severance is the retrieval-side consequence of the 43% span-integrity figure above.
| label | count | points at |
|---|---|---|
| chunk-severance | 15 | chunking |
| rank-miss | 6 | reranking |
| retrieval-miss | 0 | hybrid retrieval, embeddings |
| ceiling-bound | 0 | nothing — metric artifact |
| partial-grounding | 0 | metadata filtering, k |
Two labels never fire here, and that is a property of the chunkers rather than an accident: both tile their document contiguously with no gaps, so a query whose every relevant chunk was retrieved cannot be failing. They are kept, and their zero counts reported, because a non-zero count on a future chunker is a signal worth having.
WHY CLAUSE-AWARE
“LOSES” ON PRECISION.
Precision@5 is capped by how many chunks are relevant at all, and that count is a function of chunk size. Fixed-width chunking cuts one clause into three pieces, so three chunks count as relevant instead of one — inflating its own ceiling without retrieving one extra character of the answer. Naive chunking makes 2.70 chunks relevant per query against clause-aware’s 1.58, capping precision@5 at 0.504 versus 0.312. Measured against their own ceilings, clause-aware attains 66.7% and naive 54.8%: the strategy that looks worse on the raw metric is the more precise one.
Span coverage is reported alongside because it counts gold characters rather than chunks and so cannot be gamed by chunk size. Recall@k inherits the same flaw in its denominator and should be read the same way.
Bigger models were tried and did not win. bge-large (up to 15× the parameters of all-MiniLM-L6-v2) edged it on the clause config — MRR 0.648 vs 0.641, recall@10 0.840 vs 0.800 — but not by more than 50 queries can distinguish from noise, and it made naive chunking worse. bge-reranker-base was worse across every metric, and ms-marco-MiniLM-L-12 traded a small coverage gain for a worse MRR at 2.6× the size. The committed defaults are the smallest models tried, not merely the first ones tried.
A FRESH RUN,
UNCUT.
No API keys. Both models run locally on CPU and Qdrant runs in memory — the real client API, no Docker. First run downloads ~200 MB of weights; the seven-configuration eval then takes under ten minutes on a laptop CPU. 87 tests load no models and finish in about 25 seconds. The eval set is committed, so a fresh clone reproduces the table without the download.
WHAT THIS
DOES NOT SHOW.
- 01
Precision@k is not comparable across chunking strategies — its denominator is a function of chunk size.
- 02
This measures retrieval only. Generation faithfulness, hallucination and answer relevance are a separate layer.
- 03
The sample is small: 20 contracts, 50 queries, 1–4 per clause type, no confidence intervals. The 43% / 99.1% gap is large enough to be real; the 0.607 / 0.641 MRR gap is not.
- 04
The 50% grounding threshold is a convention, not a measurement. It is fixed across every config so it cannot be tuned, but a different threshold would move every grounding figure.
- 05
The failure taxonomy attributes, it does not prove. Only a run with the lever changed establishes that changing it fixes the query.
- 06
Retrieval is scoped to the correct document, so this measures clause localization within a contract — not document routing across a corpus.
WHY THE CORPUS
IS PUBLIC.
This is the method I used rebuilding retrieval on a production document-AI system — clause-aware chunking and cross-encoder reranking against an eval set where a domain expert labeled the correct source passage for every query. That system is under NDA, so the harness above reproduces the method on public documents instead: CUAD v1, CC BY 4.0, ships 510 commercial contracts with 13,000+ lawyer-annotated clause spans. The ground truth is labeled by someone else, not asserted by me.
MEASURE IT
BEFORE YOU
REBUILD IT.
Five business days on the pipeline you already have. You leave with a reproducible golden set, a baseline your team can re-run, one reason per failing query, and a fix backlog whose entries carry measured deltas instead of opinions.
- 01
A golden set located, not guessed
40 labeled queries with relevance judgments, approved by you. The adapter locates each answer by character offset in your own documents; a gold text matching zero times, or more than once, raises an error naming the document, the text and the match count rather than picking one for you.
- 02
A baseline on your current configuration
Precision@5, recall@10, MRR, span coverage@5, span recall@5 and complete grounding@5 — the share of questions where every clause the answer depends on actually came back. The other columns are diagnostics for why that number is what it is.
- 03
One reason per failing query
Every failure carries exactly one of five labels, assigned by rule from chunk offsets and rank positions, with the chunk ids and ranks that produced it. No model judges them, so every label can be re-derived by hand from the run's own output.
- 04
A fix backlog with measured deltas
Ranked by how many failing questions each lever accounts for. Which lever a delta belongs to is derived by diffing two configurations, so a change is never attributed to a lever that was not moved — and a lever no pair of runs isolated reads not measured rather than carrying an estimate.
- 05
The harness, and a readout
Four generated files: the client report, the developer report, per-query metrics with the full 30-deep ranking and every missed span, and the failure taxonomy with its evidence. Plus setup notes, a 30-minute readout and one consolidated revision.
No number appears in it that was not produced by a run on your corpus.
THE SCOPE CAPS ARE ENFORCED IN CODE — THE ADAPTER EXITS NON-ZERO WITH THE ACTUAL COUNTS. EXTRACTION IS RECORDED: THE EXTRACTOR, ITS VERSION AND A SHA256 OF EVERY DOCUMENT, BECAUSE OFFSETS ARE ONLY VALID AGAINST THE TEXT THAT PRODUCED THEM.Production implementation, model fine-tuning, UI work, ongoing monitoring.
Generation quality, latency and cost. This measures what retrieval puts in front of the model, not what the model does with it.