TOOL VERDICT

A reranker gained 14 points on Hit@10. The answers barely moved.

A MedCPT cross-encoder lifted Hit@10 by 14 points on 1,000 clinical QA pairs, and answer correctness moved under 4. The authors' own words: retrieval gains do not translate proportionally into answer gains.

2026-10-08

Here is a number that should make every RAG team uncomfortable. A reranker gained 14 points on Hit@10, and the answers barely moved. The retrieval metric screamed success. The thing users actually read barely changed.

The test comes from a new preprint, arXiv 2610.01324, published October 1. The setup: a MedCPT cross-encoder reranking over 1,000 clinical QA pairs. PubMedBERT dense plus BM25, fused with weighted RRF, feeding the cross-encoder. Qwen3-8B generates the answers, and a local judge scores correctness. Clean before/after, same run, both metrics.

The before/after pair

+14.0
points on Hit@10: 46.6% to 60.6%. MRR moved 0.2371 to 0.3252. The reranker reordered what retrieval found, and 14 more points of queries had the right chunk in the top 10.

That is a real retrieval win. The reranker did its job. The chunks reached the top 10.

Now the answers.

+3.8
points on answer correctness: 44.8% to 48.6%, on the same run. The better-ranked chunks reached the model, and correctness barely moved.

Same run, same chunks, same model. Retrieval up 14. Answers up under 4.

The authors said it first

This is not our spin. The authors’ own sentence: gains in retrieval do not translate proportionally into gains in answer correctness. They published the caveat alongside the win, which is more honesty than most benchmark marketing manages.

And the caveats matter for how far you can carry this result. It is a preprint, one clinical domain, Qwen3-8B answering, a local judge scoring. An 8B model on clinical notes is not your stack. What transfers is the test, not the 3.8.

But notice what the paper does not say. It does not say where the signal died. Was it the 8B model failing to use good chunks? The local judge failing to credit good answers? The chunks themselves, present but insufficient? No retrieval metric can answer that. Hit@10 was 60.6% and the answers still stalled. Only scoring the answers can tell you which stage is the bottleneck.

The verdict: score answers, not only retrieval

A reranker only reorders what retrieval found. That is its ceiling, and it is worth knowing where the ceiling is. Compare recall at your k against recall at your pool size. The gap between them is the most any reranker can ever add. If your first stage never fetches the right chunk, no reranker in the world puts it in the top 10.

Databricks’ rule of thumb is the practical version: if the answer is in the top 50 but not the top 10, test a reranker. If it is not in the top 50, fix retrieval first.

The temptation after a result like this is to conclude rerankers do not work. That is the wrong read. The right read is narrower and more useful: retrieval metrics are not answer metrics, and a pipeline scored only on retrieval is a pipeline whose real bottleneck is invisible. Score the answers. Then decide.

This post started as an Instagram post →

Discussion

Talk it through

Argue with us on Instagram.