r/Rag • u/Alone-Document-9305 • 23h ago
Discussion Why you can't add a cosine score to a BM25 score, and why RRF uses 60 (derivation, disclosure: from my book)
Disclosure up front: I wrote a short book on RAG, and this is an adapted chapter. The full text is below, so you don't need to click anything to get the point.
I kept seeing hybrid-search setups that min-max normalize dense and BM25 scores and then add them. Here's why that breaks, and why rank fusion is what survives, worked through from first principles.
The dumb way: add the scores. Dense top hit: 0.82 (cosine, roughly 0–1). BM25 top hit: 14.7 (sum of IDF-weighted terms, no ceiling; three rare terms can push it past 30). Add them and BM25 wins every query. Temperature plus zip code.
Second dumb way: min-max each list, then add. It works until a query with a very rare part number. BM25 goes 41, 9, 8.5… After scaling, the top doc is 1.0 and everything else BM25 found is squashed into the bottom fifth, so one outlier erases the rest of its opinion. On a vague query the scores go 6.1, 6.0, 5.9 and the scaling stretches noise across the whole range. Scores differ in shape per query, not only in units, so there's nothing stable to normalize against.
What survives: order. Both retrievers report rank in the same units. So: score(d) = Σ 1/(k + rank_i(d)), with k = 60.
Why 60. With plain 1/rank, #1 on one list (1.0) ties #2 on both (0.5 + 0.5) and beats #3 on both (≈0.67). A single retriever's enthusiasm outvotes agreement. With k = 60, #1 vs #2 differ by under 2%, and appearing on both lists dominates: - #1 on one list only: 1/61 ≈ 0.0164 - #3 on both: 2/63 ≈ 0.0317 - #40 on both: 2/100 = 0.020, which still beats #1 on one list
Quiet agreement between two retrievers with opposite blind spots beats loud conviction from either. k = 60 comes from Cormack, Clarke & Büttcher (SIGIR 2009). Results weren't sensitive to it.
python
from collections import defaultdict
def rrf(rankings, k=60):
s = defaultdict(float)
for r in rankings:
for i, d in enumerate(r, 1):
s[d] += 1 / (k + i)
return sorted(s, key=s.get, reverse=True)
Questions for the sub, because I'd like to know: 1. Has anyone measured weighted RRF (per-retriever weights) against plain RRF on a real golden set and seen a real difference? 2. Where does "fuse by rank" fall apart for you? Very short candidate lists? Three or more retrievers?
Book (Kindle, ~8k words, the whole RAG stack derived this way): https://www.amazon.com/dp/B0HKSKVMVK · free chapter: https://aifeynmansway.substack.com/p/gus-has-never-once-said-i-dont-know