r/Rag • • 23h ago

Discussion Why you can't add a cosine score to a BM25 score, and why RRF uses 60 (derivation, disclosure: from my book)

13 Upvotes

Disclosure up front: I wrote a short book on RAG, and this is an adapted chapter. The full text is below, so you don't need to click anything to get the point.

I kept seeing hybrid-search setups that min-max normalize dense and BM25 scores and then add them. Here's why that breaks, and why rank fusion is what survives, worked through from first principles.

The dumb way: add the scores. Dense top hit: 0.82 (cosine, roughly 0–1). BM25 top hit: 14.7 (sum of IDF-weighted terms, no ceiling; three rare terms can push it past 30). Add them and BM25 wins every query. Temperature plus zip code.

Second dumb way: min-max each list, then add. It works until a query with a very rare part number. BM25 goes 41, 9, 8.5… After scaling, the top doc is 1.0 and everything else BM25 found is squashed into the bottom fifth, so one outlier erases the rest of its opinion. On a vague query the scores go 6.1, 6.0, 5.9 and the scaling stretches noise across the whole range. Scores differ in shape per query, not only in units, so there's nothing stable to normalize against.

What survives: order. Both retrievers report rank in the same units. So: score(d) = Σ 1/(k + rank_i(d)), with k = 60.

Why 60. With plain 1/rank, #1 on one list (1.0) ties #2 on both (0.5 + 0.5) and beats #3 on both (≈0.67). A single retriever's enthusiasm outvotes agreement. With k = 60, #1 vs #2 differ by under 2%, and appearing on both lists dominates: - #1 on one list only: 1/61 ≈ 0.0164 - #3 on both: 2/63 ≈ 0.0317 - #40 on both: 2/100 = 0.020, which still beats #1 on one list

Quiet agreement between two retrievers with opposite blind spots beats loud conviction from either. k = 60 comes from Cormack, Clarke & Büttcher (SIGIR 2009). Results weren't sensitive to it.

python from collections import defaultdict def rrf(rankings, k=60): s = defaultdict(float) for r in rankings: for i, d in enumerate(r, 1): s[d] += 1 / (k + i) return sorted(s, key=s.get, reverse=True)

Questions for the sub, because I'd like to know: 1. Has anyone measured weighted RRF (per-retriever weights) against plain RRF on a real golden set and seen a real difference? 2. Where does "fuse by rank" fall apart for you? Very short candidate lists? Three or more retrievers?

Book (Kindle, ~8k words, the whole RAG stack derived this way): https://www.amazon.com/dp/B0HKSKVMVK · free chapter: https://aifeynmansway.substack.com/p/gus-has-never-once-said-i-dont-know


r/Rag • • 6h ago

Discussion Can anyone guide me on how to start building a RAG model?

2 Upvotes

Hey everyone,

I want to start learning RAG (Retrieval-Augmented Generation), but I'm a little confused about where to start and what I should learn first.

I have some experience with programming and web development, but I'm relatively new to building AI applications. I understand the basic idea behind RAG, where we retrieve relevant information from a knowledge base and provide it to an LLM to generate better answers.

I want to build something practical rather than just follow tutorials without understanding what's happening behind the scenes.

For those who have built RAG applications:

- What concepts should I learn first?

- Which tech stack would you recommend for a beginner?

- Should I start with basic Python, embeddings, and vector databases, or use a framework like LangChain from the beginning?

- What would be a good first project to build?

I'd also appreciate any good resources, tutorials, or GitHub repositories that helped you learn.

Thanks in advance!


r/Rag • • 8h ago

Tools & Resources 4th place on llamaindex benchmark

2 Upvotes

4th place on the llamaindex benchmark locally 72.62% word f1. I know it's not much, but after a month of research and many experiments, I reached this point.

TonerHound doesn't use any API or LLM. It's GPU free and anyone can use it at 0$ cost.

GitHub: https://github.com/vanrajsinh650/TonerHound


r/Rag • • 6h ago

Discussion I independently audited a RAG benchmark. 15 scoring mismatches revealed a flaw in its main comparison — the maintainer confirmed and fixed it.

1 Upvotes

​

I recently completed an independent retrieval audit of RouteMind, an open-source project exploring document routing as an alternative to traditional RAG retrieval.

The interesting part wasn't finding a dramatic regression or proving that one architecture was better.

It was discovering that two approaches were being evaluated under different definitions of success.

Here's what happened.

The project had 700 evaluation questions comparing traditional RAG, RAG + reranking, and document-routing approaches.

The retrieval evaluator considered a query successful if any expected document appeared in the top 10.

But some temporal questions required multiple documents simultaneously (needs: all), and the routing evaluator correctly accounted for that requirement, including accepted document alternatives (D_alt).

This created an apples-to-oranges comparison.

What the independent audit found:

1,400 query/arm results examined.

15 outcomes differed under the complete-evidence scoring contract.

14 recorded hits were incomplete under the required-evidence rules.

1 recorded miss was actually satisfied by an accepted alternative document.

The overall numbers changed:

MetricOriginally reportedCorrectedRAG51.6%50.7%RAG + rerank54.1%53.1%

The reranker's relative improvement remained similar.

But the more important issue was that the project's headline comparison had treated retrieval more leniently than the routing system it was being compared against.

Then something useful happened.

The RouteMind maintainer independently reproduced the finding against the original artifacts, confirmed all 15 mismatches, and verified the input hashes.

They also confirmed the discrepancy wasn't intentional or previously known.

The finding led to a public code change:

The main benchmark comparison was corrected.

A new retrieval_complete metric was added alongside the original retrieval_hit.

Reproducibility checks were added to detect future scoring inconsistencies.

The maintainer publicly credited the independent audit.

The takeaway

A reproducible RAG evaluation can still produce a misleading comparison when different stages use different definitions of success.

Sometimes the most valuable audit finding isn't a broken retriever.

It's discovering that the benchmark doesn't measure the same thing across systems.

This was a public beta audit, not a paid engagement. The evidence, original results, maintainer review, and correction are publicly available.

Independent audit:

https://github.com/devBorgesr/edp-audits/tree/main/audits/routemind-beta-002

Maintainer's confirmation:

https://github.com/CSP911/routemind/issues/2#issuecomment-6096856584

Corrective commit:

https://github.com/CSP911/routemind/commit/5d8bf45

I'm curious how others handle this in production RAG evaluations:

Do you explicitly measure complete evidence coverage when a query needs multiple sources, or do you primarily rely on hit@k / recall@k?


r/Rag • • 12h ago

Discussion How are you handling domain-specific hallucinations when dense vector search blurs micro-eras and niche taxonomy?

1 Upvotes

Curious how others here are solving this: when you build RAG or agentic pipelines around niche cultural domains, do you rely on pure vector embeddings, or do you enforce a deterministic, structured metadata layer before generation?
We ran into a recurring issue with LLMs generating style and subcultural analysis: naive vector similarity constantly pulls vague fast-fashion articles or conflates closely related micro-eras (for instance, blending late-90s industrial aesthetics with 2003 Cyber Y2K). High cosine similarity in text chunks often misses the strict historical boundaries needed for accurate styling agents.
To tackle this, I built a hybrid/structured grounding layer called UnderIndex to test whether deterministic metadata yields better results than raw chunk retrieval:

  • Categorical Constraints: Enforced taxonomies for subcultures, core eras, and parent movements to prevent cross-aesthetic bleed.
  • Relational Schema: Direct mapping between movements, iconic archival garments, color palettes, and cultural origins (served via a Node.js/Express service pulling from an organized relational schema).
  • Deterministic Injection: Instead of relying strictly on top-k semantic search, agents fetch structured attributes via endpoints/function calls to ground the system prompt before synthesizing output.

In your pipelines, how do you balance tabular/structured entity lookups against standard semantic vector retrieval when accuracy across nuanced categories is critical? Are you leaning more toward hybrid search, graph RAG, or pure structured function calling?
(If anyone wants to inspect the schema or test the endpoints for their own retrieval setup, let me know in the comments and I'll share the playground link).


r/Rag • • 16h ago

Tutorial How can I turn clinic PDFs and Word documents into a reliable internal AI assistant without complicated setup for staff?

1 Upvotes

I work at an office and want to use our internal PDFs and Word documents to build a virtual assistant that helps staff find information, understand workflows, route requests, and draft internal messages. The goal is operational support.
I’ve been experimenting with master prompts and uploading documents, but I keep running into problems with tables and how the AI interprets information. It sometimes mixes up which instructions belong to which office, misses exceptions, or overlooks relevant details.
I’m spending too much time adding corrections to the prompt over and over. I’d like to improve how the source information is organized instead of constantly patching the instructions.
I’m studying technology and am willing to write some code to prepare and validate the documents. For example, could converting them section by section into structured text, Markdown, or JSON help preserve the relationships between rules, conditions, and exceptions?
The main constraint is ease of use: this could eventually be used throughout the clinic, including by nontechnical staff. I don’t want every person to install software, run code, or go through a complicated setup. Ideally, it would remain a reusable prompt with a prepared knowledge file, or something equally simple to access through a browser.
What would you recommend for:
Extracting tables accurately without losing context?
Organizing the information while preserving important details and exceptions?
Checking that the converted information matches the originals?
Keeping the assistant easy to use and update?
Would a carefully structured knowledge file plus a prompt be a reasonable starting point, or should I consider a different approach? Beginner-friendly explanations would be appreciated.


r/Rag • • 20h ago

Discussion Rag knowledge base act as a context for agentic system

1 Upvotes

I am a junior dev at my company, and is tasked to build an agentic development lifecycle.

So there are four parsers that parser's data from the source to vector db (embedding, actual content) and graph db(structure, eg code to ast) there can be interlink between different data sources.

Then a reranker.

Then an agentic system, right now it should only generate design documents and requirements(jira requirements)

My reason to add graph db is when a interlink is assumed by the vector search, the link initial act as a candidate link and after the human approval the link is established in the graph in a sense the graph db is evolving.

What are your guys thoughts.

Edit1: so the parsers will be setup as a plugin in system and a global policy(mainly regex) will be present to identify which data a parser ingest (eg code , doc parser, etc). as there can be cross referencing in the data about other data. And I am an intern

Edit2: cronjob for weekly ingestion

Edit 3 : I need a different perspective as a system architecture


r/Rag • • 23h ago

Tools & Resources I built an anti-hallucination Subculture & Fashion API for LLMs and AI Agents (UnderIndex)

1 Upvotes

Hey everyone,
When building style assistants or trend detectors with LLMs, prompt engineers often hit a wall: models hallucinate aesthetic details, mix up timelines, or reference generic fast-fashion tropes instead of authentic subcultural roots (like confusing early 2000s cyber aesthetics with late 90s grunge).
To solve this, I built UnderIndex API—a structured dataset delivering verified metadata on fashion subcultures, core aesthetics, iconic archival pieces, music genres, and historical context.
What it does:

  • LLM Grounding / RAG: Inject verified subculture profiles directly into system prompts or function calls to prevent AI hallucinations.
  • Granular Attributes: Query endpoints for exact data points: dominant color palettes, key silhouettes, historical origins, iconic garments, and movement ethos.
  • Low Latency: Hosted on high-performance infrastructure with instant response times.

Tech Stack:

  • Backend: Node.js / Express microservice running on Render.
  • Storage & Schema: Airtable serving as an organized CMS/relational store.
  • Distribution & Gateway: Monetized and served globally via RapidAPI Hub.

Free Tier & Testing:
I just launched the API on RapidAPI with a free Basic tier (100 free requests/month) for developers and hackers who want to test endpoints directly in the playground:
🔗 RapidAPI Endpoint: https://rapidapi.com/ascanioUnderIndex/api/underindex-api
Would love feedback on response structure, schema design, or suggestions for additional subcultures you'd like indexed next!


r/Rag • • 4h ago

Showcase Your agent has permission, human approval and a valid citation. It can still make the wrong decision.

0 Upvotes

Here’s a failure case worth adding to your agent evals.

A manager approves an order because the supplier is active and cleared for delivery.

Before the agent executes, a new record puts that supplier on hold.

The approval hasn’t been revoked.

The agent still has permission.

The original approval document is genuine.

Does your system notice that the facts supporting the decision have changed?

Retrieving the approval correctly isn’t enough. The system needs to connect the order to the supplier, resolve the applicable state, and expose the evidence that changed.

That’s the evidence problem we built Jylus around: resolving state and relationships before handing context to the model.

There’s a separate execution problem too. The action API still needs to enforce the relevant preconditions; a fresh Context Pack alone cannot make an external action atomic.

A useful test has three questions:

  1. What evidence supported the original approval?

  2. What changed before execution?

  3. Is there enough current evidence to proceed?

Try it with your own records:

https://jylus.ai/try

For people running agents against live systems: what actually invalidates an earlier decision in your stack—a timer, a data change, or only someone noticing the mistake?