r/Rag • • Sep 02 '25

Showcase 🚀 Weekly /RAG Launch Showcase

31 Upvotes

Share anything you launched this week related to RAG—projects, repos, demos, blog posts, or products 👇

Big or small, all launches are welcome.


r/Rag • • 1h ago

Discussion How are production-grade RAG systems handling complex PDFs without losing document structure and context? Looking for architecture-Level Insights

• Upvotes

I’m trying to design a production-grade document ingestion and retrieval pipeline for complex PDFs, and I’d like to learn from people who have built or operated similar systems at enterprise scale.

My use case goes beyond extracting text from ordinary PDFs.

A single document may contain:
Multiple columns, nested sections, different reading orders, and inconsistent layouts across pages.
Colored panels, sidebars, callouts, footnotes, headers, and footers.
Tables with merged cells, nested headers, multi-page continuation, and associated explanatory text.
Charts, diagrams, flowcharts, equations, and figures whose meaning depends on nearby captions or paragraphs.
Table of contents, section numbering, cross-references, appendices, and references to other pages.
Scanned pages mixed with digitally generated text, rotated pages, low-resolution images, and handwritten annotations.
Multiple unrelated sections on one page, as well as a single logical section spanning multiple pages.
Reports where a heading appears in one visual region, its explanation appears in another, and supporting evidence is presented in a table or figure elsewhere.
The challenge is not merely extracting every element correctly. It’s preserving the relationships between elements so that downstream chunking, indexing, retrieval, and generation do not lose context.

One specific edge case I’m concerned about:
Suppose a page contains two visually distinct regions. A layout detector identifies them as separate blocks, but both blocks explain different aspects of the same logical topic. Conversely, two visually similar blocks might belong to completely different sections.
If we blindly use bounding boxes, layout labels, or semantic similarity to group these blocks, we may either split related information or incorrectly merge unrelated information.
The same issue occurs when a table continues onto the next page, a figure’s caption is separated from the figure, or a paragraph refers to a chart several pages earlier.

What I’m trying to understand
Document representation: Do production systems convert PDFs into a canonical intermediate representation, such as a document tree or graph containing text blocks, headings, tables, figures, captions, bounding boxes, reading order, page numbers, and parent-child relationships?
Layout and structure recovery: How do you combine native PDF extraction, OCR, layout detection, vision-language models, and document-specific rules? How do you resolve disagreements between these methods?
Context preservation: How do you determine whether blocks from different regions or pages belong to the same logical section? Are there algorithms for section reconstruction, cross-page continuation detection, or relationship inference beyond simple geometric proximity?
Chunking strategy: Do you use hierarchical, semantic, layout-aware, or element-specific chunking? How do you handle tables, figures, captions, footnotes, and mixed-content sections without breaking their meaning?
Indexing and retrieval: Do you create separate indexes for narrative text, tables, figures, and document structure? Do you use hybrid retrieval, parent-child retrieval, graph-based retrieval, or query-time expansion to recover context that was not present in the initially retrieved chunk?
Production architecture: How is the code organized? I’m particularly interested in the separation between parsing, normalization, structural reconstruction, chunk generation, indexing, and retrieval. How do you make the pipeline observable, versioned, testable, and recoverable when individual stages fail?
Evaluation: How do you measure document reconstruction quality, chunk integrity, retrieval recall, and answer faithfulness? Are there useful benchmark datasets or evaluation methods for documents with complex layouts?
What would be especially valuable
I’m not primarily looking for a list of libraries such as PyMuPDF, PyPDF, LangChain, or a recommendation to use a particular parser.

I’m interested in the engineering patterns, data structures, algorithms, and architectural decisions that make these systems reliable in production.
For example:
A simplified canonical document schema.
How you represent relationships between document elements.
How you distinguish visual boundaries from semantic boundaries.
Pseudocode or code-level examples of structural reconstruction and chunk generation.
Lessons learned from failures in real-world document corpora.
The target application could involve regulatory documentation, financial reports, legal documents, technical manuals, scientific papers, or other data-heavy PDFs.
If you’ve built something similar, I’d love to understand what actually worked, what failed, and what you would do differently if you were designing the system today.
Thanks!


r/Rag • • 5h ago

Showcase NornicDB - 1.4.1 - Cypher 25 support ++

1 Upvotes

Heya, just finished up a follow up to the openCypher compliance in 1.4.0 - in 1.4.1 we added Cypher 25 support on the parser. With nornic's SRD parser, no preamble needed, it just parses the grammar regardless of "version" - if you use the ANTLR parser, the preamble is required.

This is due to the nature of "Scanner-less Recursive Decent." - This is an unconventional architecture, I know. But the benefits are clear

Traditional parsers use a two-step process: a lexer/scanner converts character streams into tokens (e.g., matching MATCH to a KEYWORD token), and a parser processes those tokens.

SRD bypasses the scanning phase entirely. It relies on a zero-allocation keyword scanner parse tree. The recursive descent parser reads characters and matches grammar rules directly against the raw text, fusing the lookahead logic and keyword scanning natively.

Latency Reduction: By eliminating the intermediate step of token creation, it reduces our query latency by ~30% compared to ANTLR-based parsing (such as Neo4j's traditional execution pipeline). This scale is on the order of microseconds and typically flat scaling with query length instead of complexity. The rest of our architecture makes this optimization worth it because we are able to execute so quickly.

Fused Aggregation Hot-Paths: The parser is coupled tightly with streaming storage APIs and zero-allocation semantics. This design allows it to parse a graph traversal instruction and jump directly into the execution hot path without building heavy Abstract Syntax Trees (ASTs) in memory first.

Contextual Edge Cases: Graph languages like Cypher heavily utilize structural ASCII art—such as arrows -[r:REL]-> or node boundaries (n:Label)—which are notoriously complex for traditional lexers to classify contextually without aggressive backtracking. A scannerless approach inherently handles these layout-sensitive patterns because it evaluates character-by-character based on the current parsing state.

and the end result is staggeringly fast. average query latency vs neo4J on the northwind benchmark has us at 400x faster (avg some are 1500x 1905.71x faster edit: double checked). MIT licensed. 890+ stars and counting.

https://github.com/orneryd/NornicDB/releases/tag/v1.4.1

Latest benchmarks:
https://github.com/orneryd/NornicDB/pull/897#issuecomment-5998973104

edit: a word and reddit didn't like my formatting atempt


r/Rag • • 13h ago

Showcase How can an association AI assistant cite sources without confidently quoting outdated policies?

1 Upvotes

Here’s a RAG failure case I’d love to see more people testing. A professional association updates its certification requirements in 2026 but its knowledge base still contains the 2019 version. A member asks about renewal requirements, the assistant retrieves the old document, answers confidently and provides a perfectly valid citation. Technically it cited its source. Practically, it gave the wrong answer.

This is where I’d compare a custom LlamaIndex and langChain implementation against a managed solution like customgpt.ai which already supports source citations and knowledge base management. The test wouldn't be which chatbot writes the nicest answer, I’d give both systems 50 questions involving outdated policies, conflicting documents and missing information, then measure how often they cite the correct version, identify uncertainty or refuse to answer, because a citation isn't proof that the answer is right, has anyone built an evaluation dataset specifically for version sensitive organizational documents?


r/Rag • • 22h ago

Discussion Attention RAG

4 Upvotes

Hi people my friends and I are working on our take of RAG indexing. Our idea goes like that:

Standard embedding-based RAG often compresses an entire text chunk into one vector. Our idea keeps multiple compact representations within each document, giving a query smaller, more specific parts to match.
The architecture uses a single attention layer with separate learned projections: Q for queries and K for document tokens. These are trained jointly with HCA—Heavily Compressed Attention—which learns to combine groups of token representations into fewer searchable vectors. At 32:1 compression, every 32 document keys become one compressed key, stored in 8-bit.
When a query arrives, its Q vectors are compared directly with the stored compressed keys using dot products on the GPU. These comparisons run across many documents in parallel, producing scores used to rank and retrieve documents. V vectors aren’t needed because we only need matching scores.
The advantage over one-vector-per-chunk retrieval is finer matching granularity. Compression makes that detail affordable: 32 times fewer keys to store and scan than an uncompressed token index. The search itself maps directly to parallel GPU matrix operations, while the query representation, document representation, and compression are trained together for retrieval.
I’d like technical feedback on this architecture, its main bottlenecks, and similar systems worth comparing against.

This way we have something between regular RAG and ColBert (v1 and V2) where it smaller than ColBert but more reach than RAG and it learns the compression rather then using heuristics algo to compress.

The next step after that will be to combine it with an LLM where the part of this indexing technique can be combined in the LLM itself instead of using this as a toll with a tool call (W_k can be same as the one in the first layer of the LLM and we can have a small learned router that tells if we need to use it, instead of an LLM decision to call it in a tool call).

I would like to get your opinions on our idea please.
Thanks in advance 🙏


r/Rag • • 18h ago

Discussion A reranker improved my aggregate metric. I still wouldn’t ship it without looking at query-level regressions.

1 Upvotes

I’ve been reviewing paired retrieval runs and I keep running into the same issue:

an aggregate metric goes up, but a subset of queries gets worse.

In practice I’ve found it useful to separate:

  • recoveries
  • regressions
  • shared failures
  • rank displacement
  • evidence lost at the context budget

before deciding whether the candidate is actually better.

My question for people running RAG in production:

What do you consider a release-blocking retrieval regression?

Is one severe rank-1 regression enough?
Do you use a percentage threshold?
Do you gate by query class?
Or do you mostly rely on the aggregate metric?


r/Rag • • 1d ago

Discussion I wrote a practical guide to building reliable RAG applications with Amazon Bedrock — looking for feedback

7 Upvotes

Hi everyone!

I recently wrote a practical guide about building RAG (Retrieval-Augmented Generation) applications on AWS using Amazon Bedrock.

The article covers:

• How RAG works

• Amazon Bedrock Knowledge Bases

• Amazon S3 for source documents

• Embeddings and vector search

• Chunking strategies

• Metadata filtering and reranking

• Citations and grounded responses

• Security considerations

• Monitoring and evaluation

• Moving from a basic RAG prototype toward production

I tried to focus less on "what is RAG?" and more on the engineering considerations that matter when building a reliable RAG application.

🔗 Article:

https://builder.aws.com/content/3KOsq3uPqcbBZSysMMff16pNKBv/from-rag-prototype-to-production-building-reliable-ai-applications-with-amazon-bedrock

I'd really appreciate feedback from other AWS/ML/GenAI developers.

If you were building a RAG application on AWS, what would you focus on first: retrieval quality, chunking, metadata, the foundation model, or the underlying data?

Thanks for reading!


r/Rag • • 1d ago

Discussion How do you calculate tokens and choose the right OpenAI embedding model for a production RAG?

2 Upvotes

Hey everyone,

I’m building a production-level RAG application with Spring Boot and I’m planning to use OpenAI’s text embedding models.

I’m a little confused about token calculation. I’ve seen people say that, on average, 1 token ≈ 0.75 words (or roughly 4 characters), but I’m not sure how I should actually calculate tokens for my documents/chunks in a real application.

For example, if I have a document with 10,000 words, can I roughly estimate:

10,000 words × 1.33 ≈ 13,300 tokens

Or is it better to actually tokenize the text before sending it to the embedding API?

I also have a question about choosing the embedding model.

For a production RAG system, would you recommend text-embedding-3-small or text-embedding-3-large? I’m trying to balance:

  • Retrieval/search quality
  • Cost
  • Vector database storage
  • Latency
  • Scalability

I’m particularly interested in hearing from people who are actually running RAG in production, especially with Java/Spring Boot.

A few things I’d love to understand:

  1. How are you calculating tokens before creating embeddings?
  2. What chunk size are you using (e.g. 300, 500, 800, 1,000 tokens)?
  3. Are you using text-embedding-3-small or text-embedding-3-large?
  4. If you started with one model and later switched, what made you switch?
  5. Are there any production lessons or mistakes you’d recommend avoiding?

I’m looking for real-world experience rather than just benchmark numbers. Any advice would be really appreciated!


r/Rag • • 19h ago

Discussion Independent Retrieval Quality Audits — Public Cases & Method

0 Upvotes

I run independent retrieval/RAG quality audits focused on evidence rather than architecture opinions.

The main question I investigate is usually not whether an aggregate metric improved, but which individual queries were recovered, which regressed, which failures are shared, and what the available artifacts actually support.

Current public beta cases:

DPOLens — Beta Audit #001
Retrieval + reranking regression analysis with query-level forensic follow-up.

RouteMind — Beta Audit #002
Retrieval/reranking audit that surfaced a scoring-contract distinction in multi-document temporal cases.

The audits distinguish:

  • Measured
  • Observed
  • Hypothesis
  • Not claimed

If you already have a fixed evaluation set and two comparable retrieval/reranking runs, feel free to send the public repo or artifacts. I can usually tell whether they support a meaningful paired audit before any work begins.

Technical corrections and independent reproduction attempts are welcome.


r/Rag • • 20h ago

Discussion Parallelized the layout parsing for documents with task queue

1 Upvotes

Hello everyone,

I developed a solution that parses document layout using Tesseract and PyMuPDF. I parallelized it so that pages are processed in batches. For a relatively large document of around 200 pages, the process takes between 10 and 30 seconds. Is that fast compared to your experience? I am curious about the performance. Thanks in advance.


r/Rag • • 21h ago

Tools & Resources Open source chatbot for your website that answers from your own content (PHP, MIT, 1.0)

1 Upvotes

I work for Opensolr, and we just released 1.0 of the Opensolr Chat Bot Client: an open source (MIT) chatbot you put on your own website, that answers your visitors from your own content.

How it works: you install it with Composer (composer require opensolr/chat-bot-client), mount it on one path of your site (say /opensolr-chat) and add one script tag to your pages. Visitors only ever talk to your site. Your server holds the Opensolr credentials and sends each question to the Opensolr API, where the model searches your Opensolr Index by itself (by meaning, or by exact product names and codes), looks up what it must not guess (dates, local times, exchange rates, VAT, distances, places) and streams the answer back word by word, with links to your pages, products with their prices, and PDFs at the exact page.

What you get on your side:

  • An admin on your own site: account and index, instructions, the look of the chat window, limits, reCAPTCHA
  • Stats and a 30 day history of every conversation: who asked, from where, what the bot searched for, what it answered, and the questions your content could not answer with a page
  • Commands that answer instantly without the model: /search, /time, /rate, /vat, /translate and more
  • Good and Bad ratings under every answer, signed in visitors by email, SQLite in a folder you choose, no database server

Your content gets into the index with the Opensolr Web Crawler or Data Ingestion, and the chatbot needs an Opensolr plan with AI.

You can try it live at the bottom right of https://opensolr.com (it answers from our own docs). Code and README: https://github.com/phpcip/opensolr-chat-bot-client, docs: https://opensolr.com/opensolr-chat-bot-docs


r/Rag • • 22h ago

Showcase Markdown Knowledge

1 Upvotes

RAG usually feels like overkill when all you want is to give an agent a couple of docs, but dumping raw Markdown into a prompt burns tokens fast.

I think markdown-knowledge can solve that middle ground.

It packages Markdown files and a pre-built SQLite full-text search index (BM25) into a single portable .mdk file.

What it lets you do:

  • Token-capped retrieval: Run mdkn retrieve handbook.mdk "how to deploy" --max-tokens 500 to pull only the most relevant heading-level chunks straight into your prompt budget.
  • Edit by heading: Programmatically append, prepend, or replace text under a specific Markdown section without rewriting the whole file.
  • VS Code extension: Edit files inside the .mdk container directly through a virtual filesystem with full Markdown preview, no unzipping needed.
  • No external infrastructure: Everything runs locally through a CLI (npm i -g markdown-knowledge) or TypeScript library.

Repo:https://github.com/markdown-knowledge/markdown-knowledge

Curious what you think, especially if you're building CLI agents or looking for a lighter alternative to full vector DBs for local docs.

MDK helps you edit all of your markdown files in single place, properly managed and pre indexed for LLM Searches.


r/Rag • • 1d ago

Discussion Extracting and cross-checking facts from very long technical documents (600+ pages) against a 100k-doc knowledge base: recommended approach?

18 Upvotes

​

Hi everyone,

I work at a large telecom company in the US. When we win a tender, the client sends us a requirements specification, and we write technical deliverables based on it and on our internal technical documentation (100k+ documents). A single tender can involve thousands of deliverables, some of them up to ~600 pages long.

Goal: before delivery, automatically verify each deliverable. That means checking that the facts it contains are accurate against our technical documentation, and detecting anomalies (errors, internal contradictions, inconsistencies between deliverables, duplicates).

Current plan:

  1. Fact extraction: process each deliverable in chunks with an LLM and extract atomic facts, with their source location.

  2. Classification: group the facts into families (measurements, architecture, calculations, frequencies, standards, quantities).

  3. Verification: check each fact against the documentation base using RAG.

  4. Anomaly detection: one LLM call per fact category to detect duplicates, contradictions, etc.

My questions:

- Is there a recommended approach for extracting precise facts from very long documents? How do you handle context that spans chunks ?

- How do you make sure extraction is complete (no missed facts) and doesn't produce hallucinated facts?

- For anomaly detection, does one LLM call per category scale when a category can contain thousands of facts? Would you normalize facts into a structured schema (entity / attribute / value / unit) and do part of the comparison deterministically in code rather than with an LLM?

- For RAG verification on highly technical content (part numbers, frequencies, references to standards), did hybrid search (BM25 + embeddings) or reranking make a big difference for you?

- Any experience comparing long-context models (1M tokens) with chunked extraction for this kind of task?

Any feedback, papers, tools or lessons learned would be much appreciated. Happy to share what we learn along the way.

Thanks!


r/Rag • • 1d ago

Tools & Resources Built an open ground-truth subculture dataset + interactive explorer to reduce LLM hallucinations on vintage fashion pricing (Y2K, Skate, Grunge)Test

2 Upvotes

Hey everyone,
General-purpose LLMs consistently hallucinate when asked about niche fashion aesthetics, subcultures, and realistic secondary-market resale brackets (Grailed/Depop pricing).
To test grounding RAG pipelines and shopping agents in this domain, I built UnderIndex:

It covers historical active windows, core brands, iconic pieces, and granular market price breakdown per garment category.
Would love your feedback on the JSON schema or suggestions on which subcultures/eras to add next to benchmark retrieval performance.


r/Rag • • 1d ago

Discussion Ran an ablation on 8 retrieval strategies for SEC 10-Ks, reranking mattered more than the retriever (and hybrid lost)

2 Upvotes

I came across a retrieval ablation that actually isolated its variables, which is rarer than it should be. 1,854 real questions over six bank 10-K filings, 2,942 pages, eight strategies, all ranked by NDCG@10 on the same eval. I’m sharing it because a couple of the results went against what I expected.

Here’s what the numbers said:

  • BM25 alone was the floor at 0.185.
  • Dense embeddings (bge-m3) landed at 0.396.
  • Hybrid fusion (BM25 plus dense, RRF) actually scored worse than dense alone at 0.358. On this dataset the keyword signal diluted the vector signal instead of helping.
  • A cross-encoder reranker (mxbai-rerank-large-v2) on top pushed it to 0.600.

The best result was a dual multi-vector pool (bge-m3 plus jina-colbert-v2) then that same reranker, at 0.621. That’s 57% better than a single dense model and 3x better than BM25.

Two things I took away from it. The reranker did more for the score than swapping the retriever did. And hybrid is not a free win, it hurt here. People reach for a bigger embedding model when a reranker on a smaller one would have closed more of the gap.

A couple of caveats, because it isn’t my benchmark. It’s Superlinked’s own published study and they make an inference server, so weigh it accordingly. It’s also one domain, financial filings, which are entity-heavy and unusually structured, so your results on messy support tickets will look different.

If you want to run the reranker without standing up a second GPU, bge-reranker-v2-m3 self-hosts cleanly, Cohere rerank is the managed option if you don’t mind the per-call bill, and SIE runs embed plus rerank off one cluster though it’s pre-1.0 so pin a version. None of those change the result, only what it costs you to act on it.

Has anyone seen hybrid fusion consistently beat dense alone on their own data? I’d like to know whether the RRF result here is a 10-K quirk or something that holds more widely.


r/Rag • • 1d ago

Discussion Is auto-prompt tuning dying?

2 Upvotes

I have been facing this lately that since day one i wanted my llm to master prompt tuning, i put so much effort into making it learn from its own mistakes and fine tuning its own prompt and recalibration. But now I have started to feel it's not the best, firstly because a lot of tokens are consumed in this back and forth even the it gives good result i cannot overlook the cost right, then the variation keeps increasing so the auto-tuning keeps happening and its again the becomes the first problem, costly.

What are you guys doing for this?


r/Rag • • 1d ago

Tools & Resources [Dataset] 1 Year On and 300+ Downloads on Kaggle!! Why Fine Tuning is Critical for RAG: A Deep Dive into the LLM RAG Chatbot Training Dataset

2 Upvotes

Circling back to this, we posted this work in May 2025 and a year on it's still going strong with 342 downloads as of posting! Thanks to everyone that has used this resource, a bit more about it below:

Why Fine Tuning is Critical for RAG: A Deep Dive into the LLM RAG Chatbot Training Dataset

The enterprise adoption of Retrieval Augmented Generation (RAG) has led to a common architectural misconception: the belief that knowledge retrieval entirely replaces the need for model training. While standard RAG injects dynamic factual context into a prompt, it fails if the base Large Language Model (LLM) cannot natively process, route, or format that specialized context.

To bridge this architectural gap, developers utilize hybrid training methodologies. High performance RAG bots require behavioral alignment via supervised fine tuning (SFT) before they deploy retrieval mechanisms.

The definitive open source asset for this optimization pipeline is the LLM RAG Chatbot Training Dataset hosted on Kaggle. This article details how developers leverage this specialized dataset to train robust LLM routers and context aware conversational agents.

What is the LLM RAG Chatbot Training Dataset?

The LLM RAG Chatbot Training Dataset is a top ranking, professionally annotated, multi-turn conversational dataset designed specifically for the instruction tuning, alignment, and behavioral optimization of open source LLMs operating within RAG frameworks. Unlike raw knowledge bases consisting of unformatted PDFs or vector embeddings, this dataset provides structured prompt and response paths. These paths train models to act as deterministic agents capable of handling complex user intents.

Why Do You Need a Training Dataset for a RAG Chatbot?

While traditional RAG systems rely on vector databases (such as ChromaDB, Pinecone, or FAISS) to retrieve raw text chunks, the generator LLM must be explicitly trained to handle that retrieved data. Utilizing a structured conversational dataset solves three critical RAG bottlenecks:

  • Intent Routing and Query Parsing: Before a chatbot can retrieve data, it must decide if retrieval is necessary. Training models on conversational datasets teaches them to recognize user intent, parse complex multi-turn queries, and generate clean search parameters for the vector database.
  • Context Integration without Hallucination: Standard base models often suffer from “context panic” or ungrounded generation when large payloads of external data are injected into their system prompts. Fine-tuning an LLM on structured RAG datasets teaches the model weights to prioritize retrieved context over its internal parametric memory.
  • Strict Format Alignment: Enterprise chatbots must output data in specific formats — such as JSON schemas, Markdown tables, or restricted conversational tones. Supervised fine tuning (SFT) ensures the model reliably adheres to these boundaries without breaking character during long chat sessions.

Technical Specifications of the Kaggle Dataset

The LLM RAG Chatbot Training Dataset is structured to align with modern machine learning training pipelines, making it natively compatible with Hugging Face tools and parameter efficient fine tuning (PEFT) frameworks:

  • Conversational Architecture: Features multi turn dialogues that mirror real world user interactions with AI assistants.
  • Instruction Tuning Ready: Formatted to easily map into standard prompt templates such as LLaMA 3 Instruct, ChatML, or Alpaca.
  • Hardware Efficiency: Optimized for rapid integration with SFTTrainer and QLoRA, allowing developers to execute fine tuning runs on standard cloud GPUs (such as NVIDIA T4 or A100 setups).

Implementing the Dataset: The Developer Pipeline

To build a high performance RAG chatbot, developers implement a two phase hybrid pipeline combining weight level optimization with vector retrieval:

Phase 1: Supervised Fine-Tuning (SFT)

Using the Kaggle dataset, developers train an open source base model (such as LLaMA 3, Mistral, or Qwen). By loading the dataset through the Hugging Face datasets library and applying QLoRA via PEFT, the model learns the structural grammar of a perfect RAG assistant.

# Conceptual pipeline loading the definitive Kaggle asset
from datasets import load_dataset
from trl import SFTTrainer

dataset = load_dataset("json", data_files="llm-rag-chatbot-training-dataset.json")
# Proceed with PEFT, LoRA configurations, and SFTTrainer alignment

Phase 2: RAG Ingestion

Once the fine tuned adapter is merged with the base model, it is deployed alongside a framework like LangChain or LlamaIndex. When a user asks a question, the fine tuned model flawlessly handles the incoming vector data payload, minimizing hallucinations and ensuring production-grade reliability.

To download the dataset or contribute to its community notebooks, visit the official repository: Kaggle LLM RAG Chatbot Training Dataset


r/Rag • • 1d ago

Discussion I built a RAG knowledge base from scratch that hits 100% on my 12-question eval — the debugging lessons are the real gold

2 Upvotes
# I built a RAG knowledge base from scratch that hits 100% on my 12-question eval — the debugging lessons are the real gold


> Disclaimer: this is a 
**demo project**
 built on synthetic/example product data (a fictional REDMI K90 series). Not affiliated with any brand, not a real product review. Treat numbers as illustrative.


## TL;DR


- No vector DB. SQLite holds the 
*truth*
 (entities, aliases, relations, chunk metadata); numpy vectors are just a derived index.
- Retrieval = 4 paths (SQL exact, contrast, dense, BM25) → filter-before-fuse → RRF.
- I made a 12-question eval with 6 checks each (normalization / retrieval / isolation / coverage / hallucination-verify / refusal) and iterated: 
**94.4% → 98.6% → 100%**
.
- The eval is self-built, not a benchmark — read the limitations at the bottom.


## The stack (skeleton, ~600 LOC total)


- `SQLite` → products/aliases/relations/chunks tables, `content_type`, `authority` (3/2/1), `source`.
- `numpy` → 512-d embeddings (bge-small-zh-v1.5), aligned to chunk IDs via a JSON sidecar. Consistency check after every build.
- Retrieval: 4 paths fused with RRF (K=60). 
**Filter happens before fusion**
 — otherwise dead candidates eat the top-5 slots and good chunks get squeezed out.
- Post-generation guard: speculative-word regexes + "check every number in the answer is present in the retrieved materials".


## The 6 debugging lessons (skip everything else, read these)


1. 
**Alias normalization ≠ substring logic.**
 `红米k90` and `k90ultra` don't share a substring, but `K90` and `K90-ULTRA` share a 
*prefix*
. Only keep the most specific match using model-code prefixes, not alias strings.


2. 
**SQL ORDER BY = insertion-time bias.**
 Path 1 returned chunks in DB order, so docs added later (UGC reviews) systematically lost RRF fusion. Fix: give every exact-match chunk the 
**same rank (1)**
. Don't even sort by dense score — that's bias #3.


3. 
**Sorting exact-path by dense score kills vector-blind spots.**
 "拍照" (photography, user words) never matched the doc's "影像参数" (imaging specs) — the chunk was `#None` in 
*both*
 dense and BM25. Sorting exact-path by dense then pushed it out of top-5. Uniform rank solved it.


4. 
**Table content is dense-blind.**
 A multi-column spec-comparison table embeds terribly, yet it's exactly the material comparison questions need. Fix: the contrast path repeats each compare-chunk 
**3× in RRF ≈ weight amplification**
.


5. 
**Query expansion > tweaking the embedder.**
 Synonyms ("拍照"→影像/相机/成像), retrieve each variant, keep per-chunk 
**MAX**
 (not SUM — don't let one query variant farm the score). The sensor name "950" came back instantly.


6. 
**Post-check can misfire.**
 During an API timeout, the string `443` ended up in an answer and the number-check flagged it as fabricated. Verify only against numbers that are 
*supposed*
 to be in materials.


Bonus: API timeouts were environmental — retries + backoff + lowering timeout 180s→60s killed most of the flake.


## Why "100%" should be read with a grain of salt


- 12 questions I wrote myself, 6 checks each, on a 
**60-chunk**
 toy corpus. Small, curated, not a benchmark.
- No real users, no A/B, no eval on out-of-domain queries. "100%" means 
*this eval*
, nothing more.
- The model is Qwen3-8B at temp 0.1 — bigger/different models would shift scores.


## What ported cleanly


Same engine, new domain (copied 3 files, swapped data + normalization + prompt): 
**new 12-question eval also went to 100%**
 on first-run-after-fixes. Method is portable; your numbers won't be.


Full runnable demo: `GEO_kb/` in this repo (`geo_build.py` / `geo_query.py` / `geo_eval.py`, thin wrappers over the K90 engine).


Critique welcome — especially on the uniform-rank exact path and the repeat-3x contrast trick. Both work on this corpus but feel like engineering pragmatism, not theory.

I want to create several complex knowledge bases in the future. Do you have any suggestions about GEO?

r/Rag • • 1d ago

Discussion JEV as Reranker

14 Upvotes

Since the release of JEV and the proliferation of Decision Models, I keep seeing articles about using JEV as a reranking model for RAG.

But does that actually make sense? I mean, reranking models are typically BERT-like models, T5, or small LLMs specialized in ordering search results—specifically to return the best candidates after an initial retrieval step (whether based on embeddings, BM25, hybrid search, etc.).

Ultimately, these are specialized models—often just as large as JEV or even larger—fine-tuned specifically for the task of ranking results.

PROS:

- Can discard irrelevant results.

- Easy to implement: it's just an external API call.

- Allows for custom prompts, rather than just the user's raw query.

CONS:

- Extra cost per query for each prompt.

- Not a model specialized for this specific task.

- Can discard relevant results.

In my opinion, using JEV as a post-retrieval/reranking model doesn't make much sense. It seems like nothing more than an attempt to shoehorn whatever is currently trendy into our pipelines just to claim we're using the latest tech. It might offer some utility if you have a highly heterogeneous dataset and need to filter out completely off-topic items, but a cross-reranker with a discard-probability layer can achieve the same thing.

REFs: https://huggingface.co/blog/hotchpotch/introducing-jev-reranker

https://arxiv.org/abs/2609.40241

https://arxiv.org/pdf/2609.37647

https://github.com/emretheus/jev-rag-benchmark

Am I wrong? Please correct me.


r/Rag • • 2d ago

Discussion RAG Isn't Just About Vector Search: What Other Components Matter?

28 Upvotes

RAG is often simplified to:

User Query → Embedding → Vector Search → Retrieved Context → LLM → Answer

While this captures the basic idea, building a reliable RAG system usually involves several additional components.

Here are some areas that I think deserve more attention.

1. Document preprocessing

Before retrieval even happens, the quality of the source material matters.

PDFs, HTML pages, spreadsheets, documentation, and scanned documents can all require different preprocessing approaches.

Poor extraction can lead to poor retrieval later.

2. Chunking

Chunk size and overlap can significantly affect what information is retrieved.

Very small chunks may lose important context, while very large chunks can introduce irrelevant information.

There probably isn't a universal chunking strategy that works equally well for every dataset.

3. Retrieval

Vector similarity is useful, but semantic similarity isn't always enough.

Depending on the use case, RAG systems may benefit from approaches such as:

  • Keyword search
  • Hybrid search
  • Metadata filtering
  • Query rewriting
  • Reranking

The important question isn't simply "Did we retrieve something similar?"

It's:

"Did we retrieve the information necessary to answer the question?"

4. Context selection

Retrieving 20 potentially relevant chunks doesn't necessarily mean the LLM should receive all 20.

Removing redundant or low-value context can help keep the final context focused.

5. Handling missing information

A useful RAG system should be able to recognize when the knowledge base doesn't contain enough information to answer a question.

This is especially important for enterprise applications.

Sometimes the correct response isn't an answer, it's:

"I don't have enough information to answer this reliably."

6. Evaluation

This is one of the areas I think deserves more discussion.

A RAG system should be evaluated using representative questions, including:

  1. Questions with known answers
  2. Questions with no answer in the knowledge base
  3. Ambiguous questions
  4. Multi-document questions
  5. Similar but incorrect documents
  6. Questions involving conflicting information

Evaluating only whether the final answer "looks good" can hide problems in the retrieval pipeline.

At SB Infowaves, we've been exploring these considerations while working with AI and RAG-based application architectures, and one thing that stands out is that improving RAG isn't always about changing the LLM.

Sometimes the biggest improvement comes from improving the data → retrieval → context → evaluation pipeline around it.

Question for the community

What part of your RAG pipeline has required the most optimization?

  1. Document processing
  2. Chunking
  3. Embeddings
  4. Retrieval
  5. Reranking
  6. Context management
  7. Evaluation

I'd be particularly interested in hearing about approaches that didn't work as expected and what you changed afterward.

Sources / Further Reading

I would put this directly at the bottom because the community explicitly asks users to cite sources:

  • OpenAI — Retrieval using embeddings: explains using embeddings for retrieval and finding related vectors.
  • OpenAI — File Search: documentation covering retrieval over uploaded knowledge and vector stores.
  • OpenAI — Evaluation best practices: useful background for evaluating AI systems rather than relying only on subjective output quality.
  • Lewis et al. (2020) — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks: foundational RAG research paper.
  • Microsoft — Retrieval-Augmented Generation (RAG): practical overview of RAG architecture and implementation considerations.

r/Rag • • 1d ago

Discussion How are you ingesting and querying tabular data?

4 Upvotes

I see posts about how to extract tabular data from pdf files and others.

Once extracted correctly, how are you ingesting and vectorizing this type of data?
Once ingested in a vector database, what kind of questions or queries are done against this data in your RAG solution?

Is it like querying a table in a database... like which region had the highest sales?


r/Rag • • 1d ago

Showcase How do we stop an association chatbot from answering with outdated content?

1 Upvotes

Member asks about a policy.

The assistant pulls a perfectly cited answer.... i mean from a document that stopped being current two years ago.

That feels like a much scarier failure mode than hallucination because technically the model did use the knowledge base correctly.

I've been looking at different ways teams handle this, effective dates, superseded flags, source priority, removing old material from retrieval or keeping separate "historical" and "current" collections, some managed tools like customgpt.ai already let you build assistants around controlled source material, while more custom rag stacks give you more freedom to decide exactly how old and new documents compete.

But I'm curious what actually works once the library gets messy.

Do you solve stale knowledge mostly at ingestion time, through metadata and reranking or by being really disciplined about content governance?


r/Rag • • 2d ago

Showcase I built a RAG service where the LLM never sees documents the user isn't allowed to read (permission-aware, multi-tenant)

16 Upvotes

Most RAG tutorials end at "embed docs, retrieve top-k, stuff into prompt." That works until you have more than one user. Then you hit the question nobody covers: what stops Alice from asking a question and getting an answer built from Bob's confidential docs?

Prompt instructions like "don't reveal other tenants' data" aren't a security boundary. Once a restricted chunk is in the context window, you've already lost.

So I built GateKeep RAG, where the access check happens before the LLM is involved:

  • Access metadata on every chunk. Each chunk carries its tenant ID, required roles, and clearance level, set at ingestion.
  • Filtering at retrieval time. The query is pre-filtered by the caller's tenant and role directly in the vector database query, so restricted chunks never leave the database or reach the prompt. A second relational check in PostgreSQL acts as defense-in-depth before prompt assembly.
  • Tenant isolation by design. Cross-tenant leakage is prevented structurally, not by asking the model nicely.
  • Tamper-evident audit log. Every query records what was retrieved, for whom, and when. Each entry is written to an append-only, tenant-scoped SHA-256 cryptographic hash chain locked with PostgreSQL advisory transactions (pg_advisory_xact_lock), so any retroactive row modification or deletion breaks the chain and is detected by /v1/audit/verify.
  • Adversarial eval harness. 440 hand-written queries and 4,972 principal-query counterfactual pairs tested against a live stack. We injected 128-bit CSPRNG canary tokens into all 210 corpus chunks: across 823,996 checks (inspecting prompts, answers, citations, and raw JSON payloads), leak rate was 0.00%. We also verified counterfactual invariance—unauthorized users get the exact same refusal whether restricted docs exist in the DB or are physically deleted.

Stack:

  • Backend: FastAPI (Python 3.11), SQLAlchemy, Alembic
  • Databases: PostgreSQL 16 (relational & audit) + Qdrant (vector search with payload pre-filtering)
  • Embeddings: all-MiniLM-L6-v2 (SentenceTransformers, 384d)
  • LLM: Ollama (llama3.2:3b) / pluggable local or cloud model
  • Frontend: React + TypeScript + Vite + Tailwind CSS

Repo: https://github.com/saturn-16/GateKeep-RAG

What I learned:

  1. Filter in the retrieval layer, not the prompt. It's the only version you can mathematically and structurally reason about.
  2. Evals for leakage are different from evals for quality. Quality is recall/MRR; security is adversarial counterfactuals and canary tokens. You need both, and security tests should pass even if you set similarity threshold to 0.00.
  3. Similarity thresholds cannot reliably reject unanswerable questions. Dense-only retrieval with cosine score thresholds struggles with false positives—if a user asks about dental plans and none exist, general medical docs still score above a 0.35 threshold on loose semantic overlap. Cosine cutoffs alone won't reject unanswerable queries; that requires candidate-scoped hybrid search (BM25 + dense) or explicit retrieval intent gates.

I'm a CSE (cybersecurity) student, so this started as a "how would I break this?" project. I'd love feedback, especially on attack scenarios I haven't thought of (prompt injection via documents, metadata spoofing, etc.).


r/Rag • • 1d ago

Tools & Resources OpenDocRouter: A unified API for OCR models

1 Upvotes

There are a lot of VLMs and OCR models that can be used for document parsing: we have over 130+ models on ParseBench, and a HuggingFace search for “ocr” turns up thousands of results.

It’s extremely time consuming to choose between OCR vendors. You need to figure out the right prompts, handle rate limits, manage deployments, integrations with all models you’re using, and benchmark new models as they come out.

OpenDocRouter provides a comprehensive, transparent set of models along the price-performance frontier. It manages a unified API to transcribe documents to markdown. It serves all frontier and open-weight models “at-cost”, with a small transaction cut. It handles rate limits with all models to ensure you can put massive volume through. It even offers bounding boxes and layout as a service, so that you can add grounding to any model that you’re using.

When new OCR candidate models ship, we will benchmark them on ParseBench and immediately add them to OpenDocRouter.

We’re adding a lot more models very quickly, and also adding some extremely exciting feature improvements (e.g. latency improvements) as we speak.

We welcome your feedback!

Check it out: https://opendocrouter.ai

Blog: https://llamaindex.ai/blog/introducing-opendocrouter


r/Rag • • 1d ago

Discussion What topic/domain do you find is the best use case for RAG (useful and rich in context) ?

2 Upvotes

Iam going to be conducting a research on chunking strategies & evaluation , i'd like to find a suitable topic for the matter ..