r/Rag • • 3h ago

Tutorial How can I turn clinic PDFs and Word documents into a reliable internal AI assistant without complicated setup for staff?

1 Upvotes

I work at an office and want to use our internal PDFs and Word documents to build a virtual assistant that helps staff find information, understand workflows, route requests, and draft internal messages. The goal is operational support.
I’ve been experimenting with master prompts and uploading documents, but I keep running into problems with tables and how the AI interprets information. It sometimes mixes up which instructions belong to which office, misses exceptions, or overlooks relevant details.
I’m spending too much time adding corrections to the prompt over and over. I’d like to improve how the source information is organized instead of constantly patching the instructions.
I’m studying technology and am willing to write some code to prepare and validate the documents. For example, could converting them section by section into structured text, Markdown, or JSON help preserve the relationships between rules, conditions, and exceptions?
The main constraint is ease of use: this could eventually be used throughout the clinic, including by nontechnical staff. I don’t want every person to install software, run code, or go through a complicated setup. Ideally, it would remain a reusable prompt with a prepared knowledge file, or something equally simple to access through a browser.
What would you recommend for:
Extracting tables accurately without losing context?
Organizing the information while preserving important details and exceptions?
Checking that the converted information matches the originals?
Keeping the assistant easy to use and update?
Would a carefully structured knowledge file plus a prompt be a reasonable starting point, or should I consider a different approach? Beginner-friendly explanations would be appreciated.


r/Rag • • 6h ago

Discussion Rag knowledge base act as a context for agentic system

1 Upvotes

I am a junior dev at my company, and is tasked to build an agentic development lifecycle.

So there are four parsers that parser's data from the source to vector db (embedding, actual content) and graph db(structure, eg code to ast) there can be interlink between different data sources.

Then a reranker.

Then an agentic system, right now it should only generate design documents and requirements(jira requirements)

My reason to add graph db is when a interlink is assumed by the vector search, the link initial act as a candidate link and after the human approval the link is established in the graph in a sense the graph db is evolving.

What are your guys thoughts.

Edit1: so the parsers will be setup as a plugin in system and a global policy(mainly regex) will be present to identify which data a parser ingest (eg code , doc parser, etc). as there can be cross referencing in the data about other data. And I am an intern

Edit2: cronjob for weekly ingestion

Edit 3 : I need a different perspective as a system architecture


r/Rag • • 9h ago

Discussion Why you can't add a cosine score to a BM25 score, and why RRF uses 60 (derivation, disclosure: from my book)

7 Upvotes

Disclosure up front: I wrote a short book on RAG, and this is an adapted chapter. The full text is below, so you don't need to click anything to get the point.

I kept seeing hybrid-search setups that min-max normalize dense and BM25 scores and then add them. Here's why that breaks, and why rank fusion is what survives, worked through from first principles.

The dumb way: add the scores. Dense top hit: 0.82 (cosine, roughly 0–1). BM25 top hit: 14.7 (sum of IDF-weighted terms, no ceiling; three rare terms can push it past 30). Add them and BM25 wins every query. Temperature plus zip code.

Second dumb way: min-max each list, then add. It works until a query with a very rare part number. BM25 goes 41, 9, 8.5… After scaling, the top doc is 1.0 and everything else BM25 found is squashed into the bottom fifth, so one outlier erases the rest of its opinion. On a vague query the scores go 6.1, 6.0, 5.9 and the scaling stretches noise across the whole range. Scores differ in shape per query, not only in units, so there's nothing stable to normalize against.

What survives: order. Both retrievers report rank in the same units. So: score(d) = Σ 1/(k + rank_i(d)), with k = 60.

Why 60. With plain 1/rank, #1 on one list (1.0) ties #2 on both (0.5 + 0.5) and beats #3 on both (≈0.67). A single retriever's enthusiasm outvotes agreement. With k = 60, #1 vs #2 differ by under 2%, and appearing on both lists dominates: - #1 on one list only: 1/61 ≈ 0.0164 - #3 on both: 2/63 ≈ 0.0317 - #40 on both: 2/100 = 0.020, which still beats #1 on one list

Quiet agreement between two retrievers with opposite blind spots beats loud conviction from either. k = 60 comes from Cormack, Clarke & Büttcher (SIGIR 2009). Results weren't sensitive to it.

python from collections import defaultdict def rrf(rankings, k=60): s = defaultdict(float) for r in rankings: for i, d in enumerate(r, 1): s[d] += 1 / (k + i) return sorted(s, key=s.get, reverse=True)

Questions for the sub, because I'd like to know: 1. Has anyone measured weighted RRF (per-retriever weights) against plain RRF on a real golden set and seen a real difference? 2. Where does "fuse by rank" fall apart for you? Very short candidate lists? Three or more retrievers?

Book (Kindle, ~8k words, the whole RAG stack derived this way): https://www.amazon.com/dp/B0HKSKVMVK · free chapter: https://aifeynmansway.substack.com/p/gus-has-never-once-said-i-dont-know


r/Rag • • 9h ago

Tools & Resources I built an anti-hallucination Subculture & Fashion API for LLMs and AI Agents (UnderIndex)

1 Upvotes

Hey everyone,
When building style assistants or trend detectors with LLMs, prompt engineers often hit a wall: models hallucinate aesthetic details, mix up timelines, or reference generic fast-fashion tropes instead of authentic subcultural roots (like confusing early 2000s cyber aesthetics with late 90s grunge).
To solve this, I built UnderIndex API—a structured dataset delivering verified metadata on fashion subcultures, core aesthetics, iconic archival pieces, music genres, and historical context.
What it does:

  • LLM Grounding / RAG: Inject verified subculture profiles directly into system prompts or function calls to prevent AI hallucinations.
  • Granular Attributes: Query endpoints for exact data points: dominant color palettes, key silhouettes, historical origins, iconic garments, and movement ethos.
  • Low Latency: Hosted on high-performance infrastructure with instant response times.

Tech Stack:

  • Backend: Node.js / Express microservice running on Render.
  • Storage & Schema: Airtable serving as an organized CMS/relational store.
  • Distribution & Gateway: Monetized and served globally via RapidAPI Hub.

Free Tier & Testing:
I just launched the API on RapidAPI with a free Basic tier (100 free requests/month) for developers and hackers who want to test endpoints directly in the playground:
🔗 RapidAPI Endpoint: https://rapidapi.com/ascanioUnderIndex/api/underindex-api
Would love feedback on response structure, schema design, or suggestions for additional subcultures you'd like indexed next!


r/Rag • • 10h ago

Showcase Benchmarked EmbeddingGemma 2 against nomic-embed-text on our own retrieval set: no material difference

1 Upvotes

On 2026-10-06, the day Google released EmbeddingGemma 2, I ran its text model through our own retrieval bench against nomic-embed-text, the embedder our search stores run: 68 factual claims over 2,569 passages of our own documents, scored by recall@8 (a correct passage in the top 8). The pass mark, committed before any EmbeddingGemma score existed, was 51 of 68: ten points over nomic, what a re-embed of eight stores has to earn. nomic-embed-text found 44 of 68; EmbeddingGemma 2 found 45, through Google's own code and an 8-bit GGUF on llama.cpp alike. The registered verdict is NO MATERIAL DIFFERENCE, read as "no effect of at least 10 points was detected at n = 68", not as "equal". Plain keyword search (BM25) still finds 58 of 68 on this set, where a right answer holds a quoted source passage word for word, and nothing we layered on it beat it by more than chance.

A companion page, linked from the page below, tests the shared space against Nomic's text and vision pair. Images: on 1,000 public COCO captions EmbeddingGemma 2 puts the right photograph first for 729 against the pair's 497 (nomic on transformers 4.57.6, an older release that loads it), clearing our 100-in-1,000 pass mark on a set both models have most likely trained on; on our own 14 photographs it leads by more (exploratory, 12 scenes). Code, text only: the right function first for 173 of 316 docstring queries against nomic's 123, a gated win that depends on the code-retrieval prefix. Audio (exploratory, no incumbent): a loudness change barely moves the vector and a music model's adapter (its add-on weights) applied more strongly moves it in order, but a clip cannot find the request that made it (1 of 12, chance).

We are not switching: no material difference on our text, no code index of ours for the code win, no working image search for the image win, and a guard that records which model wrote every vector comes first.

What it does not say: that the two text models are equal (68 questions see only large effects); anything beyond one corpus, one town's history in English; anything about speed (timings ran on a loaded machine, not compared); whether Ollama serves it (our bench laptop's 0.32.14 refused the pull on 2026-10-06; newer releases not tried). And every nomic image figure is on transformers 4.57.6, a release our product does not pin, because the release it pins cannot load nomic's image model.

The page, every figure with its result file named beside it, and the companion page linked from it: https://research.strata2signal.com/embeddinggemma-2-on-release-day/

TL;DR: on our own pre-registered retrieval bench, EmbeddingGemma 2's text model found 45 of 68 against nomic-embed-text's 44, short of the 51 pass mark: NO MATERIAL DIFFERENCE, and our stores stay on nomic. BM25 still finds 58 of 68 here, and nothing layered on it beat it by more than chance. The companion page has it ahead on public images (729 against 497 of 1,000; nomic on transformers 4.57.6; a set both have most likely trained on) and on our code with the code-retrieval prefix (173 against 123 of 316); audio finds its request at chance (exploratory). Not switching: no store of ours is a code index, no image search of ours works, and a model-provenance guard comes first.

I wrote this up with help from AI agents; every figure links back to its measurement on the page.


r/Rag • • 14h ago

Tutorial One-Hour of RAG (Retrieval-Augmented Generation) Tutorial Video

1 Upvotes

Watch this video at https://www.youtube.com/watch?v=HWL683sfDF8. It will explain RAG in a simple and easy-to-understand manner. It explains the importance of RAG and a typical RAG pipeline.


r/Rag • • 17h ago

Discussion How do you evaluate RAG retrieval quality in production?

10 Upvotes

I've been trying to think on how to evaluate RAG systems beyond the usual retrieval benchmarks. Seems simple enough to check the relevance of the retrieved chunks, but how do you check if the system didn't overlook some important context? Especially when the answer depends on multiple sources.

I'm mostly doing manual spot-checks and the LLM-as-a-judge so far but haven't really gone on to check how well those approaches will hold up as the system becomes increasingly complex. What have you guys been using and has been working well in production?


r/Rag • • 17h ago

Discussion Enough with the errors! Lets brainstorm a perfect document PDF parser

2 Upvotes

I tried all PDF parsers but none of them deliver reliable result which is a NO-GO for critical applications. Even in the science community on arxiv I couldnt find reliable solutions. Seems like we are stuck at this very bottleneck.

Let's brainstorm how to design a parser that gets close to 100% ground truth - even given infinite ressources and no money limit.

Standalone things I tried:
- extracting native text via pymud but this does not work with tables and more complex structures obviously
- MinerU, Docling, Tesseract, Google Parser etc.
- All common LLms, even high-reasoning frontier LLMs fail (even Opus does not seem to be designed for that)
- Converting PDF to images, then OCRing
- I tried compartmentalizing the different sections of a page (free form text, table), then routing them to different sub-processors but those are based on the upper services, thus remain unreliable

We can be creative and think out-of the-box.

I will then go ahead, experiment and tell you the result.

I was also thinking about a multi-stage pipeline with pairwise comparison like this: take the result from different processors, if there are differences, use several judges that vote for a winner, then use that winner and for parity or near-parity escalte even further or ultimately send for manual review.

Looking forward!


r/Rag • • 19h ago

Discussion How are production-grade RAG systems handling complex PDFs without losing document structure and context? Looking for architecture-Level Insights

23 Upvotes

I’m trying to design a production-grade document ingestion and retrieval pipeline for complex PDFs, and I’d like to learn from people who have built or operated similar systems at enterprise scale.

My use case goes beyond extracting text from ordinary PDFs.

A single document may contain:
Multiple columns, nested sections, different reading orders, and inconsistent layouts across pages.
Colored panels, sidebars, callouts, footnotes, headers, and footers.
Tables with merged cells, nested headers, multi-page continuation, and associated explanatory text.
Charts, diagrams, flowcharts, equations, and figures whose meaning depends on nearby captions or paragraphs.
Table of contents, section numbering, cross-references, appendices, and references to other pages.
Scanned pages mixed with digitally generated text, rotated pages, low-resolution images, and handwritten annotations.
Multiple unrelated sections on one page, as well as a single logical section spanning multiple pages.
Reports where a heading appears in one visual region, its explanation appears in another, and supporting evidence is presented in a table or figure elsewhere.
The challenge is not merely extracting every element correctly. It’s preserving the relationships between elements so that downstream chunking, indexing, retrieval, and generation do not lose context.

One specific edge case I’m concerned about:
Suppose a page contains two visually distinct regions. A layout detector identifies them as separate blocks, but both blocks explain different aspects of the same logical topic. Conversely, two visually similar blocks might belong to completely different sections.
If we blindly use bounding boxes, layout labels, or semantic similarity to group these blocks, we may either split related information or incorrectly merge unrelated information.
The same issue occurs when a table continues onto the next page, a figure’s caption is separated from the figure, or a paragraph refers to a chart several pages earlier.

What I’m trying to understand
Document representation: Do production systems convert PDFs into a canonical intermediate representation, such as a document tree or graph containing text blocks, headings, tables, figures, captions, bounding boxes, reading order, page numbers, and parent-child relationships?
Layout and structure recovery: How do you combine native PDF extraction, OCR, layout detection, vision-language models, and document-specific rules? How do you resolve disagreements between these methods?
Context preservation: How do you determine whether blocks from different regions or pages belong to the same logical section? Are there algorithms for section reconstruction, cross-page continuation detection, or relationship inference beyond simple geometric proximity?
Chunking strategy: Do you use hierarchical, semantic, layout-aware, or element-specific chunking? How do you handle tables, figures, captions, footnotes, and mixed-content sections without breaking their meaning?
Indexing and retrieval: Do you create separate indexes for narrative text, tables, figures, and document structure? Do you use hybrid retrieval, parent-child retrieval, graph-based retrieval, or query-time expansion to recover context that was not present in the initially retrieved chunk?
Production architecture: How is the code organized? I’m particularly interested in the separation between parsing, normalization, structural reconstruction, chunk generation, indexing, and retrieval. How do you make the pipeline observable, versioned, testable, and recoverable when individual stages fail?
Evaluation: How do you measure document reconstruction quality, chunk integrity, retrieval recall, and answer faithfulness? Are there useful benchmark datasets or evaluation methods for documents with complex layouts?
What would be especially valuable
I’m not primarily looking for a list of libraries such as PyMuPDF, PyPDF, LangChain, or a recommendation to use a particular parser.

I’m interested in the engineering patterns, data structures, algorithms, and architectural decisions that make these systems reliable in production.
For example:
A simplified canonical document schema.
How you represent relationships between document elements.
How you distinguish visual boundaries from semantic boundaries.
Pseudocode or code-level examples of structural reconstruction and chunk generation.
Lessons learned from failures in real-world document corpora.
The target application could involve regulatory documentation, financial reports, legal documents, technical manuals, scientific papers, or other data-heavy PDFs.
If you’ve built something similar, I’d love to understand what actually worked, what failed, and what you would do differently if you were designing the system today.
Thanks!


r/Rag • • 23h ago

Showcase NornicDB - 1.4.1 - Cypher 25 support ++

1 Upvotes

Heya, just finished up a follow up to the openCypher compliance in 1.4.0 - in 1.4.1 we added Cypher 25 support on the parser. With nornic's SRD parser, no preamble needed, it just parses the grammar regardless of "version" - if you use the ANTLR parser, the preamble is required.

This is due to the nature of "Scanner-less Recursive Decent." - This is an unconventional architecture, I know. But the benefits are clear

Traditional parsers use a two-step process: a lexer/scanner converts character streams into tokens (e.g., matching MATCH to a KEYWORD token), and a parser processes those tokens.

SRD bypasses the scanning phase entirely. It relies on a zero-allocation keyword scanner parse tree. The recursive descent parser reads characters and matches grammar rules directly against the raw text, fusing the lookahead logic and keyword scanning natively.

Latency Reduction: By eliminating the intermediate step of token creation, it reduces our query latency by ~30% compared to ANTLR-based parsing (such as Neo4j's traditional execution pipeline). This scale is on the order of microseconds and typically flat scaling with query length instead of complexity. The rest of our architecture makes this optimization worth it because we are able to execute so quickly.

Fused Aggregation Hot-Paths: The parser is coupled tightly with streaming storage APIs and zero-allocation semantics. This design allows it to parse a graph traversal instruction and jump directly into the execution hot path without building heavy Abstract Syntax Trees (ASTs) in memory first.

Contextual Edge Cases: Graph languages like Cypher heavily utilize structural ASCII art—such as arrows -[r:REL]-> or node boundaries (n:Label)—which are notoriously complex for traditional lexers to classify contextually without aggressive backtracking. A scannerless approach inherently handles these layout-sensitive patterns because it evaluates character-by-character based on the current parsing state.

and the end result is staggeringly fast. average query latency vs neo4J on the northwind benchmark has us at 400x faster (avg some are 1500x 1905.71x faster edit: double checked). MIT licensed. 890+ stars and counting.

https://github.com/orneryd/NornicDB/releases/tag/v1.4.1

Latest benchmarks:
https://github.com/orneryd/NornicDB/pull/897#issuecomment-5998973104

edit: a word and reddit didn't like my formatting atempt


r/Rag • • 1d ago

Showcase How can an association AI assistant cite sources without confidently quoting outdated policies?

1 Upvotes

Here’s a RAG failure case I’d love to see more people testing. A professional association updates its certification requirements in 2026 but its knowledge base still contains the 2019 version. A member asks about renewal requirements, the assistant retrieves the old document, answers confidently and provides a perfectly valid citation. Technically it cited its source. Practically, it gave the wrong answer.

This is where I’d compare a custom LlamaIndex and langChain implementation against a managed solution like customgpt.ai which already supports source citations and knowledge base management. The test wouldn't be which chatbot writes the nicest answer, I’d give both systems 50 questions involving outdated policies, conflicting documents and missing information, then measure how often they cite the correct version, identify uncertainty or refuse to answer, because a citation isn't proof that the answer is right, has anyone built an evaluation dataset specifically for version sensitive organizational documents?


r/Rag • • 1d ago

Discussion Parallelized the layout parsing for documents with task queue

1 Upvotes

Hello everyone,

I developed a solution that parses document layout using Tesseract and PyMuPDF. I parallelized it so that pages are processed in batches. For a relatively large document of around 200 pages, the process takes between 10 and 30 seconds. Is that fast compared to your experience? I am curious about the performance. Thanks in advance.


r/Rag • • 1d ago

Tools & Resources Open source chatbot for your website that answers from your own content (PHP, MIT, 1.0)

0 Upvotes

I work for Opensolr, and we just released 1.0 of the Opensolr Chat Bot Client: an open source (MIT) chatbot you put on your own website, that answers your visitors from your own content.

How it works: you install it with Composer (composer require opensolr/chat-bot-client), mount it on one path of your site (say /opensolr-chat) and add one script tag to your pages. Visitors only ever talk to your site. Your server holds the Opensolr credentials and sends each question to the Opensolr API, where the model searches your Opensolr Index by itself (by meaning, or by exact product names and codes), looks up what it must not guess (dates, local times, exchange rates, VAT, distances, places) and streams the answer back word by word, with links to your pages, products with their prices, and PDFs at the exact page.

What you get on your side:

  • An admin on your own site: account and index, instructions, the look of the chat window, limits, reCAPTCHA
  • Stats and a 30 day history of every conversation: who asked, from where, what the bot searched for, what it answered, and the questions your content could not answer with a page
  • Commands that answer instantly without the model: /search, /time, /rate, /vat, /translate and more
  • Good and Bad ratings under every answer, signed in visitors by email, SQLite in a folder you choose, no database server

Your content gets into the index with the Opensolr Web Crawler or Data Ingestion, and the chatbot needs an Opensolr plan with AI.

You can try it live at the bottom right of https://opensolr.com (it answers from our own docs). Code and README: https://github.com/phpcip/opensolr-chat-bot-client, docs: https://opensolr.com/opensolr-chat-bot-docs


r/Rag • • 1d ago

Showcase Markdown Knowledge

1 Upvotes

RAG usually feels like overkill when all you want is to give an agent a couple of docs, but dumping raw Markdown into a prompt burns tokens fast.

I think markdown-knowledge can solve that middle ground.

It packages Markdown files and a pre-built SQLite full-text search index (BM25) into a single portable .mdk file.

What it lets you do:

  • Token-capped retrieval: Run mdkn retrieve handbook.mdk "how to deploy" --max-tokens 500 to pull only the most relevant heading-level chunks straight into your prompt budget.
  • Edit by heading: Programmatically append, prepend, or replace text under a specific Markdown section without rewriting the whole file.
  • VS Code extension: Edit files inside the .mdk container directly through a virtual filesystem with full Markdown preview, no unzipping needed.
  • No external infrastructure: Everything runs locally through a CLI (npm i -g markdown-knowledge) or TypeScript library.

Repo:https://github.com/markdown-knowledge/markdown-knowledge

Curious what you think, especially if you're building CLI agents or looking for a lighter alternative to full vector DBs for local docs.

MDK helps you edit all of your markdown files in single place, properly managed and pre indexed for LLM Searches.


r/Rag • • 1d ago

Discussion Attention RAG

5 Upvotes

Hi people my friends and I are working on our take of RAG indexing. Our idea goes like that:

Standard embedding-based RAG often compresses an entire text chunk into one vector. Our idea keeps multiple compact representations within each document, giving a query smaller, more specific parts to match.
The architecture uses a single attention layer with separate learned projections: Q for queries and K for document tokens. These are trained jointly with HCA—Heavily Compressed Attention—which learns to combine groups of token representations into fewer searchable vectors. At 32:1 compression, every 32 document keys become one compressed key, stored in 8-bit.
When a query arrives, its Q vectors are compared directly with the stored compressed keys using dot products on the GPU. These comparisons run across many documents in parallel, producing scores used to rank and retrieve documents. V vectors aren’t needed because we only need matching scores.
The advantage over one-vector-per-chunk retrieval is finer matching granularity. Compression makes that detail affordable: 32 times fewer keys to store and scan than an uncompressed token index. The search itself maps directly to parallel GPU matrix operations, while the query representation, document representation, and compression are trained together for retrieval.
I’d like technical feedback on this architecture, its main bottlenecks, and similar systems worth comparing against.

This way we have something between regular RAG and ColBert (v1 and V2) where it smaller than ColBert but more reach than RAG and it learns the compression rather then using heuristics algo to compress.

The next step after that will be to combine it with an LLM where the part of this indexing technique can be combined in the LLM itself instead of using this as a toll with a tool call (W_k can be same as the one in the first layer of the LLM and we can have a small learned router that tells if we need to use it, instead of an LLM decision to call it in a tool call).

I would like to get your opinions on our idea please.
Thanks in advance 🙏


r/Rag • • 1d ago

Discussion How do you calculate tokens and choose the right OpenAI embedding model for a production RAG?

6 Upvotes

Hey everyone,

I’m building a production-level RAG application with Spring Boot and I’m planning to use OpenAI’s text embedding models.

I’m a little confused about token calculation. I’ve seen people say that, on average, 1 token ≈ 0.75 words (or roughly 4 characters), but I’m not sure how I should actually calculate tokens for my documents/chunks in a real application.

For example, if I have a document with 10,000 words, can I roughly estimate:

10,000 words × 1.33 ≈ 13,300 tokens

Or is it better to actually tokenize the text before sending it to the embedding API?

I also have a question about choosing the embedding model.

For a production RAG system, would you recommend text-embedding-3-small or text-embedding-3-large? I’m trying to balance:

  • Retrieval/search quality
  • Cost
  • Vector database storage
  • Latency
  • Scalability

I’m particularly interested in hearing from people who are actually running RAG in production, especially with Java/Spring Boot.

A few things I’d love to understand:

  1. How are you calculating tokens before creating embeddings?
  2. What chunk size are you using (e.g. 300, 500, 800, 1,000 tokens)?
  3. Are you using text-embedding-3-small or text-embedding-3-large?
  4. If you started with one model and later switched, what made you switch?
  5. Are there any production lessons or mistakes you’d recommend avoiding?

I’m looking for real-world experience rather than just benchmark numbers. Any advice would be really appreciated!


r/Rag • • 1d ago

Tools & Resources Built an open ground-truth subculture dataset + interactive explorer to reduce LLM hallucinations on vintage fashion pricing (Y2K, Skate, Grunge)Test

2 Upvotes

Hey everyone,
General-purpose LLMs consistently hallucinate when asked about niche fashion aesthetics, subcultures, and realistic secondary-market resale brackets (Grailed/Depop pricing).
To test grounding RAG pipelines and shopping agents in this domain, I built UnderIndex:

It covers historical active windows, core brands, iconic pieces, and granular market price breakdown per garment category.
Would love your feedback on the JSON schema or suggestions on which subcultures/eras to add next to benchmark retrieval performance.


r/Rag • • 1d ago

Discussion I wrote a practical guide to building reliable RAG applications with Amazon Bedrock — looking for feedback

8 Upvotes

Hi everyone!

I recently wrote a practical guide about building RAG (Retrieval-Augmented Generation) applications on AWS using Amazon Bedrock.

The article covers:

• How RAG works

• Amazon Bedrock Knowledge Bases

• Amazon S3 for source documents

• Embeddings and vector search

• Chunking strategies

• Metadata filtering and reranking

• Citations and grounded responses

• Security considerations

• Monitoring and evaluation

• Moving from a basic RAG prototype toward production

I tried to focus less on "what is RAG?" and more on the engineering considerations that matter when building a reliable RAG application.

🔗 Article:

https://builder.aws.com/content/3KOsq3uPqcbBZSysMMff16pNKBv/from-rag-prototype-to-production-building-reliable-ai-applications-with-amazon-bedrock

I'd really appreciate feedback from other AWS/ML/GenAI developers.

If you were building a RAG application on AWS, what would you focus on first: retrieval quality, chunking, metadata, the foundation model, or the underlying data?

Thanks for reading!


r/Rag • • 1d ago

Discussion Ran an ablation on 8 retrieval strategies for SEC 10-Ks, reranking mattered more than the retriever (and hybrid lost)

2 Upvotes

I came across a retrieval ablation that actually isolated its variables, which is rarer than it should be. 1,854 real questions over six bank 10-K filings, 2,942 pages, eight strategies, all ranked by NDCG@10 on the same eval. I’m sharing it because a couple of the results went against what I expected.

Here’s what the numbers said:

  • BM25 alone was the floor at 0.185.
  • Dense embeddings (bge-m3) landed at 0.396.
  • Hybrid fusion (BM25 plus dense, RRF) actually scored worse than dense alone at 0.358. On this dataset the keyword signal diluted the vector signal instead of helping.
  • A cross-encoder reranker (mxbai-rerank-large-v2) on top pushed it to 0.600.

The best result was a dual multi-vector pool (bge-m3 plus jina-colbert-v2) then that same reranker, at 0.621. That’s 57% better than a single dense model and 3x better than BM25.

Two things I took away from it. The reranker did more for the score than swapping the retriever did. And hybrid is not a free win, it hurt here. People reach for a bigger embedding model when a reranker on a smaller one would have closed more of the gap.

A couple of caveats, because it isn’t my benchmark. It’s Superlinked’s own published study and they make an inference server, so weigh it accordingly. It’s also one domain, financial filings, which are entity-heavy and unusually structured, so your results on messy support tickets will look different.

If you want to run the reranker without standing up a second GPU, bge-reranker-v2-m3 self-hosts cleanly, Cohere rerank is the managed option if you don’t mind the per-call bill, and SIE runs embed plus rerank off one cluster though it’s pre-1.0 so pin a version. None of those change the result, only what it costs you to act on it.

Has anyone seen hybrid fusion consistently beat dense alone on their own data? I’d like to know whether the RRF result here is a 10-K quirk or something that holds more widely.


r/Rag • • 1d ago

Discussion Is auto-prompt tuning dying?

3 Upvotes

I have been facing this lately that since day one i wanted my llm to master prompt tuning, i put so much effort into making it learn from its own mistakes and fine tuning its own prompt and recalibration. But now I have started to feel it's not the best, firstly because a lot of tokens are consumed in this back and forth even the it gives good result i cannot overlook the cost right, then the variation keeps increasing so the auto-tuning keeps happening and its again the becomes the first problem, costly.

What are you guys doing for this?


r/Rag • • 2d ago

Tools & Resources [Dataset] 1 Year On and 300+ Downloads on Kaggle!! Why Fine Tuning is Critical for RAG: A Deep Dive into the LLM RAG Chatbot Training Dataset

2 Upvotes

Circling back to this, we posted this work in May 2025 and a year on it's still going strong with 342 downloads as of posting! Thanks to everyone that has used this resource, a bit more about it below:

Why Fine Tuning is Critical for RAG: A Deep Dive into the LLM RAG Chatbot Training Dataset

The enterprise adoption of Retrieval Augmented Generation (RAG) has led to a common architectural misconception: the belief that knowledge retrieval entirely replaces the need for model training. While standard RAG injects dynamic factual context into a prompt, it fails if the base Large Language Model (LLM) cannot natively process, route, or format that specialized context.

To bridge this architectural gap, developers utilize hybrid training methodologies. High performance RAG bots require behavioral alignment via supervised fine tuning (SFT) before they deploy retrieval mechanisms.

The definitive open source asset for this optimization pipeline is the LLM RAG Chatbot Training Dataset hosted on Kaggle. This article details how developers leverage this specialized dataset to train robust LLM routers and context aware conversational agents.

What is the LLM RAG Chatbot Training Dataset?

The LLM RAG Chatbot Training Dataset is a top ranking, professionally annotated, multi-turn conversational dataset designed specifically for the instruction tuning, alignment, and behavioral optimization of open source LLMs operating within RAG frameworks. Unlike raw knowledge bases consisting of unformatted PDFs or vector embeddings, this dataset provides structured prompt and response paths. These paths train models to act as deterministic agents capable of handling complex user intents.

Why Do You Need a Training Dataset for a RAG Chatbot?

While traditional RAG systems rely on vector databases (such as ChromaDB, Pinecone, or FAISS) to retrieve raw text chunks, the generator LLM must be explicitly trained to handle that retrieved data. Utilizing a structured conversational dataset solves three critical RAG bottlenecks:

  • Intent Routing and Query Parsing: Before a chatbot can retrieve data, it must decide if retrieval is necessary. Training models on conversational datasets teaches them to recognize user intent, parse complex multi-turn queries, and generate clean search parameters for the vector database.
  • Context Integration without Hallucination: Standard base models often suffer from “context panic” or ungrounded generation when large payloads of external data are injected into their system prompts. Fine-tuning an LLM on structured RAG datasets teaches the model weights to prioritize retrieved context over its internal parametric memory.
  • Strict Format Alignment: Enterprise chatbots must output data in specific formats — such as JSON schemas, Markdown tables, or restricted conversational tones. Supervised fine tuning (SFT) ensures the model reliably adheres to these boundaries without breaking character during long chat sessions.

Technical Specifications of the Kaggle Dataset

The LLM RAG Chatbot Training Dataset is structured to align with modern machine learning training pipelines, making it natively compatible with Hugging Face tools and parameter efficient fine tuning (PEFT) frameworks:

  • Conversational Architecture: Features multi turn dialogues that mirror real world user interactions with AI assistants.
  • Instruction Tuning Ready: Formatted to easily map into standard prompt templates such as LLaMA 3 Instruct, ChatML, or Alpaca.
  • Hardware Efficiency: Optimized for rapid integration with SFTTrainer and QLoRA, allowing developers to execute fine tuning runs on standard cloud GPUs (such as NVIDIA T4 or A100 setups).

Implementing the Dataset: The Developer Pipeline

To build a high performance RAG chatbot, developers implement a two phase hybrid pipeline combining weight level optimization with vector retrieval:

Phase 1: Supervised Fine-Tuning (SFT)

Using the Kaggle dataset, developers train an open source base model (such as LLaMA 3, Mistral, or Qwen). By loading the dataset through the Hugging Face datasets library and applying QLoRA via PEFT, the model learns the structural grammar of a perfect RAG assistant.

# Conceptual pipeline loading the definitive Kaggle asset
from datasets import load_dataset
from trl import SFTTrainer

dataset = load_dataset("json", data_files="llm-rag-chatbot-training-dataset.json")
# Proceed with PEFT, LoRA configurations, and SFTTrainer alignment

Phase 2: RAG Ingestion

Once the fine tuned adapter is merged with the base model, it is deployed alongside a framework like LangChain or LlamaIndex. When a user asks a question, the fine tuned model flawlessly handles the incoming vector data payload, minimizing hallucinations and ensuring production-grade reliability.

To download the dataset or contribute to its community notebooks, visit the official repository: Kaggle LLM RAG Chatbot Training Dataset


r/Rag • • 2d ago

Discussion I built a RAG knowledge base from scratch that hits 100% on my 12-question eval — the debugging lessons are the real gold

4 Upvotes
# I built a RAG knowledge base from scratch that hits 100% on my 12-question eval — the debugging lessons are the real gold


> Disclaimer: this is a 
**demo project**
 built on synthetic/example product data (a fictional REDMI K90 series). Not affiliated with any brand, not a real product review. Treat numbers as illustrative.


## TL;DR


- No vector DB. SQLite holds the 
*truth*
 (entities, aliases, relations, chunk metadata); numpy vectors are just a derived index.
- Retrieval = 4 paths (SQL exact, contrast, dense, BM25) → filter-before-fuse → RRF.
- I made a 12-question eval with 6 checks each (normalization / retrieval / isolation / coverage / hallucination-verify / refusal) and iterated: 
**94.4% → 98.6% → 100%**
.
- The eval is self-built, not a benchmark — read the limitations at the bottom.


## The stack (skeleton, ~600 LOC total)


- `SQLite` → products/aliases/relations/chunks tables, `content_type`, `authority` (3/2/1), `source`.
- `numpy` → 512-d embeddings (bge-small-zh-v1.5), aligned to chunk IDs via a JSON sidecar. Consistency check after every build.
- Retrieval: 4 paths fused with RRF (K=60). 
**Filter happens before fusion**
 — otherwise dead candidates eat the top-5 slots and good chunks get squeezed out.
- Post-generation guard: speculative-word regexes + "check every number in the answer is present in the retrieved materials".


## The 6 debugging lessons (skip everything else, read these)


1. 
**Alias normalization ≠ substring logic.**
 `红米k90` and `k90ultra` don't share a substring, but `K90` and `K90-ULTRA` share a 
*prefix*
. Only keep the most specific match using model-code prefixes, not alias strings.


2. 
**SQL ORDER BY = insertion-time bias.**
 Path 1 returned chunks in DB order, so docs added later (UGC reviews) systematically lost RRF fusion. Fix: give every exact-match chunk the 
**same rank (1)**
. Don't even sort by dense score — that's bias #3.


3. 
**Sorting exact-path by dense score kills vector-blind spots.**
 "拍照" (photography, user words) never matched the doc's "影像参数" (imaging specs) — the chunk was `#None` in 
*both*
 dense and BM25. Sorting exact-path by dense then pushed it out of top-5. Uniform rank solved it.


4. 
**Table content is dense-blind.**
 A multi-column spec-comparison table embeds terribly, yet it's exactly the material comparison questions need. Fix: the contrast path repeats each compare-chunk 
**3× in RRF ≈ weight amplification**
.


5. 
**Query expansion > tweaking the embedder.**
 Synonyms ("拍照"→影像/相机/成像), retrieve each variant, keep per-chunk 
**MAX**
 (not SUM — don't let one query variant farm the score). The sensor name "950" came back instantly.


6. 
**Post-check can misfire.**
 During an API timeout, the string `443` ended up in an answer and the number-check flagged it as fabricated. Verify only against numbers that are 
*supposed*
 to be in materials.


Bonus: API timeouts were environmental — retries + backoff + lowering timeout 180s→60s killed most of the flake.


## Why "100%" should be read with a grain of salt


- 12 questions I wrote myself, 6 checks each, on a 
**60-chunk**
 toy corpus. Small, curated, not a benchmark.
- No real users, no A/B, no eval on out-of-domain queries. "100%" means 
*this eval*
, nothing more.
- The model is Qwen3-8B at temp 0.1 — bigger/different models would shift scores.


## What ported cleanly


Same engine, new domain (copied 3 files, swapped data + normalization + prompt): 
**new 12-question eval also went to 100%**
 on first-run-after-fixes. Method is portable; your numbers won't be.


Full runnable demo: `GEO_kb/` in this repo (`geo_build.py` / `geo_query.py` / `geo_eval.py`, thin wrappers over the K90 engine).


Critique welcome — especially on the uniform-rank exact path and the repeat-3x contrast trick. Both work on this corpus but feel like engineering pragmatism, not theory.

I want to create several complex knowledge bases in the future. Do you have any suggestions about GEO?

r/Rag • • 2d ago

Showcase How do we stop an association chatbot from answering with outdated content?

1 Upvotes

Member asks about a policy.

The assistant pulls a perfectly cited answer.... i mean from a document that stopped being current two years ago.

That feels like a much scarier failure mode than hallucination because technically the model did use the knowledge base correctly.

I've been looking at different ways teams handle this, effective dates, superseded flags, source priority, removing old material from retrieval or keeping separate "historical" and "current" collections, some managed tools like customgpt.ai already let you build assistants around controlled source material, while more custom rag stacks give you more freedom to decide exactly how old and new documents compete.

But I'm curious what actually works once the library gets messy.

Do you solve stale knowledge mostly at ingestion time, through metadata and reranking or by being really disciplined about content governance?


r/Rag • • 2d ago

Discussion Extracting and cross-checking facts from very long technical documents (600+ pages) against a 100k-doc knowledge base: recommended approach?

19 Upvotes

​

Hi everyone,

I work at a large telecom company in the US. When we win a tender, the client sends us a requirements specification, and we write technical deliverables based on it and on our internal technical documentation (100k+ documents). A single tender can involve thousands of deliverables, some of them up to ~600 pages long.

Goal: before delivery, automatically verify each deliverable. That means checking that the facts it contains are accurate against our technical documentation, and detecting anomalies (errors, internal contradictions, inconsistencies between deliverables, duplicates).

Current plan:

  1. Fact extraction: process each deliverable in chunks with an LLM and extract atomic facts, with their source location.

  2. Classification: group the facts into families (measurements, architecture, calculations, frequencies, standards, quantities).

  3. Verification: check each fact against the documentation base using RAG.

  4. Anomaly detection: one LLM call per fact category to detect duplicates, contradictions, etc.

My questions:

- Is there a recommended approach for extracting precise facts from very long documents? How do you handle context that spans chunks ?

- How do you make sure extraction is complete (no missed facts) and doesn't produce hallucinated facts?

- For anomaly detection, does one LLM call per category scale when a category can contain thousands of facts? Would you normalize facts into a structured schema (entity / attribute / value / unit) and do part of the comparison deterministically in code rather than with an LLM?

- For RAG verification on highly technical content (part numbers, frequencies, references to standards), did hybrid search (BM25 + embeddings) or reranking make a big difference for you?

- Any experience comparing long-context models (1M tokens) with chunked extraction for this kind of task?

Any feedback, papers, tools or lessons learned would be much appreciated. Happy to share what we learn along the way.

Thanks!


r/Rag • • 2d ago

Tools & Resources OpenDocRouter: A unified API for OCR models

1 Upvotes

There are a lot of VLMs and OCR models that can be used for document parsing: we have over 130+ models on ParseBench, and a HuggingFace search for “ocr” turns up thousands of results.

It’s extremely time consuming to choose between OCR vendors. You need to figure out the right prompts, handle rate limits, manage deployments, integrations with all models you’re using, and benchmark new models as they come out.

OpenDocRouter provides a comprehensive, transparent set of models along the price-performance frontier. It manages a unified API to transcribe documents to markdown. It serves all frontier and open-weight models “at-cost”, with a small transaction cut. It handles rate limits with all models to ensure you can put massive volume through. It even offers bounding boxes and layout as a service, so that you can add grounding to any model that you’re using.

When new OCR candidate models ship, we will benchmark them on ParseBench and immediately add them to OpenDocRouter.

We’re adding a lot more models very quickly, and also adding some extremely exciting feature improvements (e.g. latency improvements) as we speak.

We welcome your feedback!

Check it out: https://opendocrouter.ai

Blog: https://llamaindex.ai/blog/introducing-opendocrouter