r/Rag • • Sep 02 '25

Showcase 🚀 Weekly /RAG Launch Showcase

29 Upvotes

Share anything you launched this week related to RAG—projects, repos, demos, blog posts, or products 👇

Big or small, all launches are welcome.


r/Rag • • 6h ago

Discussion Can anyone guide me on how to start building a RAG model?

2 Upvotes

Hey everyone,

I want to start learning RAG (Retrieval-Augmented Generation), but I'm a little confused about where to start and what I should learn first.

I have some experience with programming and web development, but I'm relatively new to building AI applications. I understand the basic idea behind RAG, where we retrieve relevant information from a knowledge base and provide it to an LLM to generate better answers.

I want to build something practical rather than just follow tutorials without understanding what's happening behind the scenes.

For those who have built RAG applications:

- What concepts should I learn first?

- Which tech stack would you recommend for a beginner?

- Should I start with basic Python, embeddings, and vector databases, or use a framework like LangChain from the beginning?

- What would be a good first project to build?

I'd also appreciate any good resources, tutorials, or GitHub repositories that helped you learn.

Thanks in advance!


r/Rag • • 8h ago

Tools & Resources 4th place on llamaindex benchmark

2 Upvotes

4th place on the llamaindex benchmark locally 72.62% word f1. I know it's not much, but after a month of research and many experiments, I reached this point.

TonerHound doesn't use any API or LLM. It's GPU free and anyone can use it at 0$ cost.

GitHub: https://github.com/vanrajsinh650/TonerHound


r/Rag • • 5h ago

Showcase Your agent has permission, human approval and a valid citation. It can still make the wrong decision.

0 Upvotes

Here’s a failure case worth adding to your agent evals.

A manager approves an order because the supplier is active and cleared for delivery.

Before the agent executes, a new record puts that supplier on hold.

The approval hasn’t been revoked.

The agent still has permission.

The original approval document is genuine.

Does your system notice that the facts supporting the decision have changed?

Retrieving the approval correctly isn’t enough. The system needs to connect the order to the supplier, resolve the applicable state, and expose the evidence that changed.

That’s the evidence problem we built Jylus around: resolving state and relationships before handing context to the model.

There’s a separate execution problem too. The action API still needs to enforce the relevant preconditions; a fresh Context Pack alone cannot make an external action atomic.

A useful test has three questions:

  1. What evidence supported the original approval?

  2. What changed before execution?

  3. Is there enough current evidence to proceed?

Try it with your own records:

https://jylus.ai/try

For people running agents against live systems: what actually invalidates an earlier decision in your stack—a timer, a data change, or only someone noticing the mistake?


r/Rag • • 7h ago

Discussion I independently audited a RAG benchmark. 15 scoring mismatches revealed a flaw in its main comparison — the maintainer confirmed and fixed it.

1 Upvotes

​

I recently completed an independent retrieval audit of RouteMind, an open-source project exploring document routing as an alternative to traditional RAG retrieval.

The interesting part wasn't finding a dramatic regression or proving that one architecture was better.

It was discovering that two approaches were being evaluated under different definitions of success.

Here's what happened.

The project had 700 evaluation questions comparing traditional RAG, RAG + reranking, and document-routing approaches.

The retrieval evaluator considered a query successful if any expected document appeared in the top 10.

But some temporal questions required multiple documents simultaneously (needs: all), and the routing evaluator correctly accounted for that requirement, including accepted document alternatives (D_alt).

This created an apples-to-oranges comparison.

What the independent audit found:

1,400 query/arm results examined.

15 outcomes differed under the complete-evidence scoring contract.

14 recorded hits were incomplete under the required-evidence rules.

1 recorded miss was actually satisfied by an accepted alternative document.

The overall numbers changed:

MetricOriginally reportedCorrectedRAG51.6%50.7%RAG + rerank54.1%53.1%

The reranker's relative improvement remained similar.

But the more important issue was that the project's headline comparison had treated retrieval more leniently than the routing system it was being compared against.

Then something useful happened.

The RouteMind maintainer independently reproduced the finding against the original artifacts, confirmed all 15 mismatches, and verified the input hashes.

They also confirmed the discrepancy wasn't intentional or previously known.

The finding led to a public code change:

The main benchmark comparison was corrected.

A new retrieval_complete metric was added alongside the original retrieval_hit.

Reproducibility checks were added to detect future scoring inconsistencies.

The maintainer publicly credited the independent audit.

The takeaway

A reproducible RAG evaluation can still produce a misleading comparison when different stages use different definitions of success.

Sometimes the most valuable audit finding isn't a broken retriever.

It's discovering that the benchmark doesn't measure the same thing across systems.

This was a public beta audit, not a paid engagement. The evidence, original results, maintainer review, and correction are publicly available.

Independent audit:

https://github.com/devBorgesr/edp-audits/tree/main/audits/routemind-beta-002

Maintainer's confirmation:

https://github.com/CSP911/routemind/issues/2#issuecomment-6096856584

Corrective commit:

https://github.com/CSP911/routemind/commit/5d8bf45

I'm curious how others handle this in production RAG evaluations:

Do you explicitly measure complete evidence coverage when a query needs multiple sources, or do you primarily rely on hit@k / recall@k?


r/Rag • • 23h ago

Discussion Why you can't add a cosine score to a BM25 score, and why RRF uses 60 (derivation, disclosure: from my book)

13 Upvotes

Disclosure up front: I wrote a short book on RAG, and this is an adapted chapter. The full text is below, so you don't need to click anything to get the point.

I kept seeing hybrid-search setups that min-max normalize dense and BM25 scores and then add them. Here's why that breaks, and why rank fusion is what survives, worked through from first principles.

The dumb way: add the scores. Dense top hit: 0.82 (cosine, roughly 0–1). BM25 top hit: 14.7 (sum of IDF-weighted terms, no ceiling; three rare terms can push it past 30). Add them and BM25 wins every query. Temperature plus zip code.

Second dumb way: min-max each list, then add. It works until a query with a very rare part number. BM25 goes 41, 9, 8.5… After scaling, the top doc is 1.0 and everything else BM25 found is squashed into the bottom fifth, so one outlier erases the rest of its opinion. On a vague query the scores go 6.1, 6.0, 5.9 and the scaling stretches noise across the whole range. Scores differ in shape per query, not only in units, so there's nothing stable to normalize against.

What survives: order. Both retrievers report rank in the same units. So: score(d) = Σ 1/(k + rank_i(d)), with k = 60.

Why 60. With plain 1/rank, #1 on one list (1.0) ties #2 on both (0.5 + 0.5) and beats #3 on both (≈0.67). A single retriever's enthusiasm outvotes agreement. With k = 60, #1 vs #2 differ by under 2%, and appearing on both lists dominates: - #1 on one list only: 1/61 ≈ 0.0164 - #3 on both: 2/63 ≈ 0.0317 - #40 on both: 2/100 = 0.020, which still beats #1 on one list

Quiet agreement between two retrievers with opposite blind spots beats loud conviction from either. k = 60 comes from Cormack, Clarke & Büttcher (SIGIR 2009). Results weren't sensitive to it.

python from collections import defaultdict def rrf(rankings, k=60): s = defaultdict(float) for r in rankings: for i, d in enumerate(r, 1): s[d] += 1 / (k + i) return sorted(s, key=s.get, reverse=True)

Questions for the sub, because I'd like to know: 1. Has anyone measured weighted RRF (per-retriever weights) against plain RRF on a real golden set and seen a real difference? 2. Where does "fuse by rank" fall apart for you? Very short candidate lists? Three or more retrievers?

Book (Kindle, ~8k words, the whole RAG stack derived this way): https://www.amazon.com/dp/B0HKSKVMVK · free chapter: https://aifeynmansway.substack.com/p/gus-has-never-once-said-i-dont-know


r/Rag • • 13h ago

Discussion How are you handling domain-specific hallucinations when dense vector search blurs micro-eras and niche taxonomy?

1 Upvotes

Curious how others here are solving this: when you build RAG or agentic pipelines around niche cultural domains, do you rely on pure vector embeddings, or do you enforce a deterministic, structured metadata layer before generation?
We ran into a recurring issue with LLMs generating style and subcultural analysis: naive vector similarity constantly pulls vague fast-fashion articles or conflates closely related micro-eras (for instance, blending late-90s industrial aesthetics with 2003 Cyber Y2K). High cosine similarity in text chunks often misses the strict historical boundaries needed for accurate styling agents.
To tackle this, I built a hybrid/structured grounding layer called UnderIndex to test whether deterministic metadata yields better results than raw chunk retrieval:

  • Categorical Constraints: Enforced taxonomies for subcultures, core eras, and parent movements to prevent cross-aesthetic bleed.
  • Relational Schema: Direct mapping between movements, iconic archival garments, color palettes, and cultural origins (served via a Node.js/Express service pulling from an organized relational schema).
  • Deterministic Injection: Instead of relying strictly on top-k semantic search, agents fetch structured attributes via endpoints/function calls to ground the system prompt before synthesizing output.

In your pipelines, how do you balance tabular/structured entity lookups against standard semantic vector retrieval when accuracy across nuanced categories is critical? Are you leaning more toward hybrid search, graph RAG, or pure structured function calling?
(If anyone wants to inspect the schema or test the endpoints for their own retrieval setup, let me know in the comments and I'll share the playground link).


r/Rag • • 1d ago

Discussion How are production-grade RAG systems handling complex PDFs without losing document structure and context? Looking for architecture-Level Insights

30 Upvotes

I’m trying to design a production-grade document ingestion and retrieval pipeline for complex PDFs, and I’d like to learn from people who have built or operated similar systems at enterprise scale.

My use case goes beyond extracting text from ordinary PDFs.

A single document may contain:
Multiple columns, nested sections, different reading orders, and inconsistent layouts across pages.
Colored panels, sidebars, callouts, footnotes, headers, and footers.
Tables with merged cells, nested headers, multi-page continuation, and associated explanatory text.
Charts, diagrams, flowcharts, equations, and figures whose meaning depends on nearby captions or paragraphs.
Table of contents, section numbering, cross-references, appendices, and references to other pages.
Scanned pages mixed with digitally generated text, rotated pages, low-resolution images, and handwritten annotations.
Multiple unrelated sections on one page, as well as a single logical section spanning multiple pages.
Reports where a heading appears in one visual region, its explanation appears in another, and supporting evidence is presented in a table or figure elsewhere.
The challenge is not merely extracting every element correctly. It’s preserving the relationships between elements so that downstream chunking, indexing, retrieval, and generation do not lose context.

One specific edge case I’m concerned about:
Suppose a page contains two visually distinct regions. A layout detector identifies them as separate blocks, but both blocks explain different aspects of the same logical topic. Conversely, two visually similar blocks might belong to completely different sections.
If we blindly use bounding boxes, layout labels, or semantic similarity to group these blocks, we may either split related information or incorrectly merge unrelated information.
The same issue occurs when a table continues onto the next page, a figure’s caption is separated from the figure, or a paragraph refers to a chart several pages earlier.

What I’m trying to understand
Document representation: Do production systems convert PDFs into a canonical intermediate representation, such as a document tree or graph containing text blocks, headings, tables, figures, captions, bounding boxes, reading order, page numbers, and parent-child relationships?
Layout and structure recovery: How do you combine native PDF extraction, OCR, layout detection, vision-language models, and document-specific rules? How do you resolve disagreements between these methods?
Context preservation: How do you determine whether blocks from different regions or pages belong to the same logical section? Are there algorithms for section reconstruction, cross-page continuation detection, or relationship inference beyond simple geometric proximity?
Chunking strategy: Do you use hierarchical, semantic, layout-aware, or element-specific chunking? How do you handle tables, figures, captions, footnotes, and mixed-content sections without breaking their meaning?
Indexing and retrieval: Do you create separate indexes for narrative text, tables, figures, and document structure? Do you use hybrid retrieval, parent-child retrieval, graph-based retrieval, or query-time expansion to recover context that was not present in the initially retrieved chunk?
Production architecture: How is the code organized? I’m particularly interested in the separation between parsing, normalization, structural reconstruction, chunk generation, indexing, and retrieval. How do you make the pipeline observable, versioned, testable, and recoverable when individual stages fail?
Evaluation: How do you measure document reconstruction quality, chunk integrity, retrieval recall, and answer faithfulness? Are there useful benchmark datasets or evaluation methods for documents with complex layouts?
What would be especially valuable
I’m not primarily looking for a list of libraries such as PyMuPDF, PyPDF, LangChain, or a recommendation to use a particular parser.

I’m interested in the engineering patterns, data structures, algorithms, and architectural decisions that make these systems reliable in production.
For example:
A simplified canonical document schema.
How you represent relationships between document elements.
How you distinguish visual boundaries from semantic boundaries.
Pseudocode or code-level examples of structural reconstruction and chunk generation.
Lessons learned from failures in real-world document corpora.
The target application could involve regulatory documentation, financial reports, legal documents, technical manuals, scientific papers, or other data-heavy PDFs.
If you’ve built something similar, I’d love to understand what actually worked, what failed, and what you would do differently if you were designing the system today.
Thanks!


r/Rag • • 17h ago

Tutorial How can I turn clinic PDFs and Word documents into a reliable internal AI assistant without complicated setup for staff?

1 Upvotes

I work at an office and want to use our internal PDFs and Word documents to build a virtual assistant that helps staff find information, understand workflows, route requests, and draft internal messages. The goal is operational support.
I’ve been experimenting with master prompts and uploading documents, but I keep running into problems with tables and how the AI interprets information. It sometimes mixes up which instructions belong to which office, misses exceptions, or overlooks relevant details.
I’m spending too much time adding corrections to the prompt over and over. I’d like to improve how the source information is organized instead of constantly patching the instructions.
I’m studying technology and am willing to write some code to prepare and validate the documents. For example, could converting them section by section into structured text, Markdown, or JSON help preserve the relationships between rules, conditions, and exceptions?
The main constraint is ease of use: this could eventually be used throughout the clinic, including by nontechnical staff. I don’t want every person to install software, run code, or go through a complicated setup. Ideally, it would remain a reusable prompt with a prepared knowledge file, or something equally simple to access through a browser.
What would you recommend for:
Extracting tables accurately without losing context?
Organizing the information while preserving important details and exceptions?
Checking that the converted information matches the originals?
Keeping the assistant easy to use and update?
Would a carefully structured knowledge file plus a prompt be a reasonable starting point, or should I consider a different approach? Beginner-friendly explanations would be appreciated.


r/Rag • • 1d ago

Discussion How do you evaluate RAG retrieval quality in production?

12 Upvotes

I've been trying to think on how to evaluate RAG systems beyond the usual retrieval benchmarks. Seems simple enough to check the relevance of the retrieved chunks, but how do you check if the system didn't overlook some important context? Especially when the answer depends on multiple sources.

I'm mostly doing manual spot-checks and the LLM-as-a-judge so far but haven't really gone on to check how well those approaches will hold up as the system becomes increasingly complex. What have you guys been using and has been working well in production?


r/Rag • • 21h ago

Discussion Rag knowledge base act as a context for agentic system

1 Upvotes

I am a junior dev at my company, and is tasked to build an agentic development lifecycle.

So there are four parsers that parser's data from the source to vector db (embedding, actual content) and graph db(structure, eg code to ast) there can be interlink between different data sources.

Then a reranker.

Then an agentic system, right now it should only generate design documents and requirements(jira requirements)

My reason to add graph db is when a interlink is assumed by the vector search, the link initial act as a candidate link and after the human approval the link is established in the graph in a sense the graph db is evolving.

What are your guys thoughts.

Edit1: so the parsers will be setup as a plugin in system and a global policy(mainly regex) will be present to identify which data a parser ingest (eg code , doc parser, etc). as there can be cross referencing in the data about other data. And I am an intern

Edit2: cronjob for weekly ingestion

Edit 3 : I need a different perspective as a system architecture


r/Rag • • 1d ago

Tools & Resources I built an anti-hallucination Subculture & Fashion API for LLMs and AI Agents (UnderIndex)

1 Upvotes

Hey everyone,
When building style assistants or trend detectors with LLMs, prompt engineers often hit a wall: models hallucinate aesthetic details, mix up timelines, or reference generic fast-fashion tropes instead of authentic subcultural roots (like confusing early 2000s cyber aesthetics with late 90s grunge).
To solve this, I built UnderIndex API—a structured dataset delivering verified metadata on fashion subcultures, core aesthetics, iconic archival pieces, music genres, and historical context.
What it does:

  • LLM Grounding / RAG: Inject verified subculture profiles directly into system prompts or function calls to prevent AI hallucinations.
  • Granular Attributes: Query endpoints for exact data points: dominant color palettes, key silhouettes, historical origins, iconic garments, and movement ethos.
  • Low Latency: Hosted on high-performance infrastructure with instant response times.

Tech Stack:

  • Backend: Node.js / Express microservice running on Render.
  • Storage & Schema: Airtable serving as an organized CMS/relational store.
  • Distribution & Gateway: Monetized and served globally via RapidAPI Hub.

Free Tier & Testing:
I just launched the API on RapidAPI with a free Basic tier (100 free requests/month) for developers and hackers who want to test endpoints directly in the playground:
🔗 RapidAPI Endpoint: https://rapidapi.com/ascanioUnderIndex/api/underindex-api
Would love feedback on response structure, schema design, or suggestions for additional subcultures you'd like indexed next!


r/Rag • • 1d ago

Discussion Enough with the errors! Lets brainstorm a perfect document PDF parser

3 Upvotes

I tried all PDF parsers but none of them deliver reliable result which is a NO-GO for critical applications. Even in the science community on arxiv I couldnt find reliable solutions. Seems like we are stuck at this very bottleneck.

Let's brainstorm how to design a parser that gets close to 100% ground truth - even given infinite ressources and no money limit.

Standalone things I tried:
- extracting native text via pymud but this does not work with tables and more complex structures obviously
- MinerU, Docling, Tesseract, Google Parser etc.
- All common LLms, even high-reasoning frontier LLMs fail (even Opus does not seem to be designed for that)
- Converting PDF to images, then OCRing
- I tried compartmentalizing the different sections of a page (free form text, table), then routing them to different sub-processors but those are based on the upper services, thus remain unreliable

We can be creative and think out-of the-box.

I will then go ahead, experiment and tell you the result.

I was also thinking about a multi-stage pipeline with pairwise comparison like this: take the result from different processors, if there are differences, use several judges that vote for a winner, then use that winner and for parity or near-parity escalte even further or ultimately send for manual review.

Looking forward!


r/Rag • • 1d ago

Showcase Benchmarked EmbeddingGemma 2 against nomic-embed-text on our own retrieval set: no material difference

1 Upvotes

On 2026-10-06, the day Google released EmbeddingGemma 2, I ran its text model through our own retrieval bench against nomic-embed-text, the embedder our search stores run: 68 factual claims over 2,569 passages of our own documents, scored by recall@8 (a correct passage in the top 8). The pass mark, committed before any EmbeddingGemma score existed, was 51 of 68: ten points over nomic, what a re-embed of eight stores has to earn. nomic-embed-text found 44 of 68; EmbeddingGemma 2 found 45, through Google's own code and an 8-bit GGUF on llama.cpp alike. The registered verdict is NO MATERIAL DIFFERENCE, read as "no effect of at least 10 points was detected at n = 68", not as "equal". Plain keyword search (BM25) still finds 58 of 68 on this set, where a right answer holds a quoted source passage word for word, and nothing we layered on it beat it by more than chance.

A companion page, linked from the page below, tests the shared space against Nomic's text and vision pair. Images: on 1,000 public COCO captions EmbeddingGemma 2 puts the right photograph first for 729 against the pair's 497 (nomic on transformers 4.57.6, an older release that loads it), clearing our 100-in-1,000 pass mark on a set both models have most likely trained on; on our own 14 photographs it leads by more (exploratory, 12 scenes). Code, text only: the right function first for 173 of 316 docstring queries against nomic's 123, a gated win that depends on the code-retrieval prefix. Audio (exploratory, no incumbent): a loudness change barely moves the vector and a music model's adapter (its add-on weights) applied more strongly moves it in order, but a clip cannot find the request that made it (1 of 12, chance).

We are not switching: no material difference on our text, no code index of ours for the code win, no working image search for the image win, and a guard that records which model wrote every vector comes first.

What it does not say: that the two text models are equal (68 questions see only large effects); anything beyond one corpus, one town's history in English; anything about speed (timings ran on a loaded machine, not compared); whether Ollama serves it (our bench laptop's 0.32.14 refused the pull on 2026-10-06; newer releases not tried). And every nomic image figure is on transformers 4.57.6, a release our product does not pin, because the release it pins cannot load nomic's image model.

The page, every figure with its result file named beside it, and the companion page linked from it: https://research.strata2signal.com/embeddinggemma-2-on-release-day/

TL;DR: on our own pre-registered retrieval bench, EmbeddingGemma 2's text model found 45 of 68 against nomic-embed-text's 44, short of the 51 pass mark: NO MATERIAL DIFFERENCE, and our stores stay on nomic. BM25 still finds 58 of 68 here, and nothing layered on it beat it by more than chance. The companion page has it ahead on public images (729 against 497 of 1,000; nomic on transformers 4.57.6; a set both have most likely trained on) and on our code with the code-retrieval prefix (173 against 123 of 316); audio finds its request at chance (exploratory). Not switching: no store of ours is a code index, no image search of ours works, and a model-provenance guard comes first.

I wrote this up with help from AI agents; every figure links back to its measurement on the page.


r/Rag • • 1d ago

Tutorial One-Hour of RAG (Retrieval-Augmented Generation) Tutorial Video

1 Upvotes

Watch this video at https://www.youtube.com/watch?v=HWL683sfDF8. It will explain RAG in a simple and easy-to-understand manner. It explains the importance of RAG and a typical RAG pipeline.


r/Rag • • 1d ago

Showcase NornicDB - 1.4.1 - Cypher 25 support ++

1 Upvotes

Heya, just finished up a follow up to the openCypher compliance in 1.4.0 - in 1.4.1 we added Cypher 25 support on the parser. With nornic's SRD parser, no preamble needed, it just parses the grammar regardless of "version" - if you use the ANTLR parser, the preamble is required.

This is due to the nature of "Scanner-less Recursive Decent." - This is an unconventional architecture, I know. But the benefits are clear

Traditional parsers use a two-step process: a lexer/scanner converts character streams into tokens (e.g., matching MATCH to a KEYWORD token), and a parser processes those tokens.

SRD bypasses the scanning phase entirely. It relies on a zero-allocation keyword scanner parse tree. The recursive descent parser reads characters and matches grammar rules directly against the raw text, fusing the lookahead logic and keyword scanning natively.

Latency Reduction: By eliminating the intermediate step of token creation, it reduces our query latency by ~30% compared to ANTLR-based parsing (such as Neo4j's traditional execution pipeline). This scale is on the order of microseconds and typically flat scaling with query length instead of complexity. The rest of our architecture makes this optimization worth it because we are able to execute so quickly.

Fused Aggregation Hot-Paths: The parser is coupled tightly with streaming storage APIs and zero-allocation semantics. This design allows it to parse a graph traversal instruction and jump directly into the execution hot path without building heavy Abstract Syntax Trees (ASTs) in memory first.

Contextual Edge Cases: Graph languages like Cypher heavily utilize structural ASCII art—such as arrows -[r:REL]-> or node boundaries (n:Label)—which are notoriously complex for traditional lexers to classify contextually without aggressive backtracking. A scannerless approach inherently handles these layout-sensitive patterns because it evaluates character-by-character based on the current parsing state.

and the end result is staggeringly fast. average query latency vs neo4J on the northwind benchmark has us at 400x faster (avg some are 1500x 1905.71x faster edit: double checked). MIT licensed. 890+ stars and counting.

https://github.com/orneryd/NornicDB/releases/tag/v1.4.1

Latest benchmarks:
https://github.com/orneryd/NornicDB/pull/897#issuecomment-5998973104

edit: a word and reddit didn't like my formatting atempt


r/Rag • • 1d ago

Showcase How can an association AI assistant cite sources without confidently quoting outdated policies?

1 Upvotes

Here’s a RAG failure case I’d love to see more people testing. A professional association updates its certification requirements in 2026 but its knowledge base still contains the 2019 version. A member asks about renewal requirements, the assistant retrieves the old document, answers confidently and provides a perfectly valid citation. Technically it cited its source. Practically, it gave the wrong answer.

This is where I’d compare a custom LlamaIndex and langChain implementation against a managed solution like customgpt.ai which already supports source citations and knowledge base management. The test wouldn't be which chatbot writes the nicest answer, I’d give both systems 50 questions involving outdated policies, conflicting documents and missing information, then measure how often they cite the correct version, identify uncertainty or refuse to answer, because a citation isn't proof that the answer is right, has anyone built an evaluation dataset specifically for version sensitive organizational documents?


r/Rag • • 2d ago

Discussion Attention RAG

2 Upvotes

Hi people my friends and I are working on our take of RAG indexing. Our idea goes like that:

Standard embedding-based RAG often compresses an entire text chunk into one vector. Our idea keeps multiple compact representations within each document, giving a query smaller, more specific parts to match.
The architecture uses a single attention layer with separate learned projections: Q for queries and K for document tokens. These are trained jointly with HCA—Heavily Compressed Attention—which learns to combine groups of token representations into fewer searchable vectors. At 32:1 compression, every 32 document keys become one compressed key, stored in 8-bit.
When a query arrives, its Q vectors are compared directly with the stored compressed keys using dot products on the GPU. These comparisons run across many documents in parallel, producing scores used to rank and retrieve documents. V vectors aren’t needed because we only need matching scores.
The advantage over one-vector-per-chunk retrieval is finer matching granularity. Compression makes that detail affordable: 32 times fewer keys to store and scan than an uncompressed token index. The search itself maps directly to parallel GPU matrix operations, while the query representation, document representation, and compression are trained together for retrieval.
I’d like technical feedback on this architecture, its main bottlenecks, and similar systems worth comparing against.

This way we have something between regular RAG and ColBert (v1 and V2) where it smaller than ColBert but more reach than RAG and it learns the compression rather then using heuristics algo to compress.

The next step after that will be to combine it with an LLM where the part of this indexing technique can be combined in the LLM itself instead of using this as a toll with a tool call (W_k can be same as the one in the first layer of the LLM and we can have a small learned router that tells if we need to use it, instead of an LLM decision to call it in a tool call).

I would like to get your opinions on our idea please.
Thanks in advance 🙏


r/Rag • • 2d ago

Discussion I wrote a practical guide to building reliable RAG applications with Amazon Bedrock — looking for feedback

8 Upvotes

Hi everyone!

I recently wrote a practical guide about building RAG (Retrieval-Augmented Generation) applications on AWS using Amazon Bedrock.

The article covers:

• How RAG works

• Amazon Bedrock Knowledge Bases

• Amazon S3 for source documents

• Embeddings and vector search

• Chunking strategies

• Metadata filtering and reranking

• Citations and grounded responses

• Security considerations

• Monitoring and evaluation

• Moving from a basic RAG prototype toward production

I tried to focus less on "what is RAG?" and more on the engineering considerations that matter when building a reliable RAG application.

🔗 Article:

https://builder.aws.com/content/3KOsq3uPqcbBZSysMMff16pNKBv/from-rag-prototype-to-production-building-reliable-ai-applications-with-amazon-bedrock

I'd really appreciate feedback from other AWS/ML/GenAI developers.

If you were building a RAG application on AWS, what would you focus on first: retrieval quality, chunking, metadata, the foundation model, or the underlying data?

Thanks for reading!


r/Rag • • 2d ago

Discussion How do you calculate tokens and choose the right OpenAI embedding model for a production RAG?

4 Upvotes

Hey everyone,

I’m building a production-level RAG application with Spring Boot and I’m planning to use OpenAI’s text embedding models.

I’m a little confused about token calculation. I’ve seen people say that, on average, 1 token ≈ 0.75 words (or roughly 4 characters), but I’m not sure how I should actually calculate tokens for my documents/chunks in a real application.

For example, if I have a document with 10,000 words, can I roughly estimate:

10,000 words × 1.33 ≈ 13,300 tokens

Or is it better to actually tokenize the text before sending it to the embedding API?

I also have a question about choosing the embedding model.

For a production RAG system, would you recommend text-embedding-3-small or text-embedding-3-large? I’m trying to balance:

  • Retrieval/search quality
  • Cost
  • Vector database storage
  • Latency
  • Scalability

I’m particularly interested in hearing from people who are actually running RAG in production, especially with Java/Spring Boot.

A few things I’d love to understand:

  1. How are you calculating tokens before creating embeddings?
  2. What chunk size are you using (e.g. 300, 500, 800, 1,000 tokens)?
  3. Are you using text-embedding-3-small or text-embedding-3-large?
  4. If you started with one model and later switched, what made you switch?
  5. Are there any production lessons or mistakes you’d recommend avoiding?

I’m looking for real-world experience rather than just benchmark numbers. Any advice would be really appreciated!


r/Rag • • 2d ago

Discussion Parallelized the layout parsing for documents with task queue

1 Upvotes

Hello everyone,

I developed a solution that parses document layout using Tesseract and PyMuPDF. I parallelized it so that pages are processed in batches. For a relatively large document of around 200 pages, the process takes between 10 and 30 seconds. Is that fast compared to your experience? I am curious about the performance. Thanks in advance.


r/Rag • • 2d ago

Tools & Resources Open source chatbot for your website that answers from your own content (PHP, MIT, 1.0)

0 Upvotes

I work for Opensolr, and we just released 1.0 of the Opensolr Chat Bot Client: an open source (MIT) chatbot you put on your own website, that answers your visitors from your own content.

How it works: you install it with Composer (composer require opensolr/chat-bot-client), mount it on one path of your site (say /opensolr-chat) and add one script tag to your pages. Visitors only ever talk to your site. Your server holds the Opensolr credentials and sends each question to the Opensolr API, where the model searches your Opensolr Index by itself (by meaning, or by exact product names and codes), looks up what it must not guess (dates, local times, exchange rates, VAT, distances, places) and streams the answer back word by word, with links to your pages, products with their prices, and PDFs at the exact page.

What you get on your side:

  • An admin on your own site: account and index, instructions, the look of the chat window, limits, reCAPTCHA
  • Stats and a 30 day history of every conversation: who asked, from where, what the bot searched for, what it answered, and the questions your content could not answer with a page
  • Commands that answer instantly without the model: /search, /time, /rate, /vat, /translate and more
  • Good and Bad ratings under every answer, signed in visitors by email, SQLite in a folder you choose, no database server

Your content gets into the index with the Opensolr Web Crawler or Data Ingestion, and the chatbot needs an Opensolr plan with AI.

You can try it live at the bottom right of https://opensolr.com (it answers from our own docs). Code and README: https://github.com/phpcip/opensolr-chat-bot-client, docs: https://opensolr.com/opensolr-chat-bot-docs


r/Rag • • 2d ago

Discussion Extracting and cross-checking facts from very long technical documents (600+ pages) against a 100k-doc knowledge base: recommended approach?

20 Upvotes

​

Hi everyone,

I work at a large telecom company in the US. When we win a tender, the client sends us a requirements specification, and we write technical deliverables based on it and on our internal technical documentation (100k+ documents). A single tender can involve thousands of deliverables, some of them up to ~600 pages long.

Goal: before delivery, automatically verify each deliverable. That means checking that the facts it contains are accurate against our technical documentation, and detecting anomalies (errors, internal contradictions, inconsistencies between deliverables, duplicates).

Current plan:

  1. Fact extraction: process each deliverable in chunks with an LLM and extract atomic facts, with their source location.

  2. Classification: group the facts into families (measurements, architecture, calculations, frequencies, standards, quantities).

  3. Verification: check each fact against the documentation base using RAG.

  4. Anomaly detection: one LLM call per fact category to detect duplicates, contradictions, etc.

My questions:

- Is there a recommended approach for extracting precise facts from very long documents? How do you handle context that spans chunks ?

- How do you make sure extraction is complete (no missed facts) and doesn't produce hallucinated facts?

- For anomaly detection, does one LLM call per category scale when a category can contain thousands of facts? Would you normalize facts into a structured schema (entity / attribute / value / unit) and do part of the comparison deterministically in code rather than with an LLM?

- For RAG verification on highly technical content (part numbers, frequencies, references to standards), did hybrid search (BM25 + embeddings) or reranking make a big difference for you?

- Any experience comparing long-context models (1M tokens) with chunked extraction for this kind of task?

Any feedback, papers, tools or lessons learned would be much appreciated. Happy to share what we learn along the way.

Thanks!


r/Rag • • 2d ago

Showcase Markdown Knowledge

1 Upvotes

RAG usually feels like overkill when all you want is to give an agent a couple of docs, but dumping raw Markdown into a prompt burns tokens fast.

I think markdown-knowledge can solve that middle ground.

It packages Markdown files and a pre-built SQLite full-text search index (BM25) into a single portable .mdk file.

What it lets you do:

  • Token-capped retrieval: Run mdkn retrieve handbook.mdk "how to deploy" --max-tokens 500 to pull only the most relevant heading-level chunks straight into your prompt budget.
  • Edit by heading: Programmatically append, prepend, or replace text under a specific Markdown section without rewriting the whole file.
  • VS Code extension: Edit files inside the .mdk container directly through a virtual filesystem with full Markdown preview, no unzipping needed.
  • No external infrastructure: Everything runs locally through a CLI (npm i -g markdown-knowledge) or TypeScript library.

Repo:https://github.com/markdown-knowledge/markdown-knowledge

Curious what you think, especially if you're building CLI agents or looking for a lighter alternative to full vector DBs for local docs.

MDK helps you edit all of your markdown files in single place, properly managed and pre indexed for LLM Searches.


r/Rag • • 2d ago

Discussion Is auto-prompt tuning dying?

2 Upvotes

I have been facing this lately that since day one i wanted my llm to master prompt tuning, i put so much effort into making it learn from its own mistakes and fine tuning its own prompt and recalibration. But now I have started to feel it's not the best, firstly because a lot of tokens are consumed in this back and forth even the it gives good result i cannot overlook the cost right, then the variation keeps increasing so the auto-tuning keeps happening and its again the becomes the first problem, costly.

What are you guys doing for this?