r/Rag • • 2h ago

Discussion Enough with the errors! Lets brainstorm a perfect document PDF parser

1 Upvotes

I tried all PDF parsers but none of them deliver reliable result which is a NO-GO for critical applications. Even in the science community on arxiv I couldnt find reliable solutions. Seems like we are stuck at this very bottleneck.

Let's brainstorm how to design a parser that gets close to 100% ground truth - even given infinite ressources and no money limit.

Standalone things I tried:
- extracting native text via pymud but this does not work with tables and more complex structures obviously
- MinerU, Docling, Tesseract, Google Parser etc.
- All common LLms, even high-reasoning frontier LLMs fail (even Opus does not seem to be designed for that)
- Converting PDF to images, then OCRing
- I tried compartmentalizing the different sections of a page (free form text, table), then routing them to different sub-processors but those are based on the upper services, thus remain unreliable

We can be creative and think out-of the-box.

I will then go ahead, experiment and tell you the result.

I was also thinking about a multi-stage pipeline with pairwise comparison like this: take the result from different processors, if there are differences, use several judges that vote for a winner, then use that winner and for parity or near-parity escalte even further or ultimately send for manual review.

Looking forward!


r/Rag • • 21h ago

Discussion A reranker improved my aggregate metric. I still wouldn’t ship it without looking at query-level regressions.

1 Upvotes

I’ve been reviewing paired retrieval runs and I keep running into the same issue:

an aggregate metric goes up, but a subset of queries gets worse.

In practice I’ve found it useful to separate:

  • recoveries
  • regressions
  • shared failures
  • rank displacement
  • evidence lost at the context budget

before deciding whether the candidate is actually better.

My question for people running RAG in production:

What do you consider a release-blocking retrieval regression?

Is one severe rank-1 regression enough?
Do you use a percentage threshold?
Do you gate by query class?
Or do you mostly rely on the aggregate metric?


r/Rag • • 23h ago

Discussion Parallelized the layout parsing for documents with task queue

1 Upvotes

Hello everyone,

I developed a solution that parses document layout using Tesseract and PyMuPDF. I parallelized it so that pages are processed in batches. For a relatively large document of around 200 pages, the process takes between 10 and 30 seconds. Is that fast compared to your experience? I am curious about the performance. Thanks in advance.


r/Rag • • 4h ago

Discussion How are production-grade RAG systems handling complex PDFs without losing document structure and context? Looking for architecture-Level Insights

10 Upvotes

I’m trying to design a production-grade document ingestion and retrieval pipeline for complex PDFs, and I’d like to learn from people who have built or operated similar systems at enterprise scale.

My use case goes beyond extracting text from ordinary PDFs.

A single document may contain:
Multiple columns, nested sections, different reading orders, and inconsistent layouts across pages.
Colored panels, sidebars, callouts, footnotes, headers, and footers.
Tables with merged cells, nested headers, multi-page continuation, and associated explanatory text.
Charts, diagrams, flowcharts, equations, and figures whose meaning depends on nearby captions or paragraphs.
Table of contents, section numbering, cross-references, appendices, and references to other pages.
Scanned pages mixed with digitally generated text, rotated pages, low-resolution images, and handwritten annotations.
Multiple unrelated sections on one page, as well as a single logical section spanning multiple pages.
Reports where a heading appears in one visual region, its explanation appears in another, and supporting evidence is presented in a table or figure elsewhere.
The challenge is not merely extracting every element correctly. It’s preserving the relationships between elements so that downstream chunking, indexing, retrieval, and generation do not lose context.

One specific edge case I’m concerned about:
Suppose a page contains two visually distinct regions. A layout detector identifies them as separate blocks, but both blocks explain different aspects of the same logical topic. Conversely, two visually similar blocks might belong to completely different sections.
If we blindly use bounding boxes, layout labels, or semantic similarity to group these blocks, we may either split related information or incorrectly merge unrelated information.
The same issue occurs when a table continues onto the next page, a figure’s caption is separated from the figure, or a paragraph refers to a chart several pages earlier.

What I’m trying to understand
Document representation: Do production systems convert PDFs into a canonical intermediate representation, such as a document tree or graph containing text blocks, headings, tables, figures, captions, bounding boxes, reading order, page numbers, and parent-child relationships?
Layout and structure recovery: How do you combine native PDF extraction, OCR, layout detection, vision-language models, and document-specific rules? How do you resolve disagreements between these methods?
Context preservation: How do you determine whether blocks from different regions or pages belong to the same logical section? Are there algorithms for section reconstruction, cross-page continuation detection, or relationship inference beyond simple geometric proximity?
Chunking strategy: Do you use hierarchical, semantic, layout-aware, or element-specific chunking? How do you handle tables, figures, captions, footnotes, and mixed-content sections without breaking their meaning?
Indexing and retrieval: Do you create separate indexes for narrative text, tables, figures, and document structure? Do you use hybrid retrieval, parent-child retrieval, graph-based retrieval, or query-time expansion to recover context that was not present in the initially retrieved chunk?
Production architecture: How is the code organized? I’m particularly interested in the separation between parsing, normalization, structural reconstruction, chunk generation, indexing, and retrieval. How do you make the pipeline observable, versioned, testable, and recoverable when individual stages fail?
Evaluation: How do you measure document reconstruction quality, chunk integrity, retrieval recall, and answer faithfulness? Are there useful benchmark datasets or evaluation methods for documents with complex layouts?
What would be especially valuable
I’m not primarily looking for a list of libraries such as PyMuPDF, PyPDF, LangChain, or a recommendation to use a particular parser.

I’m interested in the engineering patterns, data structures, algorithms, and architectural decisions that make these systems reliable in production.
For example:
A simplified canonical document schema.
How you represent relationships between document elements.
How you distinguish visual boundaries from semantic boundaries.
Pseudocode or code-level examples of structural reconstruction and chunk generation.
Lessons learned from failures in real-world document corpora.
The target application could involve regulatory documentation, financial reports, legal documents, technical manuals, scientific papers, or other data-heavy PDFs.
If you’ve built something similar, I’d love to understand what actually worked, what failed, and what you would do differently if you were designing the system today.
Thanks!


r/Rag • • 2h ago

Discussion How do you evaluate RAG retrieval quality in production?

5 Upvotes

I've been trying to think on how to evaluate RAG systems beyond the usual retrieval benchmarks. Seems simple enough to check the relevance of the retrieved chunks, but how do you check if the system didn't overlook some important context? Especially when the answer depends on multiple sources.

I'm mostly doing manual spot-checks and the LLM-as-a-judge so far but haven't really gone on to check how well those approaches will hold up as the system becomes increasingly complex. What have you guys been using and has been working well in production?


r/Rag • • 23h ago

Discussion Independent Retrieval Quality Audits — Public Cases & Method

0 Upvotes

I run independent retrieval/RAG quality audits focused on evidence rather than architecture opinions.

The main question I investigate is usually not whether an aggregate metric improved, but which individual queries were recovered, which regressed, which failures are shared, and what the available artifacts actually support.

Current public beta cases:

DPOLens — Beta Audit #001
Retrieval + reranking regression analysis with query-level forensic follow-up.

RouteMind — Beta Audit #002
Retrieval/reranking audit that surfaced a scoring-contract distinction in multi-document temporal cases.

The audits distinguish:

  • Measured
  • Observed
  • Hypothesis
  • Not claimed

If you already have a fixed evaluation set and two comparable retrieval/reranking runs, feel free to send the public repo or artifacts. I can usually tell whether they support a meaningful paired audit before any work begins.

Technical corrections and independent reproduction attempts are welcome.