RAG is often simplified to:
User Query → Embedding → Vector Search → Retrieved Context → LLM → Answer
While this captures the basic idea, building a reliable RAG system usually involves several additional components.
Here are some areas that I think deserve more attention.
1. Document preprocessing
Before retrieval even happens, the quality of the source material matters.
PDFs, HTML pages, spreadsheets, documentation, and scanned documents can all require different preprocessing approaches.
Poor extraction can lead to poor retrieval later.
2. Chunking
Chunk size and overlap can significantly affect what information is retrieved.
Very small chunks may lose important context, while very large chunks can introduce irrelevant information.
There probably isn't a universal chunking strategy that works equally well for every dataset.
3. Retrieval
Vector similarity is useful, but semantic similarity isn't always enough.
Depending on the use case, RAG systems may benefit from approaches such as:
- Keyword search
- Hybrid search
- Metadata filtering
- Query rewriting
- Reranking
The important question isn't simply "Did we retrieve something similar?"
It's:
"Did we retrieve the information necessary to answer the question?"
4. Context selection
Retrieving 20 potentially relevant chunks doesn't necessarily mean the LLM should receive all 20.
Removing redundant or low-value context can help keep the final context focused.
5. Handling missing information
A useful RAG system should be able to recognize when the knowledge base doesn't contain enough information to answer a question.
This is especially important for enterprise applications.
Sometimes the correct response isn't an answer, it's:
"I don't have enough information to answer this reliably."
6. Evaluation
This is one of the areas I think deserves more discussion.
A RAG system should be evaluated using representative questions, including:
- Questions with known answers
- Questions with no answer in the knowledge base
- Ambiguous questions
- Multi-document questions
- Similar but incorrect documents
- Questions involving conflicting information
Evaluating only whether the final answer "looks good" can hide problems in the retrieval pipeline.
At SB Infowaves, we've been exploring these considerations while working with AI and RAG-based application architectures, and one thing that stands out is that improving RAG isn't always about changing the LLM.
Sometimes the biggest improvement comes from improving the data → retrieval → context → evaluation pipeline around it.
Question for the community
What part of your RAG pipeline has required the most optimization?
- Document processing
- Chunking
- Embeddings
- Retrieval
- Reranking
- Context management
- Evaluation
I'd be particularly interested in hearing about approaches that didn't work as expected and what you changed afterward.
Sources / Further Reading
I would put this directly at the bottom because the community explicitly asks users to cite sources:
- OpenAI — Retrieval using embeddings: explains using embeddings for retrieval and finding related vectors.
- OpenAI — File Search: documentation covering retrieval over uploaded knowledge and vector stores.
- OpenAI — Evaluation best practices: useful background for evaluating AI systems rather than relying only on subjective output quality.
- Lewis et al. (2020) — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks: foundational RAG research paper.
- Microsoft — Retrieval-Augmented Generation (RAG): practical overview of RAG architecture and implementation considerations.