RAG Has Its Own Garbage In, Garbage Out Problem
Published ยท October 4, 2026
Every data scientist knows the phrase "garbage in, garbage out", and must have faced the issue at some point in their career. A sophisticated model trained on poor data will still produce poor predictions. RAG systems have the same problem, but it shows up in a different place.
The failure moves from training time to answer time
In classic ML, bad data damages the model during training. In RAG, the model stays the same. The damage happens at answer time. Poor documents get retrieved and placed in front of the LLM, and the LLM does what it is built to do. The LLM now has that content in its context, and it may use it as evidence when generating the answer.
This is also good news. You cannot easily repair a model trained on bad data. But you can repair a RAG system by repairing its documents.
Relevant is not the same as correct
Most RAG tuning goes into chunk sizes, embedding models, vector databases, hybrid search and rerankers. These matter, but they all do one job, which is to find the most relevant content. They have no idea whether that content is right.
Here are four ways retrieval can work correctly and still answer badly -
- Outdated documents. An old process guide matches the question best, so it ranks first.
- Near-duplicates. Five copies of the same policy fill the top results and push out other useful context.
- Contradictions. Two versions disagree, both are retrieved, and the model blends them into an answer that matches neither.
- Poorly written content. Vague or badly structured text breaks into weak chunks that give the model little to work with.
A better retriever won't fix these. It just gets better at finding the wrong thing.
Smaller and cleaner often wins
Many teams start by loading everything in, from wiki pages to PDFs to shared drives. More content feels like more capability.
In my experience, a smaller set of well-written, current, owned documents often beats a much larger messy one. Every weak document adds noise, and every noisy chunk is a chance to mislead the model.
A practical solution - let user questions decide what matters
When answers are inconsistent because the knowledge base holds overlapping and outdated documents, tuning the retriever again rarely helps. A better fix is to analyze the themes in real user questions and write a reviewed set of FAQs for the most common ones. The questions users actually ask decide what goes into the knowledge base.
The design choice that makes this work is simple - duplicate what's stable, reference what's volatile.
- Facts that rarely change - like definitions, eligibility rules and how a process works, live inside the FAQ answer.
- Details that change often - like limits, dates and thresholds, are not copied. The FAQ points to the source document by a stable ID, and the pipeline fetches the current version at query time. A "references" section at the end of each answer also lets users open the source pages or documents and read in depth.
This keeps answers short and consistent, and there is still only one place to update when something changes. The FAQs then become the preferred source in retrieval, and the old overlapping documents are deprecated for those topics.
A few details make this approach hold up over time -
- Each FAQ entry has an owner, a review date and a tag showing whether it is fully stable or depends on a live document.
- References use document IDs, not file paths, so renaming or moving a file doesn't break them.
- If a referenced document is changed or retired, the related FAQ is flagged for review.
- Question themes are re-analyzed regularly, since what users ask about shifts over time.
- Long-tail questions still fall back to the underlying documents.
RAG is partly a knowledge management problem
Document quality is not a prep step that finishes before the AI work begins. It shapes retrieval and the final answer every day. That means someone has to run the knowledge like a product -
- Ownership. Every document has a named owner.
- Metadata. Track owner, version, effective date, last review and status (active, draft, deprecated).
- Freshness rules. Filter or down-weight expired content at retrieval time, not only at ingestion.
- Deduplication. Keep one canonical version of each document.
- Source-of-truth rules. When two documents disagree, a clear rule decides which one wins.
- Deliberate removal. Decide what doesn't belong, and retire it on purpose.
- Feedback loops. Trace bad answers back to the documents behind them.
Ask where the answer failed
When an answer is wrong, ask where it failed. Was it retrieval, generation or the source document? Teams often assume the first two. In my experience, the source is more common than expected, and it is the cheapest to fix.
Track this over time. Useful measures include the share of wrong answers traced to source content, the age of the most-retrieved documents, and how many top questions are covered by a reviewed FAQ.
The real question
We often say RAG gives an LLM access to enterprise knowledge. The better question is whether that knowledge is good enough to rely on.
A useful way to think about RAG quality is that retrieval answers "Did we find the right information?" while knowledge management answers "Was the information worth finding in the first place?"
The model can retrieve what we give it. It cannot turn poorly maintained knowledge into reliable knowledge.