top of page

Please Don’t Feed the RAG Bear... Everything

Writer: Elizabeth-hadley Rich
Elizabeth-hadley Rich
Jun 30
4 min read

I keep coming back to this idea with enterprise AI:


The model does not need everything. It needs the right things.


That sounds obvious, maybe. But I think it gets lost quickly once organizations start talking about large language models, internal chatbots, RAG models, and AI-supported knowledge experiences.


Most organization's first instinct is to connect the model to as much as possible:

  • Policies

  • FAQs

  • Product documentation

  • Training materials

  • Knowledge articles

  • Old pages that might still be useful someday

  • That one PDF everyone keeps forwarding because no one knows where the real answer lives


I understand the impulse, I really do. If the model can retrieve more information, then surely it can answer more questions.


But more is not always better. Sometimes more just gives the model more ways to be wrong.


RAG is really about control


Retrieval-augmented generation, or RAG, is usually described as a way to connect a model to a defined body of knowledge.


That's true, but the part I find more interesting is control.


RAG gives us a way to create a smaller, cleaner, more intentional knowledge environment for the model to work inside.


Because not everything an organization has produced should be retrievable.

Not everything is current.Not everything is trusted.Not everything is governed.Not everything is clear enough to stand on its own.Not everything is safe to use in an answer.


This is where the work gets less flashy, but much more useful.


I saw this most clearly at the chunk level


In one enterprise AI-readiness effort, the conversation got much more concrete when we stopped talking about “content” as one giant abstract thing. We had to get closer to the retrieval layer.


Not just:

Is this page good?


But:

  • Can this chunk of information stand on its own?

  • Is the source authoritative?

  • Is it current?

  • Can a human verify the answer?

  • Is there enough metadata to help the system retrieve it appropriately?

  • Would this still make sense inside a chatbot, internal tool, support workflow, or future product experience?


And that changed the question.


The work was no longer just “clean up the content.”


It became:

What makes knowledge usable by a model?


That feels like the real AI-readiness question to me.


The model needs signals. It needs boundaries. It needs some way to know what is worth using and what is not.


Smaller can be smarter


Most organizations do not have a knowledge shortage. They have knowledge everywhere. Old documents that are technically still live. Five versions of the same answer. Policy pages that conflict with training materials. Product definitions that vary by team. PDFs that became unofficial sources of truth because someone shared them once and they never went away.


We work around this all the time, don't we?


We know who to ask. We know which page is outdated. We know which team owns the real answer. We know when the official language and the usable language are not quite the same thing.


A model does not know any of that by default.


So when we connect an LLM to a messy knowledge environment, we should not be surprised when it reflects that mess back to us.


Sometimes fluently, sometimes confidentl. Sometimes in ways that feel almost right, which may be the riskiest version of wrong.


This is where heuristics matter


A good RAG model needs more than a pile of approved documents. It needs practical rules of judgment.


Things like:

  • Use the most recent approved policy over older training content.

  • Prefer customer-facing language for customer-facing questions.

  • Do not answer compliance-sensitive questions unless the source is explicitly approved.

  • Summarize only from retrieved content.

  • Escalate when sources conflict.

  • Ask for clarification when the user’s intent is too broad.


None of this sounds especially magical. But that's kind of the point. A lot of the work that makes AI useful is not magical. It's deciding what good looks like before asking the model to produce it. It's caution over all content.


The answer is only the visible part


It's easy to look at the generated answer and think that's the AI experience. But the answer is only what we can see.


Underneath it are the quieter decisions:

  • What sources were included

  • What sources were excluded

  • How the content was chunked

  • What metadata was attached

  • What the model should do when sources conflict

  • Who owns the knowledge after launch


That last one matters more than people think. Because AI readiness is not a launch state. It's a maintenance model.


The knowledge will change. The products will change. The policies will change. The edge cases will change. The organization will change.


So the system needs stewardship. Not just implementation.


The model does not need everything


The model does not need everything. It needs the right knowledge, in the right shape, with the right rules around it. That means saying no to abandoned content, outdated pages, duplicate answers, “just in case” documents, and sources no one owns.


It also means recognizing that RAG is not where old content goes to die. It's where governed knowledge can become usable in a new way.


So no, I do not think AI readiness starts by feeding the model everything. I think it starts by deciding what is worth knowing. And then building a system that helps the model use that knowledge well.

 
 
 

Comments


bottom of page