RAG: letting AI answer based on your own documents
Be able to explain why the answer improves when the AI gets the right document — RAG in plain words.
Prerequisites
Everyday explanation
An AI knows a lot in general, but nothing about your company. It does not know your internal processes, your customer contracts, or what you decided in last week’s meeting.
The solution is simple: give it the text first.
WITHOUT: "What is our cancellation policy?"
→ "It varies between companies. Check your contract."
WITH: [here is our internal handbook: ...]
"What is our cancellation policy?"
→ "According to the handbook, it is 30 days’ notice."
Same model, same question — completely different answer. The difference is that the second time, it had the source material.
But you cannot paste everything in. If you have 500 pages of documents, they will not fit in the question. So someone must select the pages that relate to your specific question.
And that is exactly what a search engine does. That is why the search engine is the prerequisite: this technique is a search engine plus an AI, connected.
Intuition
Four steps, every time you ask a question:
| Step | What happens |
|---|---|
| 1. Search | find the parts in your texts that resemble the question |
| 2. Select | take the three to five best ones |
| 3. Paste in | put them in the prompt together with the question |
| 4. Answer | the model answers based on what is there |
The method is called RAG — retrieval-augmented generation, roughly “generation with retrieved support”.
Three benefits:
| Benefit | Why |
|---|---|
| The answer is based on your texts | not on what the model happens to remember |
| You can verify | the source is shown next to the answer |
| Updating is easy | swap the document, nothing needs retraining |
But: if the wrong parts are retrieved, the answer will be wrong — and it will still sound just as confident. The model cannot know that it received the wrong source material.
Two ways to search:
| Method | Finds | Misses |
|---|---|---|
| Keyword matching | exact same words | rephrasings: “cancellation policy” vs “rules for refunds” |
| Semantic similarity | rephrasings and synonyms | exact codes and names sometimes |
The second one is based on converting texts into number series where similar meanings yield similar numbers. The best approach is to use both.
Interactive
Do it by hand — it works, and it teaches you more than reading about it.
Preparation. Take five pages from your own internal documents or a handbook. Number the paragraphs.
Round 1 — without retrieval. Ask three specific questions about the content to a chatbot, without pasting anything in.
Round 2 — with retrieval. For each question:
- Find the paragraph that answers it yourself (you are the search engine).
- Paste in the paragraph, then the question.
- Compare the answer with Round 1.
Round 3 — wrong source material, on purpose. Paste in a paragraph that does not relate to the question and ask the question anyway.
What you will see:
| Round | Typical result |
|---|---|
| 1 | General, often evasive, sometimes invented |
| 2 | Specific and correct, with phrasing from your text |
| 3 | This is the interesting part — see below |
Round 3 is the point of the experiment. The model usually does one of three things: answers based on the irrelevant paragraph (wrong answer, confident tone), says the paragraph does not contain the answer (good!), or mixes the paragraph with its own general knowledge (worst, because it is hardest to detect).
Conclusion: the quality of the entire system depends on the right paragraph being retrieved. The AI cannot fix a bad search — and therefore it is the search part, not the model, that you most often need to improve.
Mastery means
- Explains why retrieved documents improve the answer
- Describes the steps in order
- Knows what happens when the wrong document is retrieved
Sign in to do the exercises and build your mastery up.
Sources
- Internetstiftelsen — Internetkunskap — free to read
- Skolverket — About AI in school (in Swedish) — Skolverket's open terms
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0