Retrieval Augmented Generation: The Future of Context-Aware AI

How retrieval-augmented generation works, when it beats fine-tuning, and the architecture decisions that make RAG reliable in production systems.
Retrieval Augmented Generation: The Future of Context-Aware AI
A language model on its own is a very well-read colleague who has never seen your company. It knows a great deal about the world and nothing about your prices, your policies or the booking you took yesterday. Retrieval-augmented generation (RAG) is how you introduce the two.
What RAG actually does
The pattern is four steps, and every production system is a variation on it:
- Index. Split your documents into chunks and store a vector embedding of each one.
- Retrieve. When a question arrives, embed it and pull the handful of chunks closest in meaning.
- Augment. Paste those chunks into the prompt as source material.
- Generate. Ask the model to answer only from that material, with citations.
That last instruction is the whole trick. A model told to answer from supplied sources — and to say "not in the documents" otherwise — hallucinates far less than one asked from memory.
RAG versus fine-tuning
| RAG | Fine-tuning | |
|---|---|---|
| Teaches | Facts | Style and format |
| Update cost | Re-index a document | Retrain the model |
| Citations | Natural | Impossible |
| Setup time | Days | Weeks |
| Best for | Docs, policies, catalogues, tickets | Consistent tone, strict output shapes |
Most teams that think they need fine-tuning need retrieval. The exceptions are covered on our AI model training page.
The parts that decide whether it works
Chunking. Too small and a chunk loses its context; too large and the retrieval is imprecise. Split on document structure — headings, sections, table rows — not on a fixed character count. Overlap by a sentence or two.
Hybrid search. Pure vector search misses exact matches: product codes, invoice numbers, surnames. Run keyword search alongside it and merge the results. This one change fixes more "the bot cannot find it" complaints than any other.
Re-ranking. Retrieve twenty candidates cheaply, then use a re-ranker to pick the best five. Precision at the top of the list matters more than recall across the whole index.
Metadata filters. Every chunk should carry its source, date and access level. Filter before you search — that is also how you stop one customer's data reaching another, which matters under POPIA.
Freshness. An index that updates monthly will confidently quote last month's prices. Re-index on write, not on a schedule.
A reference architecture
For a South African SME chatbot the stack usually looks like this: Postgres with the pgvector extension for storage (no separate vector database needed at this scale), an embedding model called at write time, hybrid retrieval in a single SQL query, and a serverless function that assembles the prompt and streams the answer. Cheap to run, easy to debug, and the data stays in one place you already back up.
How to measure it
Evaluate retrieval and generation separately or you will never know which one broke.
- Retrieval: for 100 real questions, was the correct chunk in the top five? Anything below 90% and no amount of prompt engineering will save you.
- Generation: was the answer supported by the retrieved text? Grade it, keep the score, watch the trend.
- Refusals: the system should decline when the answer genuinely is not in the corpus. A refusal rate of zero means it is making things up.
Frequently asked questions
How much content do I need before RAG is worth it?
Twenty solid pages is enough to beat a static FAQ. Under that, a well-written prompt with the facts inline is simpler and cheaper.
Does RAG eliminate hallucination?
It reduces it substantially and makes it detectable, because every claim should map to a citation you can check. Nothing eliminates it entirely.
What does a production RAG assistant cost to run?
For a typical SME volume — a few thousand questions a month — inference and embedding costs usually land in the low hundreds of Rand. The build is the expense, not the running.
Continue Reading
- Claude 4: Capabilities, Features & What It Means for AI Development
- AI Chatbot Training & Models Guide (2026)
- AI for Small Business in South Africa
Want one built on your documents? See AI chatbot development or get in touch.
LET’S TALK
Have a project in mind after reading this? Send a brief and I usually reply within 24 to 48 hours.

