TENDAI . G
Back to journal
AI & Technology21 May 2026

Retrieval Augmented Generation: The Future of Context-Aware AI

Tendai Gumunyu12 min read
Retrieval Augmented Generation: The Future of Context-Aware AI

How retrieval-augmented generation works, when it beats fine-tuning, and the architecture decisions that make RAG reliable in production systems.

Retrieval Augmented Generation: The Future of Context-Aware AI

A language model on its own is a very well-read colleague who has never seen your company. It knows a great deal about the world and nothing about your prices, your policies or the booking you took yesterday. Retrieval-augmented generation (RAG) is how you introduce the two.

What RAG actually does

The pattern is four steps, and every production system is a variation on it:

  1. Index. Split your documents into chunks and store a vector embedding of each one.
  2. Retrieve. When a question arrives, embed it and pull the handful of chunks closest in meaning.
  3. Augment. Paste those chunks into the prompt as source material.
  4. Generate. Ask the model to answer only from that material, with citations.

That last instruction is the whole trick. A model told to answer from supplied sources — and to say "not in the documents" otherwise — hallucinates far less than one asked from memory.

RAG versus fine-tuning

RAGFine-tuning
TeachesFactsStyle and format
Update costRe-index a documentRetrain the model
CitationsNaturalImpossible
Setup timeDaysWeeks
Best forDocs, policies, catalogues, ticketsConsistent tone, strict output shapes

Most teams that think they need fine-tuning need retrieval. The exceptions are covered on our AI model training page.

The parts that decide whether it works

Chunking. Too small and a chunk loses its context; too large and the retrieval is imprecise. Split on document structure — headings, sections, table rows — not on a fixed character count. Overlap by a sentence or two.

Hybrid search. Pure vector search misses exact matches: product codes, invoice numbers, surnames. Run keyword search alongside it and merge the results. This one change fixes more "the bot cannot find it" complaints than any other.

Re-ranking. Retrieve twenty candidates cheaply, then use a re-ranker to pick the best five. Precision at the top of the list matters more than recall across the whole index.

Metadata filters. Every chunk should carry its source, date and access level. Filter before you search — that is also how you stop one customer's data reaching another, which matters under POPIA.

Freshness. An index that updates monthly will confidently quote last month's prices. Re-index on write, not on a schedule.

A reference architecture

For a South African SME chatbot the stack usually looks like this: Postgres with the pgvector extension for storage (no separate vector database needed at this scale), an embedding model called at write time, hybrid retrieval in a single SQL query, and a serverless function that assembles the prompt and streams the answer. Cheap to run, easy to debug, and the data stays in one place you already back up.

How to measure it

Evaluate retrieval and generation separately or you will never know which one broke.

  • Retrieval: for 100 real questions, was the correct chunk in the top five? Anything below 90% and no amount of prompt engineering will save you.
  • Generation: was the answer supported by the retrieved text? Grade it, keep the score, watch the trend.
  • Refusals: the system should decline when the answer genuinely is not in the corpus. A refusal rate of zero means it is making things up.

Frequently asked questions

How much content do I need before RAG is worth it?

Twenty solid pages is enough to beat a static FAQ. Under that, a well-written prompt with the facts inline is simpler and cheaper.

Does RAG eliminate hallucination?

It reduces it substantially and makes it detectable, because every claim should map to a citation you can check. Nothing eliminates it entirely.

What does a production RAG assistant cost to run?

For a typical SME volume — a few thousand questions a month — inference and embedding costs usually land in the low hundreds of Rand. The build is the expense, not the running.

Continue Reading

Want one built on your documents? See AI chatbot development or get in touch.

RAGAILangChainVector DatabasesLLMMachine LearningPython
Work with me

LET’S TALK

Have a project in mind after reading this? Send a brief and I usually reply within 24 to 48 hours.

Start a project →

More reading

All articles