Skip to content
Writing
Esc
Type to search all articles.
    All writing

    Retrieval-Augmented Generation, Explained

    How RAG gives a language model access to knowledge it was never trained on, and the design choices that decide whether it works.

    DS

    Debanjan Saha

    · 3 min read

    – claps
    On this page

    A language model only knows what was in its training data, and only up to a cutoff. Ask about your company’s internal wiki or last week’s news and it will either say it doesn’t know or, worse, invent something plausible. Retrieval-augmented generation (RAG) fixes this by fetching relevant documents at question time and placing them in the prompt.

    The pipeline in one picture

    RAG has two phases.

    Indexing (offline):

    1. Split documents into chunks.
    2. Convert each chunk into an embedding, a vector that captures its meaning.
    3. Store the vectors in an index, usually a vector database.

    Querying (online):

    1. Embed the user’s question with the same model.
    2. Find the nearest chunks by vector similarity.
    3. Put those chunks and the question into the prompt.
    4. Let the model answer using that context.
    def answer(question, index, embed, llm, k=5):
        query_vec = embed(question)
        chunks = index.search(query_vec, top_k=k)
    
        context = "\n\n".join(c.text for c in chunks)
        prompt = (
            "Answer using only the context below. "
            "If the answer is not there, say you don't know.\n\n"
            f"Context:\n{context}\n\nQuestion: {question}"
        )
        return llm(prompt), [c.source for c in chunks]

    Returning the sources alongside the answer lets users verify claims, which is half the point of RAG.

    Chunking is where most systems succeed or fail

    Chunks that are too large dilute the signal and waste context. Chunks that are too small lose the surrounding meaning. Good defaults:

    • Split on natural boundaries such as headings and paragraphs rather than a fixed character count.
    • Add a small overlap so a sentence cut at a boundary still appears whole somewhere.
    • Attach metadata, like the title, section and date, to every chunk.

    Retrieval: similarity is not relevance

    Vector search finds text that is semantically close, which is not always what answers the question. It can miss exact identifiers, error codes or rare names, where traditional keyword search excels. Many production systems therefore use hybrid search, combining keyword scoring such as BM25 with vector similarity, then merging the rankings.

    A reranker adds a further step: retrieve a generous set of candidates cheaply, then use a more accurate model to reorder them and keep only the best few.

    Where RAG goes wrong

    • Retrieval misses. The right passage exists but isn’t returned. No amount of prompting helps if the model never sees it.
    • Distraction. Irrelevant chunks in the prompt can pull the answer off course.
    • Stale data. The index drifts out of date unless re-indexing is part of the pipeline.
    • Unfaithful answers. The model ignores the context and answers from memory, so instructions and citation requirements matter.

    Measuring it

    Evaluate retrieval and generation separately; otherwise you can’t tell which half is failing.

    StageQuestionExample metric
    RetrievalDid the right chunks come back?Recall at k, mean reciprocal rank
    GenerationIs the answer grounded in them?Faithfulness, citation accuracy
    End to endDid the user get a correct answer?Human-graded accuracy

    Build a small set of real questions with known source passages. It will teach you more than any benchmark.

    RAG or long context?

    Context windows are now large enough to paste in whole documents, so is retrieval obsolete? Not usually. Retrieval is cheaper per query, scales to corpora far beyond any window, and keeps irrelevant text out of the model’s attention. Long context shines for a handful of known documents; RAG wins when the knowledge base is large, changing or access-controlled.

    Enjoyed this?

    A clap helps others find it.

    – claps

    Discussion

    Comments (Giscus) will appear here. Set PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY and PUBLIC_GISCUS_CATEGORY_ID to enable them.

    Keep reading