RAG vs Long Context

A bigger context window — the amount of text a model can take in for one request — shrinks the case for RAG on small knowledge bases. It doesn't remove the case for RAG on large ones, ones that change often, or ones where you need to show exactly which source an answer came from. If your material is small enough to paste directly into a request and stays roughly the same, long context can genuinely replace retrieval. If it isn't, retrieval is still doing real work that a bigger window doesn't do for you.

What RAG does

RAG searches your documents for the passages relevant to the current question and hands the model only those, alongside the question itself. The model never sees the whole corpus — just a small, chosen slice of it. The RAG page covers the pipeline stage by stage.

What long context does

Long context skips the search step entirely. Instead of retrieving a subset, you hand the model a much larger piece of the material directly — sometimes the whole document, sometimes the whole corpus if it's small enough — and let the model read all of it itself. There's no separate page for this because it isn't really a technique of its own; it's the absence of one, made possible by a bigger context window.

Side by side

RAGLong context
What the model seesA small, retrieved subsetMost or all of the source material
Per-query costLower — only the relevant passages are sentHigher — the same large block gets sent on every query
LatencyRetrieval step adds a little, but the model reads lessNo retrieval step, but the model reads far more
CitationsYes — each answer can point at the passage it came fromWeaker — with no retrieval step marking what mattered, tracing an answer back to one spot is harder
FreshnessUpdate the index; the next query sees the changeRepaste or re-supply the updated material every time
Ceiling as the corpus growsScales to any size — you're always sending the same small sliceHits the window's limit eventually, no matter how large that limit is
Reliability across the inputNot applicable — the model only sees the chosen sliceUneven — models use text near the start or end of a long input more reliably than text buried in the middle

Which one your problem calls for

Long context fits when the material is genuinely small enough to paste in comfortably, doesn't change often, and you're not running enough queries for the extra tokens on every single one to add up. A single support ticket answered against one 40-page manual is a reasonable case: paste the manual, ask the question, done — building a retrieval pipeline for that is real machinery solving a problem you don't have yet.

RAG fits once the corpus is larger than comfortably fits, updates regularly, or the answer needs a traceable source — a legal research tool over millions of pages of case law that's updated weekly has no context window large enough to hold it all, and "which case does this come from" matters enough that a citation isn't optional. It also fits any high-volume system, because paying to re-read a large block of text on every query adds up in a way a one-off question never will.

Using both

They aren't strictly either/or. A common real setup still retrieves — because most corpora are too large or too dynamic for any window to hold the whole thing — but retrieves a somewhat larger, less aggressively filtered set of passages than a small-window system would need, since there's now room to be a little less precise about the cutoff. The window got bigger; the reason to be selective about what fills it didn't go away, it just moved.

In this guide
  1. What RAG does
  2. What long context does
  3. Side by side
  4. Which one your problem calls for
  5. Using both
  6. FAQ

FAQ

If a model's context window is large enough to hold your entire knowledge base, is RAG definitely unnecessary?

Not automatically. Fitting is not the only cost — every query re-reading that entire knowledge base is slower and more expensive than retrieving a few relevant passages, and it's true on every single query, not just the first one. If you're answering more than a handful of questions against the same material, that cost compounds in a way a one-off use case never hits.