RAG Interview Questions (2026)
Covers retrieval-augmented generation and the pieces it's built from — embeddings, chunking, vector databases and reranking. See also all interview topics. These assume you already know the concept — if a section here is unfamiliar, read its linked concept page first; the questions test judgment on top of the concept, not the concept itself.
In this guide
RAG
Your RAG system retrieves the correct chunk, but the answer is still wrong. Where do you look first?
At the model's use of the context, not the retriever. If the right chunk made it into the prompt and the answer is still wrong, that's a faithfulness failure — the model drifted from the retrieved text instead of using it — which is a different problem from retrieval failing to find the chunk in the first place, and needs a different fix (prompt changes, a stricter instruction to only answer from the given context) rather than touching the retriever at all.
Your knowledge base is 50 short FAQ entries today, but the team expects it to reach 5,000 within a year. Do you build the RAG pipeline now?
Not for today's size. Fifty entries still fit directly in a prompt, so build for that — not for a roadmap number that might not land as predicted. Moving from "everything in the prompt" to retrieval later is a known, straightforward step once the content actually outgrows what fits. Building the pipeline now means carrying its cost and failure surface for a problem you don't have yet, based on chunking and embedding choices made before you've seen what the real content looks like at scale.
You're building RAG over a legal contract full of clauses that reference each other. What's the risk with your chunking strategy?
A clause split from the other clause it depends on can retrieve as if it stands alone, and the model answers as though it has the full picture when it doesn't. Contracts are a case where naive fixed-size chunking is genuinely risky — chunking needs to respect the document's actual structure (clause boundaries, section references) rather than just splitting every N tokens.
You add hybrid search to catch exact product SKUs that vector search was missing. What's the tradeoff?
Two retrieval systems to tune instead of one — you now need a way to combine or weigh keyword results against semantic ones, and getting that balance wrong can make results worse than either method alone. It fixes a real gap, but it's not a free upgrade.
If reranking improves result quality, why not just rerank every chunk in the whole knowledge base instead of doing a fast vector search first?
Rerankers are too slow and expensive to run over an entire large collection — that's exactly why the pipeline has two stages. Fast vector search narrows millions of chunks down to a shortlist cheaply; reranking then spends more compute per chunk, but only on that much smaller shortlist.
How would you notice a RAG index has gone stale in production, without waiting for a user to complain?
Track the gap between when source documents change and when the index reflects that change, and alert if it grows past an expected window. Sampling real queries against known-updated documents and checking whether the new content actually gets retrieved works too — waiting for user complaints means the failure was already live for a while before anyone knew.
Embeddings
You switch to a better embedding model and retrieval quality collapses. What happened?
Almost certainly the documents are still indexed with the old model while new queries use the new one. Vectors from two different models live in different coordinate spaces — comparing them doesn't produce worse results, it produces meaningless ones. Switching embedding models means re-embedding the entire corpus. That's why the choice is worth making before you index millions of chunks, not after.
A retrieved passage scores very high similarity but doesn't answer the question at all. Is the embedding model broken?
No — it's doing exactly what it was trained to. Embeddings capture what a passage is about, so "I love this product" and "I hate this product" sit close together: same topic, opposite meaning. High similarity means topically close, not "contains the answer." Any system that ranks purely by similarity will confidently return on-topic passages that answer nothing, which is the gap reranking exists to close.
Why not just embed each whole document instead of splitting it up?
Two reasons. Embedding models cap their input, so anything past the limit is cut off rather than summarised. And below the limit, the output vector is a fixed size whether you embed a sentence or forty pages — so a long document produces its general gist, and a specific fact buried inside it stops standing out against everything else in the same vector.
Would you pick a 3072-dimension model over a 1024-dimension one?
Not automatically. More dimensions mean more storage and slower search for every query you'll ever run, and the accuracy difference may not justify that. Several current models are trained so you can truncate the vector to a shorter length and lose very little quality, which makes the dimension count a tuning decision rather than a bigger-is-better one.
Chunking
How would you chunk a 200-page product manual?
On the structure the document already has — headings and sections — rather than by counting characters. A manual has real boundaries, and splitting on them keeps each chunk about one thing. Fixed-size splitting on a structured document is how you end up with a specification table separated from the heading that says which product it describes. Blind character splitting is the fallback for unbroken walls of text, not the default.
Retrieval keeps missing an answer you know is in the corpus. How do you tell whether it's a chunking problem or a search problem?
Check whether the complete answer exists inside any single stored chunk. If it's split across two, no amount of search tuning will fix it — the thing you need was never one retrievable unit, and that's a chunking problem. If it does sit whole in one chunk that simply isn't ranking, then it's retrieval or ranking. Reading a dozen of your actual stored chunks is worth more here than any metric; broken cut points are obvious on sight and invisible in aggregate numbers.
What is chunk overlap actually for, and when doesn't it help?
It repeats the tail of one chunk at the start of the next so that a fact straddling a boundary appears whole in at least one of them, rather than being cut in half in both. It's insurance against an unlucky boundary. It doesn't rescue bad cut points — if you're splitting mid-table or mid-clause, overlap just means two chunks are each partly wrong instead of one.
Is there an optimal chunk size?
No single one — and that's a real finding, not a dodge. Around 512 tokens with a little overlap is a reasonable default across retrieval tooling. But the best size depends on the question: narrow factual lookups do better on small, focused chunks; questions that need several pieces assembled do better on larger ones. Both kinds hit the same index, so no single value is optimal for all of them. That's why tuning chunk size usually pays less than fixing where you cut and what metadata you keep.
How would you get precise retrieval and enough surrounding context at the same time?
Stop assuming the chunk you match on and the chunk you return have to be the same text. Index a small, precise chunk so matching is sharp, then when it hits, return the surrounding section for the model to read. That dissolves most of the size tension instead of trading precision against sufficiency — and it depends on having stored enough metadata to know where each chunk came from.
Vector Database
When would you not use a dedicated vector database?
When you already run Postgres and the collection is in the thousands to low millions of chunks. pgvector adds vector search to the database you already back up and monitor, and keeps your filters, joins and transactions in one place instead of split across two stores that can disagree. The honest triggers for a dedicated system are scale past what one server handles comfortably, search traffic you want isolated from your main database, or a deliberate choice to make search someone else's operational problem. None of those are about vector search being special.
What's the difference between exact and approximate nearest-neighbour search, and what do you give up?
Exact search compares the query against every stored vector, so it always finds the genuinely closest matches, at a cost that grows linearly. Approximate search builds an index that checks only a fraction of the collection, trading some recall — the share of true nearest neighbours it actually returns — for large gains in speed and memory. The consequence worth stating in an interview: a relevant chunk can be missed because of the index itself, not because the embedding was wrong.
You need results filtered to one customer's documents and ranked by meaning. What breaks?
Filtering after the search starves the result set. Ask for the closest 20, discard the 18 belonging to other customers, and you're left with 2. Filtering before the search avoids that, but it needs an index built to support it — without one, you're back to scanning everything. This is where the naive after-the-fact approach falls apart, and it's a genuine reason a database already good at the "where" clause can beat a specialized vector store.
You reach for FAISS to move fast on a prototype, instead of standing up a hosted vector database. What are you signing up for once this leaves prototype stage?
FAISS is a library, not a database — it gives you the index and the nearest-neighbour search inside your own process, with no persistence, no network service, and no access control built in. That's fine while you're the only thing reading and writing it. The moment another service needs to query it, the index needs to survive a restart, or more than one process needs write access, you're building the database part yourself instead of using one that already exists.
Reranking
The right passage isn't in your first-stage results at all. Will a reranker fix it?
No, and this is an architectural limit rather than a quality issue. A reranker reorders the shortlist it's handed; it never sees anything the first stage didn't return. So if the answer-bearing chunk is missing from your top 100 on a third of test questions, swapping rerankers cannot fix that third — ever. That's a retrieval problem living in your embeddings, chunking, keyword coverage or the query itself. It's also why the two stages have to be measured separately.
You added a reranker and end-to-end quality got worse. How is that possible?
Domain mismatch. A reranker trained mostly on general web-search relevance learns what a plausible web answer looks like — not necessarily what counts as evidence in your corpus. An error code buried in API documentation doesn't resemble a well-written web answer at all. A mismatched reranker can take a ranking that had the right passage first and push it to third, with something plausible and wrong above it. That's why the baseline to measure against is the retriever alone, not nothing.
Would you rerank the top 200 candidates instead of the top 100 to be safer?
Not without measuring — reranking more candidates can come out worse. A longer shortlist doubles the scoring work and gives the model more opportunities to promote something that merely looks right. Depth of candidates is a parameter to tune against your own data, not a dial where higher is safer.
Someone tells you reranking improves RAG accuracy by 67%. What's wrong with that claim?
It conflates two different measurements. Anthropic's contextual retrieval experiment did report a 67% relative reduction — but in the share of relevant documents missing from the top 20 retrieved chunks, falling from 5.7% to 1.9%, and for a stack that also included generated per-chunk context and keyword search. Reranking was the last step of several, and retrieval failure is not the same thing as answer accuracy. Ranking benchmarks and end-to-end answer quality get quoted interchangeably far more often than they should be.