When a retrieval-augmented system gives a weak answer, the model may be the last component that deserves blame. The source might never have entered the context at all.
Map the RAG failure chain
A typical RAG flow ingests documents, divides them into chunks, adds metadata, indexes representations, retrieves candidates and asks a model to answer from them. A failure at each stage looks different: stale ingestion, destructive chunking, poor ranking, context truncation or unsupported generation.
Keep a trace that connects the query to document version, retrieved chunks, scores, filters, final context and response. Without that chain, a plausible answer can hide a broken retriever and a wrong answer can trigger pointless prompt edits.
Measure retrieval with labeled questions
Build questions whose supporting documents are known. Measure whether the required source appears in the top results, then use ranking measures when order matters. Microsoft’s RAG evaluator guidance separates document-retrieval metrics from text-quality judgments of retrieved context; that distinction is useful because relevance labels enable more precise diagnosis.
Test filters and permissions as first-class behavior. A retriever that finds a relevant but unauthorized document has failed even if the answer is correct. Include duplicated pages, outdated revisions, tables and questions requiring evidence from more than one chunk.
Judge the answer on independent axes
Groundedness asks whether claims stay within the supplied context. Relevance asks whether the answer addresses the question. Completeness asks whether it covers the critical expected information. These dimensions can move separately: a cautious answer may be grounded but incomplete, while a polished answer may be relevant yet unsupported.
For high-impact use, attach citations to specific claims and verify that each citation entails the nearby statement. A list of sources at the bottom is not enough if readers cannot tell which passage supports which conclusion.
Stress the boundaries
Ask questions the corpus cannot answer and require an explicit limitation rather than an improvised response. Add prompt injection embedded in retrieved documents, conflicting versions, misspellings, paraphrases and time-sensitive questions. Evaluate whether the system identifies authority and date instead of treating every chunk as equal.
Run separate language slices when your users query in one language and the corpus uses another. Cross-language retrieval can fail before generation, and a fluent response may conceal the miss.
Turn diagnosis into changes
If recall is low, inspect ingestion, chunk boundaries, metadata, embeddings and query rewriting. If retrieved evidence is good but groundedness is poor, change the generation contract and citation requirements. If answers are grounded but incomplete, revisit context assembly and multi-document coverage.
Compare changes against a frozen set and a fresh holdout set. This prevents optimization for familiar questions while preserving a path for genuine generalization.
How to put the idea into practice
Begin with one bounded workflow related to rag evaluation: test retrieval before blaming the model. Write a one-page baseline before changing the system: current completion time, quality checks, common failure categories, escalation path and the person responsible for the outcome. Select a representative sample rather than only the cleanest examples. Include ordinary cases, difficult edge cases and at least one case where the correct result is to stop or ask for more information.
Run the candidate beside the existing process before allowing it to replace that process. Review both successful and failed outputs, because a lower error rate can still conceal a new high-impact failure. Record the exact configuration used for each test and keep artifacts that let another reviewer reproduce the result. At the end of the pilot, decide whether to expand, revise or stop using thresholds agreed in advance—not a retrospective impression of the best demonstration.
Questions to ask a vendor or internal team
Ask what evidence supports the central claim, which system and data versions produced that evidence, and what conditions were excluded. Request results for the languages, input types and risk categories your deployment will encounter. Ask how changes are announced, how regressions are detected and how a customer can export logs needed for an incident review.
Also ask what the system does when confidence is low, a dependency fails or the request falls outside its supported scope. A dependable product should have a defined failure state, not merely a more polished answer. Ownership matters: identify who can pause the workflow, who approves exceptions and who informs affected users if a material error escapes into production.
A practical checklist
- Define the user task and the failure that matters before choosing a model or tool.
- Keep a small, versioned test set drawn from real work, including awkward and adversarial cases.
- Record model, prompt, tools, retrieval settings, data version, latency and cost for every run.
- Require human confirmation for irreversible, high-impact or externally visible actions.
- Review failures by category, not just by one average score, and add regressions to the test set.
Related reading from Meydo Journal
Primary sources
- Retrieval-Augmented Generation evaluators — Microsoft Learn
