RAG retrieval debugging

RAG Retrieval Debugging With Query and Source Traces

Table of Contents

A wrong RAG answer can survive several successful API calls. The vector database returns results, the model responds, and the application displays citations, yet unsupported evidence prevents grounded generation.

I approach RAG retrieval debugging by finding the first broken stage, rather than immediately changing the prompt or model. You need a trace that connects the original query, eligible sources, retrieved candidates, final context, and generated claims.

Start by preserving that chain before changing anything.

Key Takeaways

  • Record candidate lists and final context separately; retrieved evidence can disappear during reranking or prompt assembly.
  • Apply authorization and version rules consistently across dense and sparse retrieval.
  • Fix missing evidence before adding a reranker, then verify each answer claim against its cited passage.

Capture a Trace Before Changing Anything

A final-answer log is insufficient. I recommend using one request identifier across query processing, retrieval, reranking, context assembly, and generation.

A query card leads to three document cards and a small stack of results.

Record the Query and Every Candidate List

Preserve the original question alongside any rewritten query. Query expansion can drop an error code, change a product name, or erase a date constraint.

Record the embedding model, query vector, and index revision. Log effective metadata filtering rules, candidate count, chunk IDs, scores, ranks, and stage durations. For hybrid search, retain dense and sparse results before fusion.

OpenTelemetry’s GenAI conventions provide a reference for telemetry naming. Your application still needs retrieval-specific fields that explain why a particular chunk qualified.

Preserve the Source and Final Context

Every chunk should resolve to a document ID, immutable revision, content hash, and visible location, such as a page or section. Save the exact context supplied to the model, including ordering and truncation.

A document URL alone cannot reproduce an incident after the document changes.

For agent-driven retrieval, AI agent audit logging extends this record to tool arguments, permissions, retries, and returned state.

Protect traces with access controls, redaction, and retention limits. Raw queries and source passages can contain customer data.

Locate the First Broken Stage

I would first classify whether the incident is a retrieval failure or a generation failure. The RAG research survey examines retrieval and generation challenges separately; production debugging needs that distinction too.

Follow the expected supporting passage through the pipeline.

Trace ObservationLikely Failure AreaNext Check
Supporting passage is absent from the indexIngestion or source coverageInspect parsing and indexing manifests.
Passage exists but isn’t eligibleMetadata or authorizationValidate scope and version rules.
Eligible passage never enters candidatesFirst-stage retrievalCompare lexical and dense results.
Passage enters candidates but disappearsRanking or context assemblyInspect rank changes and truncation.
Complete evidence reaches the model, but the answer contradicts itGeneration or groundingEvaluate claims against supplied context.

These are starting points, not automatic verdicts. Several stages can fail in the same request.

Replay the failed question with verified supporting passages while holding the prompt and model configuration fixed. If the answer improves, retrieval or context selection deserves attention. If it remains wrong, evaluate grounded generation against the supplied evidence.

Unauthorized evidence must remain unavailable. An accurate answer isn’t a valid reason to weaken access boundaries.

Check Parsing, Chunk Boundaries, and Versions

Before tuning similarity scores, inspect the source text stored in the index. Clean-looking PDFs can produce broken reading order, detached table headers, or missing footnotes.

Keep Conditions With the Rule

Heading-aware chunking usually provides a more useful baseline than splitting every document at an arbitrary token boundary.

Keep a policy rule with its exceptions. Preserve table headings with relevant rows. Attach section context when a paragraph depends on a preceding definition.

I wouldn’t prescribe one chunk size across FAQs, API references, and contracts. Choose a document-specific chunking strategy instead. Compare configurations against labeled questions and inspect whether retrieved text chunks contain complete supporting evidence. Overlap can reduce boundary losses, but duplicate passages can also crowd the candidate list.

Validate Revisions Across Both Indexes

Dense and sparse indexes for a knowledge base must reference the same eligible source revisions. Otherwise, hybrid retrieval can combine current lexical matches with obsolete vector results.

Check document IDs, revision IDs, lifecycle status, effective dates, and indexing manifests. Normalize filter values and reject chunks with missing authorization metadata.

RAG document versioning strategies help separate historical evidence from current answers. Historical queries need explicit validity rules, not whichever near-identical passage ranks highest.

Investigate filter effects using an authorized test corpus. Don’t disable tenant boundaries against live production data.

Tune Hybrid Search and Reranking Separately

Hybrid search broadens candidate coverage. A reranker changes candidate ordering. Those jobs need separate measurements.

Fuse Rankings Without Mixing Raw Scores

Dense retrieval supports semantic search by matching a query vector with passages about paraphrases and related concepts. Sparse retrieval with BM25 protects literal terms such as PostgreSQL’s 23505 error code, API parameters, and version strings.

Elasticsearch and OpenSearch are common choices for BM25-based retrieval. When paths are independent, run them concurrently and apply identical permission and validity constraints.

As an experimental baseline, you could retrieve 20 candidates per path, deduplicate by chunk ID, and combine rankings with Reciprocal Rank Fusion (RRF). This is a starting configuration, not a universal setting.

RRF combines rank positions rather than directly mixing BM25 and cosine similarity scores. I favor that baseline over guessing raw-score weights before evaluating real queries.

Add a Reranker Only When Evidence Is Present

A cross-encoder evaluates the query and candidate passage together. It can distinguish a directly relevant answer from text that merely discusses the same subject.

Milvus documents WeightedRanker and RRFRanker for combining search results. Those fusion mechanisms combine rankings, unlike a model that reads query-passage pairs.

Compare ranks before and after the cross-encoder. If it demotes the supporting passage, investigate missing surrounding context, text truncation, or domain mismatch.

A reranker cannot recover evidence absent from its candidate set. Measure candidate recall before paying for a second ranking stage.

Reranking adds compute, latency, and possibly API costs. Keep the reranker only when measured quality gains justify that overhead.

Use Vector-Space Views Carefully

A projected view of the vector space can help expose duplicate clusters, isolated document groups, or queries landing near unrelated material.

A display shows point clusters, scattered outliers, and one separated candidate group.

When reviewing vector embeddings in Milvus or another vector database, compare groups by source family, language, revision, and ingestion batch. Unexpected separation can suggest parsing differences or inconsistent embedding configurations.

I treat the visualization as a diagnostic aid, not a relevance verdict. Two-dimensional and three-dimensional projections distort the original geometry. Apparent proximity doesn’t establish citation support.

Cosine similarity also measures representation-level closeness, not factual applicability, authorization, or freshness. Inspect actual passages and full-dimensional retrieval results before changing the embedding model.

A visually convincing cluster can still contain the wrong policy version.

Trace Answer Claims to Exact Source Passages

Citation presence and citation support are different checks. An answer can link to a legitimate document that never makes the statement.

Extract factual claims and check each against its cited passage. Include dates, quantities, permissions, conditions, and exceptions in that review.

Maintain a controlled chain connecting each claim to its chunk ID, source revision, and visible location. The model may select source identifiers, but the application should validate and resolve them. Don’t allow invented URLs or unknown IDs.

Document-level citations are too coarse when an exception elsewhere changes the rule. Users should be able to inspect the passage that supports each statement.

Distinguish unsupported claims from contradicted claims. Missing evidence may require abstention or another retrieval step. A contradiction demands investigation.

Reasoning-capable models still need these checks. Grounded generation doesn’t validate an incorrect premise. Fluent explanations aren’t proof of factual support or protection from LLM hallucinations.

Replay Failures as Regression Tests

A trace becomes more useful when it becomes a repeatable test. Add confirmed production failures alongside representative questions, exact identifiers, ambiguous wording, historical requests, and unanswerable queries.

For each case, record the authorized supporting passages and expected answer constraints, including checks for unsupported claims. Keep index revisions and retrieval settings identifiable so comparisons remain meaningful.

I would track evaluation metrics separately: candidate recall, supporting-passage rank, final-context coverage, answer correctness, citation support, and permission violations.

Recall at k measures how much labeled relevant evidence appears among the first k results. Mean reciprocal rank tracks the position of the first relevant result. Neither establishes answer faithfulness.

LangSmith’s RAG evaluation tutorial separates retrieval relevance from answer evaluation. That prevents a fluent response from masking weak retrieval.

The broader RAG evaluation metrics and testing guide connects these checks to a repeatable testing process.

Change one variable at a time. Track p50 and p95 latency, candidate counts, context tokens, and cost alongside quality. Improving an aggregate score while breaking tenant isolation or doubling tail latency isn’t an acceptable release.

Frequently Asked Questions

How Can You Separate Retrieval Failures From Hallucinations?

Inspect the exact context supplied to the model. Missing, outdated, or incomplete evidence points upstream. Use citation support to assess the evidence behind an answer: unsupported claims may indicate missing evidence, while claims that contradict complete context point toward generation or grounding. Both failures can occur together, so evaluate retrieved passages and generated claims separately.

Should Every RAG Application Use Hybrid Search?

No. Add hybrid search when evaluation exposes exact-term misses or mixed lexical and semantic needs. Dense-only retrieval can be sufficient for some corpora. Hybrid search adds indexing, synchronization, and diagnostic work, so use it when measurements justify the added complexity.

What Should You Fix Before Adding a Reranker?

Validate source coverage, extraction, chunk boundaries, authorization, revision eligibility, and candidate recall. A reranker can help when supporting evidence enters the candidate pool but ranks below less useful passages. It can’t repair missing documents or invalid permissions.

Make the Evidence Chain Reproducible

RAG retrieval debugging becomes practical when every answer has a reproducible evidence chain. Preserve the query, eligible source revisions, candidate lists, final context, and claim-level support.

I would fix the earliest verified failure, then replay the incident before adding another component. Successful API calls aren’t enough; the trace must show that the right evidence reached the answer.

RAG Retrieval Debugging With Query and Source Traces mailbox@3x

Oh hi there!
It’s nice to meet you.

Sign up to receive awesome content in your inbox, every month.

We don’t spam! Read our privacy policy for more info.

You might also like

Picture of Evan A

Evan A

Evan is the founder of AI Flow Review, a website that delivers honest, hands-on reviews of AI tools. He specializes in SEO, affiliate marketing, and web development, helping readers make informed tech decisions.

Your AI advantage starts here

Join thousands of smart readers getting weekly AI reviews, tips, and strategies — free, no spam.