RAG Retrieval Evaluation and RepairOperated by Reality Contact, LLC

Specific answer

Retrieval relevance and groundedness measure different RAG failures

Retrieval relevance and groundedness require different records and identify different RAG failures.

Retrieval relevance asks whether the returned passages help answer the query. Groundedness asks whether the answer's claims are supported by the context it received. A repair needs both checks because either stage can pass while the other fails.

Evaluate the evidence before evaluating the prose

A relevant passage addresses the user's question and comes from an accepted source, even if the final answer misstates it. Compare each returned passage with the query and expected source, then record its usefulness and rank. Also inspect whether the context window preserved the passage after reranking, truncation, or prompt assembly. This separates an index or ranking problem from an answer-generation problem.

A grounded answer can still be unhelpful when it faithfully summarizes an irrelevant passage. A relevant retrieval can still produce an ungrounded answer when the model adds a fact that does not appear in context. Keep a claim-to-passage table for the answer and a query-to-passage table for retrieval. The two tables make the failure location visible without relying on an overall impression of quality.

Use labels that support an engineering decision

Define relevance against the buyer's source authority, not whichever passage looks plausible. Define groundedness at the claim level and mark unsupported, contradicted, and partially supported statements separately. When an automated evaluator is used, store its prompt, model, and output beside a human-reviewed sample so changes in the evaluator do not masquerade as product improvement.

The acceptance table should show the retrieval result, cited evidence, answer result, and limitation for each case. That makes the next action concrete: repair ingestion, adjust metadata, tune ranking, change context assembly, constrain the answer, or hold the case because the corpus lacks an authoritative source.

Where the service stops

Reality Contact, LLC evaluates and repairs the bounded retrieval path, but does not certify model accuracy, source truth, security, compliance, or correctness for queries outside the agreed cases. The buyer approves the source authority and acceptance cases, reviews the reported limitations, and decides whether and how to deploy the changed retrieval configuration. This is software evaluation and implementation; it does not replace legal, financial, medical, security, compliance, or professional advice. No result establishes corpus-wide correctness, source truth, or reliable behavior for queries outside the buyer-approved case set.

Sources: LangSmith RAG evaluation tutorial; LangSmith evaluation concepts.

Free failed-answer trace

A person returns a scored trace for one failed case showing the query, retrieved passages, expected source, first observable failure point, and a testable acceptance condition. The trace arrives within two business days after a readable case is received.

Do not send private links or files through this form. If the service fits, a person will reply with a secure intake method and written deletion terms before you share private material.

Questions about this answer

RAG retrieval relevance vs groundedness?

Retrieval relevance asks whether the returned passages help answer the query. Groundedness asks whether the answer's claims are supported by the context it received. A repair needs both checks because either stage can pass while the other fails.

What should I send for the free check?

Do not send private links or files through this form. If the service fits, a person will reply with a secure intake method and written deletion terms before you share private material.

What does Reality Contact, LLC do?

Reality Contact, LLC evaluates and repairs the bounded retrieval path, but does not certify model accuracy, source truth, security, compliance, or correctness for queries outside the agreed cases. The buyer approves the source authority and acceptance cases, reviews the reported limitations, and decides whether and how to deploy the changed retrieval configuration.

Operated by Reality Contact, LLC.

Results apply only to the buyer-approved acceptance cases and recorded configuration.

First-party pseudonymous attention analytics · Privacy and opt-out