RAG Retrieval Evaluation and RepairOperated by Reality Contact, LLC

Specific answer

How to build a RAG evaluation set that can guide a repair

Known failures, important tasks, expected sources, and explicit pass conditions form a small retrieval evaluation set.

A useful retrieval evaluation set is a versioned collection of representative queries tied to expected evidence and observable pass conditions. It should contain ordinary cases, known failures, confusing neighbors, and deliberate no-answer cases rather than a pile of unreviewed prompts.

Select cases from real decisions and known failure modes

Begin with the tasks that matter to the buyer: questions that precede an action, recur in support, or already produced a wrong answer. Add nearby queries that use different wording, queries where two documents conflict, and queries for which the corpus contains no support. A compact set of twenty reviewed cases usually teaches more than hundreds of synthetic questions whose expected source nobody has confirmed.

For each case, store the raw query, approved query variants, expected document identifiers, the passage or facts a reviewer considers sufficient, forbidden unsupported claims, and a reason the case matters. Keep the source snapshot or version beside the case. Otherwise a later document change can look like a retrieval regression even when the indexed record itself changed.

Score stages separately and preserve disagreements

Use separate labels for indexing coverage, retrieval relevance, ranking, context assembly, citation match, and answer groundedness. A single overall score hides which component needs repair. Binary pass conditions work for source presence, while graded judgments can capture rank or passage usefulness. Document the judge, rubric, and model version when an automated evaluator contributes a label.

Review disagreements rather than averaging them away. If a subject-matter reviewer expects one policy page while the system retrieves a newer page, resolve source authority before tuning. The final set should be small enough to inspect, stable enough to rerun, and broad enough to expose regressions when a repair helps one family of questions at another's expense.

Where the service stops

Reality Contact, LLC evaluates and repairs the bounded retrieval path, but does not certify model accuracy, source truth, security, compliance, or correctness for queries outside the agreed cases. The buyer approves the source authority and acceptance cases, reviews the reported limitations, and decides whether and how to deploy the changed retrieval configuration. This is software evaluation and implementation; it does not replace legal, financial, medical, security, compliance, or professional advice. No result establishes corpus-wide correctness, source truth, or reliable behavior for queries outside the buyer-approved case set.

Sources: LangSmith evaluation concepts; Braintrust evaluation guide.

Free failed-answer trace

A person returns a scored trace for one failed case showing the query, retrieved passages, expected source, first observable failure point, and a testable acceptance condition. The trace arrives within two business days after a readable case is received.

Do not send private links or files through this form. If the service fits, a person will reply with a secure intake method and written deletion terms before you share private material.

Questions about this answer

how to build a RAG evaluation set?

A useful retrieval evaluation set is a versioned collection of representative queries tied to expected evidence and observable pass conditions. It should contain ordinary cases, known failures, confusing neighbors, and deliberate no-answer cases rather than a pile of unreviewed prompts.

What should I send for the free check?

Do not send private links or files through this form. If the service fits, a person will reply with a secure intake method and written deletion terms before you share private material.

What does Reality Contact, LLC do?

Reality Contact, LLC evaluates and repairs the bounded retrieval path, but does not certify model accuracy, source truth, security, compliance, or correctness for queries outside the agreed cases. The buyer approves the source authority and acceptance cases, reviews the reported limitations, and decides whether and how to deploy the changed retrieval configuration.

Operated by Reality Contact, LLC.

Results apply only to the buyer-approved acceptance cases and recorded configuration.

First-party pseudonymous attention analytics · Privacy and opt-out