How to Evaluate Agentic RAG Systems.

By
Adithya Kamath
September 29, 2026
•
8 min read

01

‍Blog content test

Blog content test

How to evaluate and improve an agentic RAG system

A question-answering system can sound smart long before it is trustworthy. It searches your internal documents and uses what it finds to write a fluent answer. The first good result feels like magic; repeated confident errors turn it into a governance problem.

The usual response is to build a benchmark: collect questions with known-good answers, run the system, score the outputs, and watch the number move over time. That benchmark works for ordinary retrieval-augmented generation. An agentic system requires evidence about the run itself.

An agentic RAG system does not follow one fixed path. Depending on the question, it may choose a structured query, free-text retrieval, business filters, reranking, context selection, and a final synthesis step. Two runs of the same question can take different routes, and both can be valid if they satisfy the same required constraints. A final answer can be right for the wrong reason, or look credible while using the wrong evidence.

An acceptable answer must also come from an acceptable run. Did the system take a valid path, use the right evidence, honor the required constraints, and behave consistently enough that we can trust it tomorrow?

That is a higher bar. 

A run is not an answer. It is an artifact.

The key shift is to treat every evaluation run as an evidence trail.

For each question, record the route constraints the system was expected to satisfy, the route it actually took, the queries and filters it issued, the sources it selected, and what happened to those sources before the answer was written. Capture the final answer too, but do not let it stand alone.

A weak answer may begin with retrieval. Other failure modes include:

  • the agent chose free-text search for a question that needed a structured count;
  • a required date or campaign filter was never applied;
  • the system found the right document, then dropped it while reranking or packing context;
  • the model had the evidence but failed to use or cite it; or
  • the infrastructure was unavailable and the system safely declined to answer.

Those are different problems with different owners. A single "answer quality" score turns them all into one undifferentiated bad result. A good harness goes further: it clusters every failure into a named cause, such as wrong route, filter gap, unproven denominator, dropped source, over-abstention, or run blocker, and attaches a prescribed next action to each. The output is never "question 14 failed." It is "question 14 failed as a structured filter gap, along with three siblings, and here is what to inspect."

The practical rule underneath all of it: unknown is not a pass. If a required trace is missing, the system has not proved the behavior, so mark it for review. If token usage is unavailable, report it as unavailable, not as zero. That discipline is not pedantry. It is how an evaluation system avoids lying to the people who use it.

The first question should test the ground beneath the system

There is a subtle way evaluation goes wrong before the first real question is asked.

When a credential expires or a database connection breaks, a well-behaved system may not throw an obvious error. It fails closed: "I could not find enough information to answer safely." That is the correct response when the evidence truly is not there, and also the response when retrieval is broken. If you only inspect final answers, a dead run looks like a cautious product.

The answer is a preflight canary. Before the suite runs, send one known-answerable synthetic question through the exact path the real questions use. If it cannot retrieve and summarize a single indexed document, stop the run and stamp it with a named blocker, such as expired credentials or a failed tool connection, before a single product question is scored. Blocked stays blocked all the way through the pipeline: it can never be averaged into a quality score or classified as a product regression.

This small control separates "the product failed" from "the ground the product stands on failed." Without it, infrastructure failures quietly contaminate the benchmark and create false stories about regressions.

Structured questions need structured proof

Some questions are not semantic questions. They are inventory questions.

"How many confidential information memoranda are in this campaign?" cannot be validated by a convincing paragraph containing a number. The system must show that it found the correct campaign, searched the right source population, applied the right document-type and date filters, and knows whether its count is complete.

That needs a contract. For structured questions, the evaluation specifies, in machine-checkable form, the required constraints and allowed route class, the source denominator, the required filters and metadata, the coverage expectations, and the fallback policy: if the structured route fails, may the system fall back to generic search, or must it refuse rather than guess? Require an exact sequence only when order is a business or safety requirement, including a compliance control.

The distinction matters most when a system hedges its way into a bad answer:

"This may be incomplete, but here are 12 companies."

 This answer shape is dangerous because it admits the denominator is unproven, then launders the uncertainty through a list. The harness detects the combination of hedging language and an enumerated result, then treats it as a hard failure that is worse than a clean abstention.

Refusal, meanwhile, earns real credit. In our scoring, a clean fail-closed abstention is worth more than half of a full pass because honest refusal is genuinely useful behavior, but never enough to compete with a proven answer. That weighting is a harness choice, not a universal rule. The verdict ladder has room for both truths.

Watch the evidence funnel

One of the most useful diagnostics is also the easiest to miss.

A system can retrieve the right document and still produce a bad answer, because the document disappears later in the pipeline:

retrieved → filtered → reranked → final context → cited

The harness tracks stable source identifiers and counts at every stage and computes explicit drop records between them. It also records why a source was excluded. If the expected source appears at retrieval but vanishes before the final context, the search index was not the problem. The reranker or context selection was, and the harness says so by name. If the source survives into the final context but never appears in the answer, the failure moved downstream into synthesis or citation.

Two refinements keep this check honest. Loosely written expectations ("the report, or the data-room folder if present") are filtered out so only hard expected sources can trigger an alarm. And routes whose correct answer is "nothing matches" are exempt from the empty-context failure. A system should never be punished for correctly proving an absence.

This is what turns improvement work from "how do we make this answer better?" into "where did the evidence stop being available, and what is the smallest change that fixes that mechanism?"

Score behavior

Once the evidence trail exists, scoring can reflect the actual standard. In our weighting, contract and trace behavior control roughly seven of every ten points; required answer content controls only three. Those weights are a local choice, not a portable standard. The incentive is deliberate: a candidate cannot win by writing smoother sentences. Take the wrong route or drop the expected source and fluent prose does not help; it loses points.

Comparison against a baseline is deliberately asymmetric. An improvement requires a meaningful score or verdict gain. A regression is any score drop, verdict drop, trace-health drop, or any new trace failure, even if the total score went up. A baseline row missing from the candidate run counts as the worst regression of all. But one regressed row should veto the entire candidate only when it creates a defined critical failure, such as an authorization breach or required-source loss. Other regressions should be reviewed against severity, run-to-run variance, and workload-specific thresholds. Nine better answers do not make one newly broken critical contract acceptable.

A consistency gate repeats the candidate and compares its route constraints with the evidence and facts used in each answer. Any variation that breaks the contract fails the gate, while different valid execution paths remain acceptable. Three runs is our current starting point, not a universal requirement; the trial count should scale with the system's nondeterminism and the risk of the workload. A separate generalization check flags candidates whose gains come only from answer markers, with no improvement in contract or trace behavior. That pattern usually indicates benchmark fitting rather than a better system.

AutoRAG: automate the experiments instead of judgement

This harness eventually becomes the foundation for an improvement loop we call AutoRAG, though its boundaries matter more than the name.

The loop takes a diagnosed failure, proposes a small configuration change with a stated hypothesis, such as "inventory questions are falling through to generic retrieval because the structured route is not enabled," never "make question 7 pass," and tests it. Each candidate runs in a dedicated environment with its settings injected. The environment is discarded after the test, so there is never any doubt the change actually took effect.

It never tests the failing question alone. Neighbors are selected by measured similarity, including shared query type, shared required route, the same source section, or shared tags, because a change that repairs one question and breaks a sibling is not a fix. If no neighbors qualify, the loop refuses to declare generalization on a sample of one. Promotion tests also include a protected holdout and regression set that the optimizer did not use to generate the candidate.

Two defenses target self-deception directly. A patch review scans added lines for question identifiers or other verbatim benchmark text. Finding a match is a warning signal, not proof of contamination, but it forces a human look because a real fix addresses a class of failures rather than naming the test. An optional LLM judge compares baseline and candidate answers for semantic direction, with instructions to penalize fabrication or degraded grounding. The judge is advisory by construction. It can keep a candidate under review but cannot rescue one that deterministic evidence has condemned. Otherwise, fluency wins arguments against facts.

What the loop can never do is promote itself. A winning candidate becomes a promotion candidate, with a report. Humans still own the meaning: whether the contracts reflect real business intent, whether the cost and latency trade is acceptable, whether the behavior generalizes, and whether the change ships. Promotion also needs a rollback path and post-release monitoring. The machine does the repetitive investigation. People retain the right to decide what counts as evidence and what risk is acceptable.

The standard is evidence

The point of evaluation is not to reward an answer that sounds right. It is to prove that the system reached it through an acceptable route, used the required evidence, and satisfied the contract for the question.

That changes how improvement is judged. A higher score is meaningful only when the harness can show what changed, where the failure was fixed, and whether the gain survives repeated and neighboring tests without creating a critical regression. Otherwise, the result is just a better-looking output.

Better answers matter. But an agentic system becomes trustworthy only when you can inspect how it got there, reproduce the behavior, and trust it to refuse when the evidence is missing. Evaluation should make that standard enforceable.