Skip to content
Discovery AI Engineering 5 min read · Updated 5 Aug 2026

Evals: Testing a System That Never Repeats Itself

intermediate ai-agentstestingci

Every engineer who ships an LLM feature arrives at the same afternoon. A prompt gets a small edit — one clarifying sentence, a reordered instruction — and the outputs look better on the three examples anyone checked. It goes out. A week later something unrelated is worse, and nobody can say when it started, because there was never a moment where the system said this got worse.

That moment is the entire product. Everything called an “eval” is machinery for producing it.

The evaluation loop around a change to an LLM system A change to a prompt, model or tool runs against a kept regression set of cases and their expectations. Graders score each case individually, using exact, programmatic or model-graded checks. The scores are compared against the last known-good baseline. If no case regressed the change ships. If any case regressed the merge is blocked, naming the specific case that broke, and the change returns for a fix. Separately, every failure observed in production is written back into the regression set as a new case, so the set grows from real behaviour rather than from cases imagined up front. Change prompt or model Regression set cases you kept Graders per case Compare vs last green 1 2 3 exact · programmatic · model-graded Ship it nothing regressed Block the merge name the case that broke fix, then run it again Production surprise the one nobody imagined every failure becomes a case eval run --set core --gate --baseline main

The prompt is not the asset. The set of cases you refused to lose is the asset — the prompt is just the thing you are allowed to change because of it.

The eval loop. The forward path is cheap: a change runs against the kept set, graders score each case, and the scores are compared with the last known-good baseline. The two return edges are what make it work — a regression blocks the merge and names the case, and every production surprise is written back into the set. The set grows from real behaviour rather than from cases someone imagined up front.

Why the usual assertion breaks

A unit test works because the same input produces the same output, so equality is a valid question. Generative output has neither property. Ask the same question twice and you get two different sentences, both correct. Assert on the string and you have written a test that fails for the wrong reason forever; delete the assertion and you have written no test at all.

The way out is to stop asserting on the output and start asserting on a property of it. “Contains the order number.” “Cites only documents that were retrieved.” “Refuses.” “Parses as JSON matching this schema.” “Does not name a competitor.” Each of those survives rewording, which is precisely what makes it testable.

The three graders, cheapest first

Programmatic. A function over the output: schema validation, a regex, a numeric tolerance, checking that every cited id appears in the retrieved set. Deterministic, free, instant. Most of what people reach for a model to judge is actually one of these in disguise — “is the JSON valid” is not a matter of opinion.

Human. A person reading outputs and labelling them. Slow and expensive, and irreplaceable at exactly one job: producing the ground truth that tells you whether your other graders are any good.

Model-graded. Another model, given the input, the output and a rubric, returning a judgement. This is the one that unlocks the properties that matter most — faithfulness, tone, whether an answer actually addressed the question — and it is the one that quietly goes wrong.

Keep model graders narrow. One rubric, one question, a small output — pass, fail, and one line of reason. A grader asked to return a score from one to ten will hand you 7 for everything, and the difference between 7 and 8 will not survive a model upgrade.

The set is built from failures, not imagination

The instinct is to sit down and write cases. That produces a set that covers what you already understood, which is the part that was never going to break.

The set that finds things is grown. Every time something goes wrong in production — a bad answer someone reported, a hallucinated citation, a tool call with nonsense arguments — the input becomes a case and the correct behaviour becomes its expectation. That is the long return edge in the diagram, and it is the difference between a set that keeps earning its runtime and one that goes quiet after a month.

Start smaller than feels responsible. Twenty to fifty real cases that run on every commit beat five hundred that run when someone remembers. Two cohorts is usually the right shape: a core set that gates every merge, and a broad set that runs nightly and is allowed to be slow.

Grade per case, never on the average

An aggregate score is how a regression hides. Ninety-four percent this run, ninety-three last run — inside that single point of movement, one case that used to work now leaks a customer’s data, and three unrelated cases got marginally better.

Compare case by case against the baseline and report the diff as these specific cases changed state. The useful CI output is not a percentage; it is invoice_with_credit_note: pass → fail. That names the thing to look at, which is the only job the eval has.

For the genuinely flaky cases — and there will be some — run them a handful of times and gate on the pass rate rather than the single result. A case that passes three times in five was never a binary test, and pretending otherwise just teaches everyone to re-run CI until it goes green.

What to fix when it goes red

Resist the reflex to edit the prompt. A red case is evidence, and the first question is which part of the system produced it: retrieval that returned the wrong documents, a tool that returned an error the model then narrated, a context window that truncated the relevant part, or the prompt itself. Only the last one is a prompt problem, and it is not usually the one.

This is also the point of running evals on the pieces separately. A retrieval eval that scores whether the right document came back at all will tell you in one run what a whole-system eval takes an afternoon to isolate.

What you have actually built

Not proof that the system is correct — evals do not offer that and a set that claimed to would be lying about coverage.

What you have is a ratchet. A behaviour, once someone cared about it enough to write it down as a case, cannot silently go away again. That is a smaller promise than “tested”, and it is the one that turns a prompt everyone is frightened to touch into a system somebody can actually change.

Quick answers

What is an eval in AI engineering?
An eval is a test for a system whose output is not deterministic. Instead of asserting an exact string, it runs a fixed set of inputs through the system and scores each output with a grader — a rule, a program, or another model — then compares the scores against a known-good baseline.
How are evals different from unit tests?
A unit test asserts equality and gives a binary answer. An eval scores a distribution: the same input can produce different wording every run, so the assertion has to be about a property of the output rather than its exact text. Evals are also graded per case rather than aggregated, because an average hides the one case that broke.
How many test cases does an eval set need?
Fewer than people expect to start — twenty to fifty real cases will find most regressions, and a set small enough to run on every commit is worth more than a large one that runs weekly. Grow it from production failures rather than trying to imagine coverage in advance.
Should you use an LLM to grade LLM outputs?
Yes, for properties a program cannot check, like whether an answer is faithful to its sources. But treat the grader as code under test: check it against human labels on a sample, keep its model and prompt pinned, and prefer a cheap programmatic check whenever one exists.

References

Related Discoveries