Evals: Testing a System That Never Repeats Itself
Every engineer who ships an LLM feature arrives at the same afternoon. A prompt gets a small edit — one clarifying sentence, a reordered instruction — and the outputs look better on the three examples anyone checked. It goes out. A week later something unrelated is worse, and nobody can say when it started, because there was never a moment where the system said this got worse.
That moment is the entire product. Everything called an “eval” is machinery for producing it.
The prompt is not the asset. The set of cases you refused to lose is the asset — the prompt is just the thing you are allowed to change because of it.
Why the usual assertion breaks
A unit test works because the same input produces the same output, so equality is a valid question. Generative output has neither property. Ask the same question twice and you get two different sentences, both correct. Assert on the string and you have written a test that fails for the wrong reason forever; delete the assertion and you have written no test at all.
The way out is to stop asserting on the output and start asserting on a property of it. “Contains the order number.” “Cites only documents that were retrieved.” “Refuses.” “Parses as JSON matching this schema.” “Does not name a competitor.” Each of those survives rewording, which is precisely what makes it testable.
The three graders, cheapest first
Programmatic. A function over the output: schema validation, a regex, a numeric tolerance, checking that every cited id appears in the retrieved set. Deterministic, free, instant. Most of what people reach for a model to judge is actually one of these in disguise — “is the JSON valid” is not a matter of opinion.
Human. A person reading outputs and labelling them. Slow and expensive, and irreplaceable at exactly one job: producing the ground truth that tells you whether your other graders are any good.
Model-graded. Another model, given the input, the output and a rubric, returning a judgement. This is the one that unlocks the properties that matter most — faithfulness, tone, whether an answer actually addressed the question — and it is the one that quietly goes wrong.
Keep model graders narrow. One rubric, one question, a small output — pass,
fail, and one line of reason. A grader asked to return a score from one to
ten will hand you 7 for everything, and the difference between 7 and 8 will not
survive a model upgrade.
The set is built from failures, not imagination
The instinct is to sit down and write cases. That produces a set that covers what you already understood, which is the part that was never going to break.
The set that finds things is grown. Every time something goes wrong in production — a bad answer someone reported, a hallucinated citation, a tool call with nonsense arguments — the input becomes a case and the correct behaviour becomes its expectation. That is the long return edge in the diagram, and it is the difference between a set that keeps earning its runtime and one that goes quiet after a month.
Start smaller than feels responsible. Twenty to fifty real cases that run on every commit beat five hundred that run when someone remembers. Two cohorts is usually the right shape: a core set that gates every merge, and a broad set that runs nightly and is allowed to be slow.
Grade per case, never on the average
An aggregate score is how a regression hides. Ninety-four percent this run, ninety-three last run — inside that single point of movement, one case that used to work now leaks a customer’s data, and three unrelated cases got marginally better.
Compare case by case against the baseline and report the diff as these
specific cases changed state. The useful CI output is not a percentage; it is
invoice_with_credit_note: pass → fail. That names the thing to look at, which
is the only job the eval has.
For the genuinely flaky cases — and there will be some — run them a handful of times and gate on the pass rate rather than the single result. A case that passes three times in five was never a binary test, and pretending otherwise just teaches everyone to re-run CI until it goes green.
What to fix when it goes red
Resist the reflex to edit the prompt. A red case is evidence, and the first question is which part of the system produced it: retrieval that returned the wrong documents, a tool that returned an error the model then narrated, a context window that truncated the relevant part, or the prompt itself. Only the last one is a prompt problem, and it is not usually the one.
This is also the point of running evals on the pieces separately. A retrieval eval that scores whether the right document came back at all will tell you in one run what a whole-system eval takes an afternoon to isolate.
What you have actually built
Not proof that the system is correct — evals do not offer that and a set that claimed to would be lying about coverage.
What you have is a ratchet. A behaviour, once someone cared about it enough to write it down as a case, cannot silently go away again. That is a smaller promise than “tested”, and it is the one that turns a prompt everyone is frightened to touch into a system somebody can actually change.
Quick answers
- What is an eval in AI engineering?
- An eval is a test for a system whose output is not deterministic. Instead of asserting an exact string, it runs a fixed set of inputs through the system and scores each output with a grader — a rule, a program, or another model — then compares the scores against a known-good baseline.
- How are evals different from unit tests?
- A unit test asserts equality and gives a binary answer. An eval scores a distribution: the same input can produce different wording every run, so the assertion has to be about a property of the output rather than its exact text. Evals are also graded per case rather than aggregated, because an average hides the one case that broke.
- How many test cases does an eval set need?
- Fewer than people expect to start — twenty to fifty real cases will find most regressions, and a set small enough to run on every commit is worth more than a large one that runs weekly. Grow it from production failures rather than trying to imagine coverage in advance.
- Should you use an LLM to grade LLM outputs?
- Yes, for properties a program cannot check, like whether an answer is faithful to its sources. But treat the grader as code under test: check it against human labels on a sample, keep its model and prompt pinned, and prefer a cheap programmatic check whenever one exists.
References
Related Discoveries
Lumi's weekly note
A short email when we publish something useful. No spam, unsubscribe anytime.