Evaluation

A green test suite is not a good answer

Why we stopped trusting unit scores for an agent, and started grading real shopper outcomes instead.

HawkShift Research··7 min read

In short

  • A scripted test suite proves an agent doesn’t crash. It can’t prove a shopper got what they asked for, because it never asks anything its author didn’t expect.
  • Many of our serious bugs lived in the seams between parts: the database, the real API, the route a request actually took. Unit tests don’t cross seams.
  • Replaying old conversations is a weaker check than it sounds: the same agent, run twice on the same input, disagrees with itself.
  • We now grade outcomes with a simulated shopper and a judge, run real probes against real services, and always keep a control.
93%of turns unanswered on a build our scripted suite scored as unchanged
4real defects live behind a fully green 41-case suite
3defects found by the first real API call, past 47 passing tests
0 / 71failed model outputs rescued by simply asking again

All from our own testing of Concierge in 2026.

The scoreboard lied

For a while, our main quality check for Concierge was a suite of scripted conversations with pass-or-fail rules: no apologising for things that weren’t wrong, no robotic phrasing, no stock phrases that make an assistant sound like a form. It was fast, cheap and reassuring.

Then we built a second check that worked differently. A model played an ordinary shopper, wrote each message in reaction to what Concierge had actually just said, and a separate judge scored each exchange on one question: did the shopper get what they asked for?

On the same build, the scripted suite scored 51 passes and 3 failures, identical to the previous version. The simulated shopper found that 93% of turns didn’t answer the question. Both numbers were accurate. Only one described the product.

A smaller example made the same point more sharply. A change was reported as passing 15 out of 15 checks. In the same period, the share of turns that returned products when products were requested fell from 33% to 26%. The check counted “we don’t stock dehumidifiers” as a pass, because it contained no forbidden phrase. The shopper who asked for a dehumidifier got nothing.

Why suites pass

A scripted suite is a list of situations its author already thought of. That makes it good at one job, catching the return of a known failure, and structurally blind to everything else. An agent that talks to the public spends most of its time in situations nobody wrote down.

Rules phrased as “must not contain” make this worse. They reward silence. The cheapest way to never say the wrong thing is to say very little, and an agent optimised against a list of bad phrases drifts toward answers that are safe, polite and useless.

The bugs live in the seams

The second lesson came from ordinary engineering, not from the model. Three times we shipped, or nearly shipped, a green build that was broken in production.

What passedWhat was broken
40 checks on a routing changeA word in a pattern misread stated budgets; a gate never reached the new code; a stale result set leaked into comparisons.
41 cases for price watchesThe real database dropped a field the in-memory one kept; a declined alert re-fired every 15 minutes; a temporary product reference was stored as permanent; two callers passed identity in different shapes.
47 tests for an order-history readerThe real API wrapped its reply in an envelope; it rejected a query that asked for the same field twice; its real “access denied” error matched none of the phrases we looked for.
Each test suite was correct about what it tested. None of them crossed the boundary where the defect lived.

The pattern is consistent. A unit test proves a function does the right arithmetic. It can’t prove the request ever reaches that function, or that a real database returns the shape a fake one does. A fixture records what you already believe about an API; only the live API can reject your query.

The watch bug is worth a closer look. The alert’s identity included a revision number that changed every time the system looked at it. So a shopper who dismissed one price drop would be offered the same drop again a quarter of an hour later, with a fresh identity each time. In a test with one check, the revision never moved. The fix was to tie an alert’s identity to the moment the condition actually became true, not to the bookkeeping around it.

Replay can’t replay

The obvious next idea is to record real conversations and replay them against every new build. We built that, and learned to be careful with it.

Language models are not fully deterministic, and neither is an agent that calls them in a loop. When we re-ran 585 recorded turns twice on the same build, 35 disagreed with themselves. And 21% of turns no longer matched what had originally been logged, not because anything broke, but because the system had moved on since.

So a replay can’t be compared with the old logs. It has to be compared with a fresh baseline of the current build, with the self-disagreeing turns set aside and reported, never quietly dropped. A replay percentage quoted without its excluded count is not a number we trust.

Replay also taught us what a good signal looks like. In one change, every final result matched and every unit test passed. The only evidence of a bug was that the agent now made a different number of model calls on five turns. A check moved to the wrong place had sent those turns down a different path that happened to land on the same answer. A changed call count is what a wrong branch looks like when the final label is unchanged.

Grading outcomes

The questions that matter are about the shopper’s result, not the agent’s phrasing:

  • Did they get what they asked for? Products when they wanted products, a comparison when they asked to compare, a clear answer when they asked a question.
  • Was every fact grounded? Each price and claim must trace back to something the system actually looked up.
  • Did the conversation hold? A follow-up like “the cheaper one” has to resolve to the product the shopper meant.

The simulated shopper matters because it reacts. A scripted suite sends message two no matter what message one produced. A simulated shopper who gets a bad answer asks again, gets frustrated or changes course, exactly the paths where real conversations go wrong.

It isn’t free: a hundred-turn run takes about twenty minutes and real model calls. We run it before calling anything done, not on every save.

Measuring honestly

A few habits made the difference between numbers we could act on and numbers that flattered us.

Measure the raw output, not the cleaned-up one

Our system tidies small mistakes in model output before using it. That is good engineering, and it hid a problem: 34 of 34 outputs were “valid” after tidying, but only 16 needed no tidying at all. Reporting only the first number would have hidden a model getting worse behind a cleaner getting better. We report both.

Check that a retry can succeed before adding one

When a model output fails validation, asking again is the reflex. We counted: 71 failed outputs were retried and none recovered. Sampling the same input four times gave the same wrong value every time. The error was a property of the input, not bad luck, and the right fix was to handle it directly. Some failures really are random: one kind succeeded about nine times in ten on byte-identical input, so a single retry recovers it. The only way to know which kind you have is to sample.

Interleave A and B

Two fixes for a random failure looked perfect in sequence, 0 failures in 20 each, and did nothing when the variants were interleaved. The service underneath drifts over minutes, and running all of A and then all of B measures the drift. We now alternate them.

What we do now

  1. Grade what the shopper gotScore outcomes, never the absence of bad phrases. A quiet agent passes every “must not say” rule.
  2. Cross the seamBefore calling a change done, run it once against the real database, the real API and the real route.
  3. Baseline, don’t compare to logsReplay against the current build run twice, and publish how many turns disagreed with themselves.
  4. Watch the call countA turn that made a different number of model calls took a different path, even if the answer looks the same.
  5. Sample before retryingIf the same input fails the same way every time, a retry only doubles the cost.
  6. Keep a controlAlternate variants. Any before-and-after measured in sequence is partly a measurement of the clock.

None of this makes testing easier. It makes the tests about the thing we actually care about, which is a person asking for something and getting it.

HawkShift Research · . Numbers are from our own tests while building Concierge; small samples are marked as small.