Grounding

Why a better prompt rarely fixes an agent

When the data an agent sees contradicts its rules, rewriting the rules is wasted work. Fix what it sees.

HawkShift Research··6 min read

In short

  • When an agent’s data and its instructions disagree, the data wins. Rewording the instruction is wasted effort.
  • Several of our worst behaviours were correct inferences from incomplete inputs. Adding one true fact fixed them with no new rules.
  • Notes written for one situation, stated unconditionally, create the behaviour they describe. A note that assumes its own effect prevents it.
  • Where a sentence sits matters as much as what it says. And a rule should forbid a false claim, never a line of reasoning.

The reflex

When an AI agent does something wrong, the natural fix is to tell it not to. Add a sentence to the instructions, test it once, move on. We did this for months while building Concierge, and we kept a record of which fixes survived.

Most didn’t. The ones that did had something in common: they changed what the agent could see, not what it was told. This article collects the clearest cases, with the numbers. Samples are small, and we say so where they are; we measured each by changing one thing at a time and re-running the same request.

The catalogue that wasn’t

A shopper asked for laptops with 32GB of memory. Concierge found eight and said: “These eight are the only 32GB laptops in the catalogue.” The store had more.

The obvious fix was a rule: search results are not the whole catalogue; don’t claim they are. It changed nothing. The model kept making the claim.

Then we looked at what the model was actually given. Every catalogue search returns a page of at most eight products. But nothing in the model’s input said it was a page. It received an unqualified list of eight laptops. From its point of view, “these are the only eight” wasn’t a hallucination. It was the correct conclusion from the data it had.

We added one true field to the search result, saying whether more matches existed. No new rule, no rewording. Faced with a result that said “more available”, the agent searched again on its own, 4 times out of 4.

A rule can bound what the data leaves open. It can’t overrule what the data says.

The cap that chose the answer

A related bug was hidden in a setting that looked like housekeeping. To keep the model’s input small, the list of products handed to it at the answering step was trimmed to the top two.

Two things followed, both measured. First, every search was silently rewritten into a two-product answer: the context carried eight products, the step that wrote the reply saw two, and the reply mentioned two. Second, and less obviously, two products is exactly the shape of a comparison. Ordinary “find me” requests kept drifting into “here’s how these two compare” framing, because a pair of products looks like a comparison request.

Raising the cap to a full page of eight fixed both. The model handled the longer list cleanly and named all eight. A size limit on an input is also a decision about the output.

Notes that presuppose

Instructions often arrive as short notes attached at a particular step. Two of ours were written for a specific situation, but phrased as if they were always true, and both created the problem they were near.

“The shopper sees the product cards below your reply”

This note was meant to stop the agent from reading out every name and price the cards already showed. But the cards only appear if the agent’s reply chooses to present them. The model read the note as “display is handled”, wrote a nice paragraph, and never presented the cards. It even repeated the false premise back: “the cards you see already cover office chairs.”

We rewrote that one sentence as the causal fact: the shopper sees these results only if your reply presents them. Nothing else changed.

“Show me monitors under $500”: cards shown0/43/4
Follow-up “now show me office chairs”: cards shown3/55/5
Figure 1. One sentence rewritten. Small samples (4 and 5 live requests per arm).

“Keep your recommendation even if a later step fails”

This was a sensible rule for one situation: don’t drop your advice because, say, adding to the cart failed. But it was delivered on every turn, and it presupposes that a recommendation exists. So the agent made one. A plain “show me …” request came back with unrequested verdicts and sweeps through every product’s specifications.

We replayed a captured request and removed one block of instructions at a time. Removing that note took unrequested verdicts from 4 out of 4 to 0 out of 4. Removing any of the other blocks left it at 4 out of 4. Adding prose elsewhere asking for “proportionate” answers did nothing.

Where a sentence sits

Position turned out to matter as much as wording.

Concierge can remember a shopper’s preferences, if they allow it, and should check them before searching. With the fact that memory was available sitting in the context the agent read first, it consulted memory on 0 of 6 searches where it should have. One sentence placed directly after the shopper’s message, asking it to check: 6 of 6, with no change in the typical number of model calls. The model acts on what is next to the decision it’s making.

The reverse also happened. We added a line to an always-visible list of optional procedures, noting that loading one was cheap. The agent read cheapness as encouragement: a plain “hi” went from one model call to two, every time. Moving the same sentence into the description of the tool that loads procedures restored one call, 3 out of 3. Permission stated where a choice is made reads as a nudge. Stated in the tool’s own description, it reads as mechanics.

Bound claims, not thoughts

The last pattern is about what a rule forbids. Concierge had the rule “you cannot derive tax”, written to stop it from inventing a tax rate. A shopper then gave it their local rate and asked for a total. It refused, three times in a row. Deleting that one sentence restored the right answer.

The rule was forbidding an operation, calculating tax, when the actual danger was a false claim, a made-up rate. “Don’t invent a tax rate” is correct. “You can’t calculate tax” teaches the agent to be useless even when the shopper hands it everything it needs.

When we audited the instructions, 32 of 33 prohibitions bounded a claim. Exactly one bounded an operation, and it was the bug. We now scan for that shape automatically: a new rule that bans a way of reasoning fails the check.

Before you write a rule

  1. Read the input firstBefore blaming the model, look at exactly what it was given for that decision. Check whether its mistake is the right conclusion from those inputs.
  2. Add the missing fact, not a ruleIf a list is a page, say so. If a count is partial, say so. One true field beats a paragraph of policy.
  3. Treat every cap as a decisionTrimming an input for size also chooses the content and even the shape of the answer.
  4. Never presupposeA note must not assume a state that only the agent’s own output creates, or assume work that may not exist.
  5. Put rules next to the choiceGuidance works when it sits beside the decision it governs, not in a preamble.
  6. Forbid false claims, not reasoningAsk of every rule: does this ban a statement, or a thought? Only the first is safe.
  7. Measure one change at a timeReplay the same request, remove or change one block, and alternate the variants.

None of this makes instructions unimportant. It puts them in their place: they can shape how an agent uses what it knows. They can’t change what it knows. For that, fix the input.

HawkShift Research · . Numbers are from our own tests while building Concierge; small samples are marked as small.