Agents

A follow-up needs evidence

Teaching an agent that keeps watching after the shopper leaves to only speak when something really changed.

HawkShift Research··6 min read

In short

  • Some shopper requests outlive the conversation: “tell me if this drops under $100”, “let me know if something better comes out”.
  • An agent that wakes up later must be handed fresh evidence, not just the old question. Without it, the only honest answer is “I don’t know”, and the dishonest one is a made-up change.
  • “Found nothing new” and “couldn’t look” are opposite messages, and the system has to keep them apart.
  • How often an agent looks again is a cost decision, not just a timing one. A shopper who asked for weekly should not pay for every fifteen minutes.

Work that outlives the chat

Most of what Concierge does happens while the shopper is watching. Some of it can’t. “Tell me when this drops under $100.” “Let me know if a lighter version comes out.” “Keep an eye on this until I buy it.” Each is a small promise to keep working after the tab is closed.

Promises like these are where an assistant earns trust or loses it. A message that arrives at the right moment with a real reason is the best thing an assistant can do. A message that arrives with nothing to say is spam, and a message that invents a reason is worse.

No news is not a reason to message a shopper.

Two kinds of promise

We found it useful to separate follow-ups into two kinds, because they fail in different ways.

A watchA revisit
Example“Tell me if it drops under $100.”“Tell me if something better appears.”
Who decides it’s timeSoftware: a number crossed a line.The model: something changed enough to matter.
Costs per checkA price lookup.A search and a model call.
Typical failureFiring on a stale or repeated signal.Waking up with nothing new to judge.
In a watch, the condition is data. In a revisit, the model is the condition, so the model has to be given something to judge.

A watch fits the split we describe in The model judges. The infrastructure checks. Software compares a fresh price with the shopper’s number; the model only writes the message. A revisit is harder, because deciding whether a new product is “better” is a judgment, and judgment needs material.

A wake needs evidence

Our first version of a revisit woke the model up with the shopper’s original request and the products it had found at the time. That sounds reasonable until you try to answer as the model: here is what was true three weeks ago; is there something better now? There is no way to know. The honest reply is “I don’t know”, every time, forever. The other possible reply is an invented change, which is the one thing a follow-up must never contain.

So before a revisit reaches the model, the system does the looking:

  1. It re-runs the shopper’s own request against the store’s catalogue, as it stands now.
  2. It compares the result with a snapshot taken when the promise was made.
  3. It hands the model three concrete lists: what appeared, what is gone, and what moved in price or stock.

Now the model has a real decision to make: is anything that appeared actually better for this shopper, given what they said they wanted? If not, it stays quiet, and staying quiet is the correct result.

This changed how we evaluate new kinds of follow-up. Before adding one, we ask a single question: what fresh evidence will it carry when it wakes? If the answer is none, it isn’t a feature yet.

Nothing vs. couldn’t look

Two cases look identical in a careless implementation and mean opposite things.

  • “Nothing appeared” is only a real statement if we know what was there before. If a promise has no snapshot to compare against, the wake is marked as incomplete and the model is told so. It must not tell the shopper that nothing changed.
  • “Nothing found” must never stand in for “the search failed”. A lookup that errors is reported as an error, not as an empty list. One means the market is quiet; the other means we don’t know.

The same trap exists for watches. One store platform answers an unauthorised stock request with an empty result, not an error. Treated as data, that reads as “every product was delisted”, and every watch on that store would have become impossible to keep. Only an explicit “this product no longer exists” counts as gone.

The cadence is the budget

A watch is cheap to check. A revisit is not: each one re-runs a search and asks a model to judge the result. So how often it runs is a spending decision.

We found this out from a bug. After each check, revisits were scheduled again using the fastest interval any follow-up was allowed, fifteen minutes. A shopper who had asked Concierge to look again weekly was quietly committing us to 96 model calls a day on their behalf. Nothing looked wrong, because every one of those calls worked.

Now a single rule decides when any follow-up is next due, it respects what the shopper asked for, and a revisit never runs more than once a day. Watching a market also needs its own permission, separate from permission to message the shopper, because it spends the store’s resources for as long as the promise stands.

Say it once

A price that drops below a shopper’s number should produce one message. Our first implementation produced one every fifteen minutes.

The cause was subtle. Each alert had an identity, so the system could tell whether it had already been sent. That identity included a revision number that changed every time the system touched the watch, including every routine check. So when a shopper declined an alert, the next check produced a “new” alert about the same drop.

Now an alert’s identity is tied to the crossing itself. It only changes when the condition is seen to be false again, the price goes back up, or the shopper genuinely edits the watch. One drop, one message.

Old facts, said honestly

Prices on a catalogue are refreshed on a schedule. If a follow-up demands a price read within the last hour and the catalogue refreshes every few hours, the follow-up can never fire, and it fails silently: no error, just a watch that stays quiet forever.

The fix wasn’t to accept stale prices everywhere. It was to separate saying from doing. A message can rest on a reading that is a few hours old, as long as it says so: “$279 when we checked at 9am, down from $319 when you asked.” Anything that commits the shopper to a price, like changing their cart, still needs a fresh reading.

The rules we follow now

  1. Hand over evidence, not a questionA follow-up wakes with what changed since the promise was made, computed by software.
  2. Silence is a valid resultIf nothing that matters changed, nothing is sent.
  3. Keep “nothing” and “couldn’t look” apartA failed lookup is reported as a failure, and a missing snapshot as a gap.
  4. Price the scheduleEach check that calls a model has a cost; the interval is set by what the shopper asked for, with a daily floor.
  5. One crossing, one messageAlert identity follows the event, not the bookkeeping.
  6. Name the age of every factA message can use an older reading if it says when it was taken.

HawkShift Research · . Numbers are from our own tests while building Concierge; small samples are marked as small.