The problem
Ask a language model about a pair of running shoes and it will give you a confident answer. It understands that “for rain” means waterproof, and that “like my last pair” is a question about memory. That part it does beautifully.
Ask it what the shoes cost right now, or whether a size 10 wide is in stock today, and the confidence stays the same while the answer quietly goes stale. A price is not knowledge. It is a reading, and it is only as good as the moment it was taken.
For a shopping assistant this is not a small flaw. A wrong adjective costs nothing. A wrong price costs a sale, a refund or a customer’s trust, and a store owner reasonably expects the assistant on their site to be at least as accurate as the product page it sits on.
The split
So we drew a line through Concierge. On one side is everything that needs judgment: understanding what a shopper means, deciding which products truly fit, and explaining the choice in plain words. That side belongs to the model.
On the other side is everything that needs to be true: prices, stock, what a shopper has allowed, what actually happened. That side belongs to ordinary, deterministic software. It doesn’t guess, and it doesn’t improvise.
The model is allowed to be clever. It is never allowed to be the record.
In practice the model never types a price from memory. It asks for products through tools, receives them with their current price and stock attached, and writes its answer from that. The product cards a shopper sees are drawn from the same checked data, not from the model’s sentence, so even a clumsy sentence can’t put a wrong number on a card.
One request, slowed down
Here is a single question, and who is responsible for each step.
Figure 1. Judgment and truth alternate. The shopper sees one smooth answer.
Two details matter here. The search runs only over the store’s own catalogue, so Concierge can’t recommend something the store doesn’t sell. And the live check happens on the products that are about to be shown, not on the whole catalogue, which keeps it fast enough to do on every answer.
A price has an age
Once prices come from software instead of the model, the next question is how fresh they are. This turned out to be subtler than it looks.
Different sources report different times. One tells you when it last checked a price. Another tells you when the price last changed, which can be weeks ago on a price that is still perfectly correct. A third reports when a search index was built. In a test fixture these look identical. In production they mean opposite things, and treating “last changed” as “last checked” makes a correct price look ancient.
We now stamp every fact with the moment we read it, and keep the source’s own timestamp alongside as a separate field. Freshness is always measured from the read.
Then there are two limits, not one:
- To commit to a price, such as adding to a cart at that price, the reading has to be very recent.
- To mention a price, an older reading is acceptable, as long as the message says how old it is: “$129 when we checked this morning”.
We learned this the hard way. With a single strict limit, a price watch on a catalogue that is refreshed every few hours can never fire. Not rarely: never. And it fails silently, with no error anywhere, just a watch that stays quiet forever. Relaxing the one limit would have fixed the watch and let a cart ride a stale price. Two limits fix both.
Unknown stays unknown
Keeping facts outside the model has a quieter benefit: an unknown can stay unknown. If a check fails, Concierge can say so, instead of filling the gap with something plausible.
That only works if the software itself is honest about failure, and this is where many of our bugs hid. “We found nothing” and “we couldn’t look” are opposite statements, but a careless function returns an empty list for both. One real example: a store’s stock service answers an unauthorised request with an empty result. Read naively, that says every product has been delisted. Read correctly, it says we lack a permission.
So the rule we build to is that every lookup returns one of three things: a result, an honest empty result, or a failure with a reason. The model is told which one it got, and the answer it writes has to match.
Permission comes first
The split matters most when Concierge does something instead of saying something. A shopper who asks it to watch a price is asking for ongoing work, and ongoing work needs clear limits. So every task carries three separate permissions:
Our first design was a ladder: notify, then observe, then act, each level including the ones below. It read nicely and it was wrong. On a ladder, “change my cart but don’t tell me” is simply a level, and a combination you can’t describe is a combination you can’t refuse. With three independent permissions, the rule is easy to state and to enforce: changing a cart requires permission to tell the shopper. Concierge can never change a cart silently.
Every action also writes a receipt: what was done, on whose permission, based on which reading. A shopper or a store can read it later, and so can we when something looks wrong.
Claims are explicit
The same principle applies to what the model says about its own choices. Early on, the interface decided which product was “recommended” by reading the model’s prose and matching product names. If the answer discussed a product, the card got a badge.
That conflated two different things. Focus is the product the answer is talking about. A recommendation is a claim that this product is the best choice. An answer can focus on a product to explain why it is not the right one. Now a recommendation only appears when the model states it as a separate, explicit field. Ambiguous prose publishes no claim at all, and nothing downstream is allowed to guess one.
What we learned
The hardest lesson was about where to fix things. When the model got something wrong, the tempting fix was to write another instruction. It rarely worked. What the model sees shapes its answer far more than what it is told. The fixes that lasted changed the inputs, not the wording. We wrote that story up separately in Why a better prompt rarely fixes an agent.
The second lesson was that this split is not a safety tax. It is what made the rest possible. Price watches, cart actions and follow-ups that happen days later all depend on facts with an age and actions with a permission. Without the line, none of them would be safe to build.
For engineers: what a turn looks like on the wire (illustrative)
// streamed to the widget as server-sent events event: step data: {"label":"search_catalog","topic":"running shoes"} event: step data: {"kind":"check","what":"price_and_stock","items":3,"asOf":"2026-09-18T14:02:11Z"} event: step data: {"kind":"action","what":"add_to_cart","grants":["deliver","mutate"],"receipt":"rcpt_…"} event: done data: {"text":"…","products":[…],"recommended_product_id":"…"}
The answer text streams separately from the product data. The cards are rendered from products, never parsed out of text.
The same line gets more important as shopping spreads out. Shoppers will increasingly arrive through AI assistants, carry their preferences between stores, and ask agents to act for them over days, not seconds. Every one of those needs the same promise: the thinking can be flexible; the facts can’t.
HawkShift Research · Updated . Numbers are from our own tests while building Concierge; small samples are marked as small.











