Research · HawkShift

Notes from building an honest shopping agent.

What we measure while building Concierge: what worked, what didn’t, and the questions we haven’t answered yet.

  1. 1Real tests, real numbersEvery claim comes from a measurement we ran.
  2. 2Failures includedWhat broke is often the most useful part.
  3. 3Plain words firstReadable by a store owner, with the details for engineers.
ArchitectureFeatured

The model judges. The infrastructure checks.

Why Concierge never lets a language model be the source of truth for a price, a stock level or a cart, and what that split made possible.

HawkShift Research··7 min read
Read the paper
All research

Recent work.

26page setups tested on an iPhone to find what Safari puts behind the clock
Web platform

What iOS 26 Safari puts behind the clock

Why most websites get a solid strip behind the iPhone status bar and Dynamic Island in iOS 26, tested one change at a time, and how we designed around it.

·9 min read
1shopper is enough for demand to show, never their words
Memory & privacy

Insights without surveillance

How a store learns what shoppers want without ever seeing one shopper’s conversation.

·9 min read
93%of turns unanswered on a build our old suite scored as unchanged
Evaluation

A green test suite is not a good answer

Why we stopped trusting unit scores for an agent, and started grading real shopper outcomes instead.

·7 min read
No newsis not a reason to message a shopper
Agents

A follow-up needs evidence

Teaching an agent that keeps watching after the shopper leaves to only speak when something really changed.

·6 min read
Inputsbeat instructions, again and again
Grounding

Why a better prompt rarely fixes an agent

When the data an agent sees contradicts its rules, rewriting the rules is wasted work. Fix what it sees.

·6 min read
−52%wait for the model’s first response, by connecting while the shopper types
Speed

Where the first second goes

Taking apart a first answer, millisecond by millisecond, and what we could and couldn’t make faster.

·6 min read
Yoursshoppers see, correct and remove what Concierge remembers
Memory & privacy

Memory that belongs to the shopper

Designing memory a shopper can read, correct and take back, and why consent has to be part of the design.

·6 min read
Open questions · what we’re working on

Shopping is changing. These are the hard parts.

01

Can preferences travel without being exposed?

A shopper’s sizes and tastes could follow them between stores. How do we let that happen only with their consent, and without stores seeing each other?

02

What should a store say to someone else’s AI?

More shoppers will arrive through AI assistants. How does a store give those assistants the same checked truth it gives a person?

03

Can one agent shop the whole web honestly?

Helping a shopper across many stores means being clear about where every fact came from, and whose store it is.

04

When should an agent act on its own?

Watching a price is easy. Deciding when to message, or change a cart, is about permission, evidence and trust.

These are research directions, not products.

Follow along

New research, when it’s ready.

No schedule and no filler. Just an email when we publish something worth reading.