Why the first answer
A shopper who opens a chat on a store is deciding, in the first second or two, whether it is worth their time. Later answers are forgiven more easily: the conversation is going, and there are product cards to look at. The first answer has no such cover.
So we took apart the first answer of a brand-new conversation, the slowest case, to see where the time goes. Everything below is measured on a connected test store, running the production code, with the timings the system records for each step rather than a stopwatch around the whole thing.
The anatomy of a first answer
| Milestone | Time | What has happened by then |
|---|---|---|
| Request admitted | ~155 ms after sending | The request has arrived, been checked and been accepted. |
| Answer stream open | ~167 ms after sending | The connection that will carry the answer back to the widget is ready. |
| Model called | 250–400 ms after sending | Everything before the model is done: the store’s settings, the shopper’s allowed memory, the tools for this turn. |
| First words from the model | a further 490–1,140 ms | The model has read the question and started answering. This varies from request to request. |
Two things stand out. The work we control, everything before the model, is a few hundred milliseconds. And the model step is not only the largest part but also noisy: the same question can take twice as long from one moment to the next.
What can’t be warmed
The most common suggestion we heard was to “warm up the model” so it would be ready when the shopper asked. It sounds reasonable, and it misunderstands what the model is doing during those 500 to 1,100 milliseconds. It isn’t loading. It is reading this shopper’s question and working out an answer, and it can’t start that until the question exists.
The second suggestion was to warm the prompt cache. Language model providers can reuse the processing of a prompt’s unchanging opening, which saves time and money. We checked: on the first turn of a brand-new conversation, 94% or more of the prompt was already served from cache. The instructions and tool descriptions are the same for everyone, so they are already warm. There was nothing left to buy.
The only levers on the model step itself are how much it is asked to think and which model is used. Those are quality decisions as much as speed decisions, and we treat them that way.
What can: the handshake
Before any data flows between our servers and a model provider, the two sides open a secure connection: a network round trip or two to agree on encryption. Normally this happens when the shopper presses send. It doesn’t depend on the question at all.
The widget already told our servers when a shopper started typing, so it could prepare a conversation. We attached one more job to that signal: open the connection to the model provider now, so it is ready and waiting when the question arrives.
A few details made this safe to ship:
- It spends nothing. The warm-up asks the provider for a trivial listing, never for an answer, so no tokens are used.
- It happens at most once every ten seconds, however fast someone types.
- It fails silently. If warming fails, the real request simply opens its own connection, as before.
The trap in the test
There is an easy way to build this warm-up that works perfectly in every log and does nothing at all.
A warm-up that opens its connection from its own private pool will connect, receive a response and report success. But the real request draws from a different pool, so the warm connection sits unused while the real request does its own handshake. Every metric that asks “did the warm-up succeed?” says yes.
So the warm-up has to share the real request’s pool, and it has to fully read its own response, or the connection is never returned for reuse. The lesson is about the measurement: the right question is never “did the warm-up run?” but “did the real request reuse a connection?” That is the second row of Figure 2, and it is the number we check.
Photos
Shoppers can send Concierge a photo: “something like this, but cheaper”. Looking at an image adds a separate step, where a vision model describes what it sees before the main agent decides what to do.
That step takes about 4.6 seconds on average for a typical product photo, and tall screenshots take longer, because the time depends mostly on how much description is written, not on the image size. We run other preparation at the same time, but that work is fast, so the overlap only hides between 0.1 and 0.5 seconds. The shopper still waits about four seconds.
Shortening the description is the only real lever, and it trades away exactly the detail that makes photo search useful. So for now we keep the detail and are honest about the wait: the widget shows a “Looking at your photo…” step while it happens.
Feeling fast
A few design choices don’t change the measured time but change how it feels:
- Show real progress. Concierge streams short progress steps, such as “Searched the catalogue”, as they happen. They are true descriptions of work done, not a spinner with words on it.
- Stream the answer. Words appear as they are written, so the first sentence can be read while the rest is composed.
- Withdraw a rejected draft cleanly. If a streamed draft fails a check and is rewritten, the rejected text is cleared, not left on screen with the new answer appended below it. We learned this one from a bug.
What we learned
- Measure per step, on the wireTotal time tells you something changed; per-step timings tell you what.
- Check the cache before optimising itOurs was already above 94% on a first turn. Warming it would have been wasted work.
- Separate what needs the question from what doesn’tAnything that doesn’t depend on the shopper’s words can start while they are still typing.
- Measure the effect, not the attempt“The warm-up succeeded” can be true and useless; “the real request reused a connection” is the real signal.
- Interleave comparisonsNetwork timings drift; run A and B alternately after the same idle period.
The biggest part of the first second still belongs to the model, and that is as it should be: it is the part doing the thinking. Our job is to make sure nothing else is in its way.
HawkShift Research · . Numbers are from our own tests while building Concierge; small samples are marked as small.











