Skip to content
Back to the workshop

The Cache That Never Hit

A prefix cache rewards stable prefixes and charges for everything else

E
EugeneBuilding Cleo
4 min read

Model providers offer prompt caching for long requests. The mechanism is a prefix cache. You mark a point in the request, and if the next request begins with exactly the same content up to that point, the provider reads it from cache at a fraction of the normal price instead of processing it again. Writing to the cache costs more than an ordinary request and reading from it costs much less. At the five-minute cache lifetime, a write costs about twelve and a half times as much as a read.

Cleo places several of these markers. One sits after the standing instructions, which change rarely. Another sits at the end of the conversation history, so that on a follow-up turn the conversation so far can be read from cache and only the newest message is new.

I measured that second marker across two weeks of real traffic, grouping follow-up turns by the gap since the previous turn. It had never produced a hit. Not once. Turns two minutes apart rewrote the same history as turns forty minutes apart.

Two reasons a prefix changes

The first reason was placement. Between the stable instructions and the conversation sat a block of context that changes on every turn: the current time, material retrieved for this particular message, and a few facts about where the user is and what state their account is in. All of it useful. All of it different every time. And because the cache matches a prefix, anything that changes before the conversation invalidates everything after it. The conversation could never be read from cache, because the bytes in front of it were never the same twice.

The second reason was the window. Only the most recent part of a long conversation is sent with each request, and that window slid forward by one message every turn. So even with stable content in front of it, the first message of the history changed on every turn, and the prefix broke there instead.

Moving the part that changes

The per-turn context now travels as the opening text of the final user message, after the conversation marker rather than before it. It is attached to each request and never saved into the conversation, so it does not pile up in the history. If a user types text that imitates the wrapper around that context, it is stripped, so the boundary cannot be impersonated from inside a message.

The window now moves in steps of ten messages instead of one. For nine turns out of ten, the start of the history is identical to the previous turn and the prefix holds. On the tenth it moves and the cache is written once. There was a subtle race here: a message still being saved could shift the count and flip the window's starting point between two requests. The client now sends its untrimmed count, so the server anchors on the same number either way.

Choosing the lifetime

The third change was the cache lifetime. The conversation marker used the five-minute cache, while the instructions already used the one-hour cache. The traffic showed about a quarter of follow-up turns arriving between five minutes and an hour after the previous one, which is a person reading, thinking, doing something else, and coming back. Those turns were paying for a full write. The conversation marker now uses the one-hour lifetime as well.

A longer lifetime costs more to write, so I set the condition for undoing it before shipping. If, within two days, follow-up turns were not showing cache reads, or the per-turn context was not landing in its new place on nearly every request, the lifetime would go back to five minutes and the other two changes would stay. Writing down the revert condition in advance is the only way I know to keep an optimisation honest.

What to take from it

The cache was never broken. It did exactly what its documentation says: match the prefix, byte for byte. What was wrong was my picture of the prefix, which included things I thought of as context and the cache saw as content. If you use prompt caching, read the actual request in order, mark every part that can change between turns, and make sure all of it sits after the last marker you want to hit. Then measure hits by the gap between turns rather than in aggregate, because an overall cache rate can look healthy while the cache you care about has never worked.

E

Written by Eugene

Building Cleo, an AI marketing operating system. These posts cover the architecture decisions, technical challenges, and lessons learned along the way.

More from the workshop