Skip to content
Back to the workshop

The Row That Never Finished

A concurrency gate with no age limit fails slowly and silently

E
EugeneBuilding Cleo
4 min read

When a tool in Cleo starts a piece of work, it writes an activity row with the status running. When the work finishes, the row is updated to completed or failed. Those rows drive several things: progress indicators, the history of what has been made, and a concurrency gate on image work.

The gate is simple. Making an image is slow and expensive, so each workspace may have at most two image jobs running at once. Before starting a new one, the code counts that workspace's running rows. If there are already two, it waits for a slot, for up to sixty seconds, and then proceeds anyway rather than blocking the user indefinitely.

Every part of that is defensible. Put together, it hid a problem for months.

Rows that outlive their process

A row is only marked completed if the process that wrote it lives long enough to update it. Serverless functions get terminated: a timeout, a deploy, a crash in a dependency. When that happens mid-job, the row stays running forever. Nothing reconciles it.

When I went looking, production held twenty-one of these rows. The oldest was from March.

Now look at the gate again. It counted running rows without regard to their age. Two orphans in one workspace and the gate considered that workspace permanently full. Every image job after that waited the full sixty seconds for a slot that would never free up, and then went ahead.

Because the gate proceeded rather than throwing, nothing ever reached the error tracker. The images were made. They were simply a minute slower than they should have been, every time, for as long as the orphans existed. The workspaces that had actually crossed the threshold turned out to be my own internal ones, so no customer was yet paying that minute. But the trap was set for anyone who accumulated two.

Three fixes that overlap on purpose

I fixed it three ways, deliberately redundant.

The gate now ignores running rows older than fifteen minutes. No image job takes that long, so an older running row is an orphan by definition. This makes the gate correct on its own, without depending on anything else having run.

A nightly job moves stale running rows to a final state, so the table tells the truth for everything else that reads it.

A migration cleared the twenty-one that already existed.

All three change the status. None of them delete. Completed rows are the history people see of what they have made, and other tables refer to them. Deleting stale rows would have been the obvious tidy-up, and it would have quietly broken references elsewhere.

The comment that described a clean-up that did not exist

There was one more thing. The migration that created this table, months earlier, carried a comment saying stale rows were removed after twenty-four hours by a scheduled job. No such job existed. It never had.

That comment did real harm, in the way documentation can. Anyone reading the table definition would believe the problem was handled. And had anyone implemented the comment as written, deleting every row older than a day, almost the entire table would have gone, including the history people rely on. The accurate description now lives in a comment on the table itself, in the database, where the next person to read the schema will find it.

What I check for now

Two things. Any gate that counts work in flight needs an age limit, because in flight is a state processes can enter and fail to leave. And any fallback that degrades instead of failing needs a measurement of its own, because a system that waits a minute and then carries on is a system nobody will ever report as broken. It is merely slow, consistently, and consistent slowness rarely gets investigated.

E

Written by Eugene

Building Cleo, an AI marketing operating system. These posts cover the architecture decisions, technical challenges, and lessons learned along the way.

More from the workshop