Scheduled social posts in Cleo are published by two cooperating pieces. A scheduler runs every minute, finds posts that are due, and claims each one by writing a timestamp into its row. It then starts a durable background run for the post, passing that same timestamp along as a claim token. When the run starts, it reads the row and checks that the claim is still the one it was given. If the row now carries a different claim, another process has taken over, and this run exits quietly as stale.
It is a sound design for preventing a post from going out twice. Under one condition I had not tested for, it could also stop a post from going out at all.
The same instant, different text
The scheduler wrote the claim as an ISO timestamp from JavaScript, which ends in Z: something like 09:30:00.374Z. It sent that exact string as the token.
The database stored it as a timestamp with time zone, which is correct, and when the run read it back, Postgres rendered it in its own format: 09:30:00.374+00:00.
Those are the same instant. They are not the same string. The run compared them as strings, found them different, concluded the claim had been taken by someone else, and exited as stale. Every run, for every post. A recovery job that retries stuck posts saw them remain unpublished, tried a few times, and gave up.
Why the tests passed
The tests for this path used made-up tokens: tidy strings, identical on both sides because the test wrote both sides. They exercised the logic of the comparison perfectly and the reality of the comparison not at all. The bug only exists where a value makes a round trip through a system with its own opinion about formatting.
The new tests carry the real pair, a Z timestamp and its +00:00 rendering, copied from production. They fail against the old comparison.
The fixes
The claim check now parses both values and compares instants, not text. The scheduler also sends the token in the form the database stored it, so even a naive comparison would agree. Either change alone would have fixed it. Both together mean the next change to either side cannot quietly bring it back.
The give-up path got attention too, because its silence turned a small bug into a long one. When the recovery job abandons a post now, it tells the workspace through the same notice and owner email every other settled failure uses, and it logs at error level rather than as a warning. A post that will not go out is now something the owner hears about the same day.
Two smaller things fell out of it. Scheduling a post on a channel that is not connected now says so at the moment of scheduling, rather than accepting it as scheduled and failing later. And the failure notice reads "An Instagram post", not "A Instagram post". A small thing, but that notice is the one sentence the owner reads about the problem.
The rule
Never compare timestamps as strings. They have too many valid spellings: Z or +00:00, with or without fractional seconds, three digits or six, a T or a space. Any value that passes through a database, a queue or a serialisation boundary can come back spelled differently while meaning exactly the same thing. Parse, then compare. And when you test code that compares values across a boundary, use values that have actually crossed it.