When an email campaign finishes in Cleo, it gets a final status. For a long time that status was nearly always "sent", and the reason was structural rather than optimistic.
The dispatcher sends to recipients in chunks through the email provider. When a recipient was refused, the error was pushed into a local array and logged as a warning. The warning left a breadcrumb in the error tracker and nothing countable anywhere else. At the end, the finalising step counted only successful sends and wrote "sent" regardless. A campaign where the provider refused every single recipient would show as sent, with a zero beside it.
Failures as data
The first change was to stop discarding failures. Each refused recipient now writes a durable event row with a sanitised reason and a flag for whether the failure is worth retrying. The row never holds the email body or the address, only what is needed to count and classify.
Finalising now reconciles those failure rows against the success rows and picks the state the send actually earned: sent, partly failed, or failed. A recipient who failed once and succeeded on a retry counts as sent. A campaign with nobody left to send to is sent, not failed, because having nothing to do is different from trying and having nothing work.
Which failures deserve a second attempt
Only transient provider failures are retried: rate limits, timeouts, server errors. A refused address is recorded and left alone. Anything unclassified defaults to no retry, because a retried email is not free for the person receiving it.
Getting that classification right had a trap in it. The first version treated any message containing a five-hundred-series code as transient, since an HTTP 5xx means a server problem. But SMTP also uses 5xx codes, and in SMTP they mean permanent refusal. "550 5.1.1 User unknown" is a dead address, and the loose match would have sent to it again. Permanent signatures are now checked first and win. The numeric match is gone, replaced by explicit HTTP status codes that cannot collide with SMTP's 55x family. Rate limiting outranks both, because however a provider words it, it is always transient.
Reserve, then settle
Billing needed the same honesty. Approving a send used to check whether the workspace could afford it, reserve nothing, and charge at the end. Approval now reserves the energy up front. When the send finishes, the reservation is settled against what was actually delivered. A failed campaign costs nothing. A partial one costs only what reached people. The refund is tied to the original debit by a stable key, so even if two finalising calls race, the database refuses the second credit.
While working the edges of that change, before it went for review, I found one more defect, and it was mine. A campaign's statistics were stored in a single JSON field, recomputed and written wholesale every time a delivery event arrived from the provider's webhook. I had just added the reservation to that same field. A "delivered" event arriving mid-send would recompute the counters, overwrite the field and erase the reservation. Finalising would then find no reservation, fall back to the old charge-at-the-end path, and bill the send a second time.
The statistics write now merges: recomputed counters win, and everything else in the field is preserved. To be sure the regression tests meant something, I put the defect back and watched all three fail.
One migration, applied first
A last practical note. The new final states needed a database constraint change before the code could write them. Had the code gone out first, writing "partly failed" against the old constraint would have been refused and left the campaign stuck in "sending" for good, which is the very class of problem this work was meant to end. The migration was tested in a rolled-back transaction against the real schema and applied ahead of the deploy.
A status is a claim the product makes to the person using it. "Sent" is a strong claim. It should only appear when the evidence supports it, and the bill should follow the evidence too.