IOSOR Learn

Old webhooks must drain before you cut keys

Drain in-flight DLR on the old endpoint before you revoke keys. Cut only after quiet, then re-prove day-1 runway and ordered failover.

Cutting keys while the old webhook still holds in-flight DLR drops delivery truth mid-air. Buyers see “sent” with no final status; finance sees open holds that never close; ops loses the trail that proves failover order. Drain first. Cut second.

IOSOR rotation treats the old endpoint as a queue that must go quiet — not as a switch you flip when the new URL answers one smoke. In-flight reports are part of the message lifecycle; orphan them and you invent a week of ghost rows.

Inventory in-flight DLR on the old endpoint

Export open deliveries and pending final statuses tied to the old webhook URL. Tag each row with age and message id. That inventory is the drain backlog — not a vague hope that “retries will finish somehow”.

Share the export with finance before anyone schedules a key revoke. If the backlog is larger than one quiet shift, extend the drain window in writing instead of cutting into a storm of late DLR.

Drain until quiet, then cut keys

Keep the old endpoint Live only for completion events until the inventory hits zero or an agreed residual floor. Do not revoke messaging or webhook secrets while the drain list still shows fresh arrivals. Quiet means no new finals for a measured window — not “looks calm in Slack”.

When quiet is proven, revoke old keys in the same change window as the sole-owner URL flip. Partial revoke that leaves a half-live secret is how late DLR arrives on a dead path.

Keep failover order honest during drain

While draining, do not invent a second Live path that bypasses the ordered backup story. Failover stays primary-then-backup; drain is not permission to fan out. Document which path owns in-flight rows so incident week does not argue about which rail “should have” answered.

If primary fails mid-drain, follow the ordered backup path and re-inventory — do not cut keys as a panic move to “simplify”.

Re-prove day-1 runway after the cut

After keys are cut, run the day-1 runway checks that must be green: heartbeat, send proof, webhook finals on the new endpoint, wallet hold close. Only then reopen pilot volume. A drained old path plus a red runway is still a blocked launch.

Export the first quiet hour on the new URL so ops can show buyers that DLR did not vanish with the old secret.

Related ops paths

Start with IOSOR

Inventory in-flight DLR on the old endpoint, set a quiet drain threshold, and keep failover order honest until the queue is empty. Cut keys only after the drain report is green, then re-prove day-1 runway on the new path.

IOSOR takeaway

Drain first, then cut. Dropping keys while DLR is still in flight orphans delivery truth and breaks buyer tickets. Failover order must stay honest during the drain — a silent backup that never fires is not a plan.

Do not archive the old endpoint until quiet. Do publish the drain report and the post-cut runway check before Live volume rises.

Was this guide helpful?

Related guides