IOSOR Learn
Handling Active Traffic with a Stale Webhook Heartbeat
Learn how to manage active SMS and OTP traffic when your webhook heartbeat goes stale, avoiding false positive failovers on the IOSOR platform.
Handling Active Traffic with a Stale Webhook Heartbeat.
Analyzing Traffic OK with Stale Webhook Heartbeat
When your core SMS and OTP traffic flows normally but your webhook heartbeat goes stale, you face a silent observability failure. Buyers must distinguish between a complete platform outage and a localized delivery path failure. If DLRs are successfully processed but the heartbeat endpoint fails to respond, your automated systems might trigger unnecessary failovers.
Ledger Actions and Prepaid Hold Mechanics
To keep your E.164 routing active during these incidents, IOSOR maintains strict ledger rules. Every JIT number assignment requires a prepaid hold to secure the resource. Your account must maintain the USD 20 prepaid floor to prevent automatic outbound suspension.
Diagnostic Steps for Webhook Delivery
Verify that your application is receiving actual OTP and Verify traffic even if the heartbeat is dead. Check your webhook logs for 504 gateway timeouts or 403 forbidden errors. Often, a stale heartbeat is caused by a routing misconfiguration on the buyer's firewall rather than an IOSOR platform issue. Ensure that your endpoints can handle concurrent DLR payloads without dropping the lightweight heartbeat ping that monitors system health.
Mitigating False Positives in Production
Do not rely solely on a single heartbeat ping to declare a routing disaster. Implement a multi-factor health check that combines heartbeat status with real-time DLR success rates. If your DLR delivery rate remains above 95%, keep your active routes open. This prevents costly and unnecessary failover actions that disrupt active E.164 sessions and trigger redundant JIT provisioning fees.
Observability and Failover Resources
To build a resilient integration, review our detailed guides on webhook management and automated failover strategies:
- Heartbeat and smoke gates before paging humans
- Monitoring Consumer Webhook Endpoint Health Metrics
- Failover incident export at 02:00
These resources help you configure advanced thresholds and export incident
Start with IOSOR
Audit your webhook alert gates inside the IOSOR console before turning heartbeat delays into public incident reports. Validate whether active OTP DLR flows are still delivering to prevent false-alarm failovers. If live delivery metrics remain green, update your automated status rules to flag webhook transport issues without tearing down healthy SMS routes.
IOSOR takeaway
A stale webhook heartbeat is an observability warning, not an automatic confirmation of carrier downtime. Treating every silent heartbeat ping as a full system outage causes unnecessary routing failovers while real DLR traffic continues to clear successfully.
Do cross-verify synthetic heartbeats against actual OTP delivery throughput before publishing external status incidents or altering active route assignments. Don't rely on a single heartbeat check as a binary smoke-test checkbox for total platform failure.
Was this guide helpful?
Related guides
- Status Page Must Match the Send Pause
Learn how to automatically align your public status page with active send pauses in IOSOR to maintain trust and prevent unnecessary API retries.
- Buyer Incident Language vs Internal Smoke Signals
Learn how to translate internal CPaaS telemetry and stale heartbeats into clear, buyer-facing traffic_ok status updates without exposing raw infrastructure logs.