IOSOR Learn

DID incident week: messaging down is not Activated

How to handle your first DID messaging incident during an outage, manage prepaid holds without shop-stock fiction, and communicate honest statuses.

DID incident week.

Messaging down means routing failure, not a shop restock

When messaging fails on a newly provisioned number, your first instinct might be to check inventory or look for restock alerts. In white-label CPaaS operations, there is no idle stock pool or physical shelf. Numbers are instantiated via JIT provisioning. If inbound SMS or OTP delivery halts, the issue sits in routing tables, webhook dispatchers, or upstream gateway handshakes—never in a 'sold out' bin. Treat every outage as a live network exception rather than a merchandising error.

Immediate freeze on assignments and send queues

As soon as clients report dropped DLRs or silent OTP flows, immediately freeze automated number assignment and high-volume send queues. Letting scripts continue to allocate routes during an active degradation compounds the blast radius. Put a temporary hold on the prepaid balance allocation for affected sub-accounts. Communicate clearly that the incident is under active engineering review, keeping your minimum USD 20 prepaid floor intact while support teams trace the HB and API payload logs.

Verifying readiness before blaming the network

Before escalating an incident, verify that the affected number meets baseline protocol requirements. Many perceived outages stem from skipped validation steps outlined in the DID messaging readiness before production guide. Check 10DLC registration status, brand compliance, and webhook URL responsiveness. If headers return 5xx errors, the bottleneck is on the application endpoint, not the carrier network.

Swapping, refunding, or releasing failed assets

If an underlying routing path is permanently degraded and cannot be recovered within SLA limits, do not leave the client hanging. Execute a clean swap or issue an automated credit. Review the protocol for a DID order fail refund and swap to ensure balance adjustments clear correctly. Prepaid holds must be released immediately so the tenant can provision a working asset without double-paying for failed infrastructure.

Financial predictability past the honeymoon phase

Operational incidents often coincide with scaling milestones. Once a tenant moves past initial testing and approaches the soft review near USD 1,000/month, traffic patterns shift from sporadic OTP bursts to sustained A2P campaigns. Keep a close eye on your DID Second Month: Full MRC when the UTC Calendar Rolls cycles to ensure recurring charges and usage top-ups reconcile cleanly without triggering false-positive fraud suspensions during active troubleshooting.

Start with IOSOR for native white-label reliability

When DLR or the messaging webhook dies, freeze the send queue on that DID. Do not keep MT because the number row still says assigned. Export the freeze time, last good DLR, and a messaging-down status. Resume only after a live smoke on the same digits. This is not a shop unavailable badge and not an invoice dispute.

IOSOR takeaway

Messaging-down is a freeze, not an inventory outage.

Do: halt queues and tell tenants messaging is down. Don’t: keep sending, or relabel the DID as missing stock.

Was this guide helpful?

Related guides