IOSOR Learn
DID incident week: messaging down is not Activated
How to handle your first DID messaging incident during an outage, manage prepaid holds without shop-stock fiction, and communicate honest statuses.
DID incident week.
Messaging down means routing failure, not a shop restock
When messaging fails on a newly provisioned number, your first instinct might be to check inventory or look for restock alerts. In white-label CPaaS operations, there is no idle stock pool or physical shelf. Numbers are instantiated via JIT provisioning. If inbound SMS or OTP delivery halts, the issue sits in routing tables, webhook dispatchers, or upstream gateway handshakes—never in a 'sold out' bin. Treat every outage as a live network exception rather than a merchandising error.
Immediate freeze on assignments and send queues
As soon as clients report dropped DLRs or silent OTP flows, immediately freeze automated number assignment and high-volume send queues. Letting scripts continue to allocate routes during an active degradation compounds the blast radius. Put a temporary hold on the prepaid balance allocation for affected sub-accounts. Communicate clearly that the incident is under active engineering review, keeping your minimum USD 20 prepaid floor intact while support teams trace the HB and API payload logs.
Verifying readiness before blaming the network
Before escalating an incident, verify that the affected number meets baseline protocol requirements. Many perceived outages stem from skipped validation steps outlined in the DID messaging readiness before production guide. Check 10DLC registration status, brand compliance, and webhook URL responsiveness. If headers return 5xx errors, the bottleneck is on the application endpoint, not the carrier network.
Swapping, refunding, or releasing failed assets
If an underlying routing path is permanently degraded and cannot be recovered within SLA limits, do not leave the client hanging. Execute a clean swap or issue an automated credit. Review the protocol for a DID order fail refund and swap to ensure balance adjustments clear correctly. Prepaid holds must be released immediately so the tenant can provision a working asset without double-paying for failed infrastructure.
Financial predictability past the honeymoon phase
Operational incidents often coincide with scaling milestones. Once a tenant moves past initial testing and approaches the soft review near USD 1,000/month, traffic patterns shift from sporadic OTP bursts to sustained A2P campaigns. Keep a close eye on your DID Second Month: Full MRC when the UTC Calendar Rolls cycles to ensure recurring charges and usage top-ups reconcile cleanly without triggering false-positive fraud suspensions during active troubleshooting.
Start with IOSOR for native white-label reliability
When DLR or the messaging webhook dies, freeze the send queue on that DID. Do not keep MT because the number row still says assigned. Export the freeze time, last good DLR, and a messaging-down status. Resume only after a live smoke on the same digits. This is not a shop unavailable badge and not an invoice dispute.
IOSOR takeaway
Messaging-down is a freeze, not an inventory outage.
Do: halt queues and tell tenants messaging is down. Don’t: keep sending, or relabel the DID as missing stock.
Was this guide helpful?
Related guides
- Second-owner DID handover: who may assign and release
Master operational boundaries, JIT provisioning, and prepaid financial thresholds during second-owner DID handovers.
- Spend Cap Per DID: Rent Plus MT Burn On One Number
Control per-number exposure in your white-label CPaaS with a combined spend cap for MRC and outbound mobile terminated traffic.
- Inbound webhook routing on DID: MO without owner loses STOP
Route inbound webhooks to the owning account securely. Prevent orphan MO events and missed opt-outs in white-label prepaid CPaaS.