IOSOR Learn

UCS-2 concatenation silently burns prepaid SMS segments

One emoji or a Unicode character can flip GSM-7 to UCS-2, split a “single message” into billed segments, and leave finance reading composer counts instead of prepaid segment lines.

Product typed one SMS. The prepaid wallet debited three units. That is concatenation after a silent encoding flip — not a ledger glitch. GSM-7 holds a modest Latin set; anything outside it (emoji, many scripts, a “smart” quote) forces UCS-2, drops the per-segment budget, and splits the body. Finance that still reports “messages sent” will never reconcile the wallet.

IOSOR runs white-label prepaid SMS on one ledger. Catalog live means the send path is ready; in setup is a request, not cheaper encoding. Near USD 1,000+ monthly usage, segment count and encoding become commercial review material. No platform subscription to keep the account alive — fund prepaid and see units.

Why finance must count segments, not messages

A composer that shows “1 message” answers a product question. The wallet answers a money question: how many billed segments left. Those numbers diverge when encoding changes or length crosses a boundary. Pair with SMS segment accounting — pricing explains the unit; this page explains the silent burn.

If finance cannot export destination, encoding, length, segment count, and unit rate, the row is a receipt, not a control. See prepaid spend control before promising a monthly SMS budget from “message” totals.

GSM-7 versus UCS-2: the silent encoding flip

GSM-7 is efficient and brittle. UCS-2 is honest and expensive. The flip is invisible in QA if authors test English-only Latin: one emoji recodes the entire body; a smart quote or checkmark does the same; template variables that look short in English expand elsewhere and drag encoding with them.

UCS-2 is the encoding the body actually used — not an “international surcharge.” Show it next to character count in the composer and the API estimate, not only after the debit.

Concatenation overhead that the composer hides

Encoding Single-part limit Multipart limit What the header steals
GSM-7 160 chars 153 chars Concatenation header
UCS-2 70 chars 67 chars Same header, smaller budget

Templates and locales that cross the boundary

Segment burn hides where support looks last: a variable three characters in the authoring locale and thirty for the recipient; a legal footer with curly quotes; an emoji added “for warmth” on failed DLR retries; a multilingual template QA’d only on the writer’s keyboard.

Red flags

  • Composer or API returns “messages” instead of segments
  • Encoding hidden until month-end
  • Emoji allowed in OTP templates with no segment warning
  • Ledger rows that cannot show encoding + segment count
  • Bulk sends billed as a flat guess, reconciled later
  • Catalog in setup promised as if encoding math already applied

Start with IOSOR

Open the IOSOR console and activate the pre-flight encoding gate on all outbound campaign templates. Set up webhooks to inspect message payloads for non-GSM-7 characters like smart quotes or emojis before queueing dispatches. Place an automated hold on any dispatch where variable expansions force UCS-2 concatenation beyond your target segment budget.

IOSOR takeaway

Measuring campaigns by message count rather than billed segments guarantees unexpected budget burn. When a single non-GSM-7 character triggers UCS-2 encoding, your multi-part limit drops from 153 to 67 characters, instantly doubling or tripling segment consumption across localized templates and automated retries.

Do audit pre-flight segment calculations inside your delivery pipeline and strip unexpected unicode artifacts automatically before billing. Don't trust basic UI message counters or assume localized template variables will fit within single-segment boundaries.

Was this guide helpful?

Related guides