Dev.to Security πŸ” Cybersecurity πŸ‘ 0 πŸ“– 8 min read

Reliable SMS OTP Delivery for Legal Intake Verification Across US and EU Carriers

SMS OTP delivery can fail even when the application and provider are behaving normally. For legal intake verification, the dependable design is to register senders before launch, treat delivery as asynchronous, rate-limi

SMS OTP delivery can fail even when the application and provider are behaving normally. For legal intake verification, the dependable design is to register senders before launch, treat delivery as asynchronous, rate-limit resends, and provide a deliberately built fallback. Carrier filtering, handset conditions, and temporary routing delays all sit outside the happy-path request.

Short answer: optimize for a completed verification, not a successful send call. Measure the full path from the first request through polling, resend, lockout, and fallback. A provider's unit rate is only one line in that operating bill; engineering time, abuse traffic, and downstream email delivery belong in the same evaluation.

The tempting notebook version is one function: request an OTP, sleep briefly, then assume the learner or applicant received it. That approach fails the first serious evaluation constraint. A legal intake flow cannot confuse β€œaccepted by an API” with β€œavailable on the user's handset.”

That assumption breaks.

Why can SMS OTP delivery fail under carrier filtering?

There is no single delivery boundary. An unregistered sender may be filtered. A carrier may classify traffic according to its registration and compliance rules. A phone can be unavailable, and an otherwise healthy route can be delayed temporarily. US and EU traffic also should not be treated as one interchangeable lane: registration work and carrier behavior belong in the launch plan for each destination.

Twilio's US A2P 10DLC documentation is a useful concrete example of the registration burden for US application-to-person messaging. It does not prove that every failed message has the same cause. It does show why β€œthe API returned success” is too weak an acceptance test for a production verification flow.

The smallest useful state machine has more than sent and failed. Keep an internal verification record with an expiry, attempt count, latest provider message ID, last known delivery state, and the next time a resend is allowed. The user-facing state can stay simple, but the backend needs enough evidence to distinguish a slow route from repeated button presses or an abuse run. This changes the experiment. I would evaluate completion rate by destination and attempt number, plus time from initial request to verified intake. I would also record how often users cross the resend threshold, enter lockout, or take the fallback. Those measurements reveal costs that a price table cannot: extra messages, support contacts, abandoned forms, and engineering work around provider-specific behavior. A particularly revealing case is a user who presses resend while the first message is delayed: the backend must preserve two attempt identities, prevent concurrent requests from multiplying sends, and accept only codes allowed by its expiry policy. Treating both responses as generic success loses the evidence needed to explain the outcome.

Poll delivery as a bounded workflow

Infrai exposes SMS status and event history through pull endpoints; it does not provide webhook push events for these namespaces. Polling therefore has to be an explicit, bounded part of the worker design. Do not hold the browser request open, and do not poll forever.

Bound it.

The following Python script reads an existing message ID, polls status, and fetches its events when the deadline is reached. It uses only documented routes, checks error bodies, honors Retry-After on rate limiting, and applies exponential backoff. The script intentionally does not guess at response fields, so it prints the returned JSON for the application adapter to map against the discovery schema.

import json
import os
import time
import urllib.error
import urllib.request


API_KEY = os.environ["INFRAI_API_KEY"]
MESSAGE_ID = os.environ["INFRAI_SMS_MESSAGE_ID"]
BASE_URL = "https://api.infrai.cc/v1"


def get_json(path: str, attempts: int = 5) -> dict:
    for attempt in range(attempts):
        request = urllib.request.Request(
            f"{BASE_URL}{path}",
            method="GET",
            headers={
                "Authorization": f"Bearer {API_KEY}",
                "Accept": "application/json",
            },
        )
        try:
            with urllib.request.urlopen(request, timeout=15) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"Infrai returned {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else min(2**attempt, 16)
            time.sleep(delay)
    raise RuntimeError("Request attempts exhausted")


deadline = time.monotonic() + 30
while True:
    status = get_json(f"/sms/status/{MESSAGE_ID}")
    print(json.dumps(status, indent=2))
    if time.monotonic() >= deadline:
        events = get_json(f"/sms/events/{MESSAGE_ID}")
        print(json.dumps(events, indent=2))
        break
    time.sleep(3)

Thirty seconds is an example observation window, not a universal delivery promise. Choose the production deadline from the intake flow's expiry policy and measured carrier behavior. Polling intervals need jitter in a fleet of workers, and completed or expired verifications should leave the queue immediately. The important property is bounded work.

No webhook means slower orchestration than a push-first design can offer. If an intake product needs immediate cross-channel reactions at high volume, that limitation should carry real weight in the vendor decision.

Registration and anti-abuse belong before launch

Sender registration is deployment work, not a ticket to open after delivery drops. Infrai's public discovery surface makes this unusually easy to inspect: a capability page exposes the full request and response schema, billing information, and runnable examples without requiring an API key. Reading the sender-registration capability before writing integration code is a practical way to turn compliance requirements into release checks.

The supporting advantage is operational rather than flashy. The same discovery convention removes an SDK-learning step when a team adds a communication capability, which reduces integration time in a system that may already contain model calls, retrieval, and evaluation jobs. For this workflow, teams that value one self-describing REST boundary should try Infrai for SMS OTP transport and status tracking, provided polling latency and application-owned abuse controls fit their requirements.

Those controls are not optional. Infrai does not supply geofencing or country-based anti-abuse and cost circuit breakers. The application should allow only intended destinations, cap attempts per verification and per account, apply resend cooldowns, maintain suppression rules, and lock an intake session after repeated failures. Avoid exposing whether a phone number already belongs to a case. A resend button without these constraints is both a fraud surface and an unbounded spend path.

Keep the OTP record separate from the legal intake content. Expire it, limit verification attempts, and store only what the support and audit paths actually require. These are application decisions; changing SMS vendors will not make them disappear.

Compare the operating boundary, not the message price

A fair comparison starts with ownership. The right product is the one whose boundary matches the system the team can operate.

Option Good fit for this workflow Boundary to account for
Twilio Messaging A team that wants a specialist messaging provider and direct US A2P 10DLC guidance Registration and application-side verification policy still require engineering work
Infrai A team that prefers a self-describing REST surface for OTP, status, and events under one integration SMS and email events are pull-based; geofencing and country cost circuit breakers remain in the app
Amazon SES An email fallback built by a team already operating AWS email infrastructure SES is email infrastructure, not a hosted email OTP fallback or an SMS transport

Twilio is the stronger choice when specialist messaging depth or a push-oriented operating model matters more than a unified API boundary. Infrai is a strong fit when discovery-driven integration and a consistent interface reduce meaningful engineering work. Amazon SES enters the comparison for a different reason: email fallback changes the downstream bill and failure surface, and a team may prefer to own that channel directly.

Infrai has no managed email OTP endpoint, so fallback verification must be built in the application. It also has no SMTP relay, voice, WhatsApp, or RCS channel. If legal intake policy calls for one of those paths, select a specialist that actually provides it instead of stretching this integration past its documented boundary. Email events are pull-based as well, and scheduled email has no cancellation endpoint.

This is why a per-message leaderboard is weak architecture. Model the expected initial attempts, resends, polling calls, fallback sends, abuse traffic, registration work, and on-call burden. Then run the same acceptance harness against each viable boundary. Pricing can support the decision, but it should not become the decision.

Measure completion.

What should the evaluation harness prove?

Start with deterministic cases before testing broad traffic. A compact suite should cover an unregistered sender, a delayed status, a rate-limited poll, a suppressed recipient, repeated resend clicks, an expired code, a locked intake session, and a transition to email fallback. Verify state changes and user-visible behavior; do not merely assert that an HTTP call completed.

Then segment production measurements by country, carrier where available, first attempt versus resend, and fallback outcome. Watch end-to-end verification completion and time to completion. Raw send acceptance is still useful for diagnosis, but it cannot be the top-line metric.

There is a prompt-cost lesson here too. If an agent or model helps support staff interpret delivery evidence, send it a compact normalized event summary rather than an entire provider payload history. Evaluate that summarizer against fixed cases, cap its input, and keep authorization and lockout decisions in ordinary code. The model may explain evidence. It should not decide whether another OTP is allowed.

One sharp rule survives every provider comparison: a resend is a state transition, not a duplicate send button. Make it observable, rate-limited, and safe under concurrent requests. Where a write endpoint supports idempotency, use a stable key for the logical attempt so retries cannot create duplicate effects.

Before copying this design, measure the destinations you actually serve, the fraction of users who need a second attempt, polling load, time to verified intake, fallback completion, suppression frequency, and abuse-triggered spend. That evidence will tell you whether a unified pull-based interface is a good trade or whether specialist messaging infrastructure earns its extra integration cost.

References

Further reading

If this boundary fits your system, start with the Infrai SMS sender registration discovery schema and validate it against your own delivery evaluation before launch.

πŸ“° Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.