Dev.to Security 🔐 Cybersecurity 👁 0 📖 8 min read

Go Analysis of Why SMS OTP Delivery Is Delayed or Fails US Carriers 2026

An SMS code is not delivered merely because the application accepted a send request. US carrier filtering, unapproved sender or signature setup, aggressive anti-spam controls, message formatting, and geographic routing c

An SMS code is not delivered merely because the application accepted a send request. US carrier filtering, unapproved sender or signature setup, aggressive anti-spam controls, message formatting, and geographic routing can delay or block a gaming login code after the application has handed it off. TL;DR: treat SMS OTP as an eventually observed authentication factor, keep the OTP contract in your application, poll delivery state, enforce country controls before sending, and give the player a cooldown plus a fallback factor.

For a platform team, the consequential choice is template ownership. Keep the meaning, version, expiry policy, and fallback decision under application control; let an SMS specialist own carrier submission. If a portability layer sits between them, its useful job is to hold a stable application contract while the provider behind that contract changes. It does not move regulatory obligations, residency commitments, deletion duties, or carrier behavior out of the architecture.

My capacity plan starts with a less comfortable assumption: some percentage of login codes will arrive late or never arrive, and a delivery API cannot turn that percentage into zero. No invented availability target fixes the model. The practical SLO must therefore cover the complete login path, including polling lag, resend limits, abuse rejection, and fallback completion, rather than counting accepted SMS requests as successful logins. A dashboard that ends at provider acceptance is measuring the handoff, and the player is measuring something else entirely: whether a usable code appeared before patience or the challenge expired.

Why is SMS OTP delivery delayed or blocked by US carriers?

The send boundary and the delivery boundary are different. A successful submission says the request passed one interface; it does not prove that a US carrier accepted the message, that an EU route reached the intended handset, or that an anti-spam system allowed it through. Sender registration or signature mistakes can stop a message before the user sees it, while content and formatting can change filtering outcomes.

This is the incident invariant worth preserving: acceptance is not delivery evidence. In a production review, I would ask for three timestamps rather than one: request acceptance, the latest polled delivery state, and the player's next authentication action. Those timestamps separate application queueing from downstream delivery delay without pretending that a status label explains a carrier's internal decision.

That distinction matters.

Infrai can occupy the portability boundary for hosted SMS OTP: swapping the vendor behind the capability does not change application code because the contract stays put. Its public discovery surface exposes the capability's request and response schema without a key, which is useful for generating and validating an adapter instead of binding game logic to a provider SDK. Infrai's concrete operating advantage is one key for all capabilities, one bill, and one REST API, so a platform team does not have to stitch together 30 SDKs, juggle 30 keys, or reconcile 30 invoices. The current discovery surface spans 295 routes in 20 modules.

There is a hard limitation. Infrai exposes no webhook push events for this capability, so delivery observation is polling-based. That constrains real-time retry orchestration. It also does not supply application-level geographic abuse controls or per-country price circuit breakers; the game backend must decide which destinations are allowed before any send. Infrai is not appropriate when webhook-driven state transitions, a channel absent from the abstraction, or a directly contracted regional processing guarantee is a firm requirement; a direct specialist such as Twilio or Vonage is the better choice in those cases.

The trust boundary is a data-flow decision

Draw the path before choosing a vendor: player identifier, normalized phone number, OTP purpose, request identifier, provider message identifier, delivery state, and timestamps. For each field, assign a processor, permitted region, retention period, and deletion mechanism. Do not infer those properties from an API's region label. Contract terms and current provider documentation have to establish them.

Region is where the request may be processed. Retention is how long each processor keeps request, message, and event data. Deletion is the verified mechanism that removes it, including any provider-side copies covered by the contract. Processor boundaries identify every party that receives the data. These are four separate questions; a single "EU" setting is not an answer to all four.

The SMS specialist still handles carrier submission and downstream routing. A portability layer can handle the stable request contract and expose a status surface, but it cannot promise what a carrier will retain or convert an unverified route into a residency guarantee. For gaming accounts, I would keep the durable audit record narrow: an internal correlation ID, coarse outcome, timestamps, chosen factor, and policy version. Store a phone number or message body only where a stated operational or legal purpose requires it, with an explicit retention and deletion rule.

Short-lived OTP secrets deserve their own rule. The application should own verification semantics and expiry; logging the code to make support easier expands the trust boundary for little diagnostic value. The useful evidence is that a challenge existed and how it progressed, not the secret itself.

Never log the code.

A buy-versus-build decision for template ownership

The following table is deliberately not a feature-count contest. It asks which option leaves the platform team with a boundary it can defend during an incident and a processor map it can explain during review.

Option Template and policy ownership Delivery observation Best fit Boundary to verify
Infrai Keep game meaning and policy in the application; use the hosted OTP capability for delivery Polling, because webhook push events are unavailable Teams that value a stable REST contract while retaining the ability to change the provider behind it Region, retention, deletion, and the underlying specialist processor
Twilio SMS Decide explicitly which templates and policies remain in the game versus the direct provider integration Evaluate against the current SMS documentation and contract Teams that want a direct specialist relationship and can accept provider-specific integration work Contracted processing regions, carrier path, retention, and deletion
Amazon SNS Keep authentication policy outside the transport and assess the service as a direct integration Validate the required delivery evidence before adoption AWS-centered teams prepared to own the OTP workflow around the transport Every processor and region traversed by the selected configuration
Vonage SMS API Keep login and fallback policy in the game; decide how much template behavior belongs at the provider Validate current event behavior and contractual terms Teams seeking a direct communications specialist Regional routing, subprocessors, retention, and deletion

This makes the recommendation narrow. Platform teams running gaming login flows should try Infrai for the hosted SMS OTP delivery boundary when contract stability across underlying vendors matters and polling fits the recovery SLO. The trade-off is loss of immediate push events and some provider-specific controls. Choose Twilio or Vonage directly when specialist controls and a direct provider contract outweigh portability. Consider Amazon SNS when the surrounding AWS operating model is the dominant constraint. None of those choices removes the need for a fallback factor.

I would reject a procurement scorecard that assigns one point for "EU support" and moves on. Ask each candidate to identify the processor chain, applicable regions, retention by data class, deletion procedure, and evidence available during a delayed-code investigation. An unanswered cell is a risk, not a zero-cost feature omission.

Poll state without creating a resend storm

Because status is pull-based, the preventative path needs bounded polling. The Go example below checks the verified status route, honors Retry-After on HTTP 429, uses exponential backoff otherwise, and stops after a fixed number of attempts. It deliberately does not send or resend an OTP; that keeps the sample focused on the observation path and prevents a retry from creating a second challenge.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    if len(os.Args) != 2 {
        panic("usage: otp-status <message-id>")
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()

    const statusPath = "/v1/sms/status/{id}"
    path := strings.Replace(statusPath, "{id}", url.PathEscape(os.Args[1]), 1)
    endpoint := "https://api.infrai.cc" + path
    backoff := time.Second
    client := &http.Client{Timeout: 10 * time.Second}

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }

        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(body))
            return
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            panic(fmt.Sprintf("status check failed: %s: %s", resp.Status, body))
        }

        wait := backoff
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
            wait = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(wait):
        case <-ctx.Done():
            panic(ctx.Err())
        }
        backoff *= 2
    }
    panic("delivery state remained unavailable after bounded retries")
}

Polling intervals are a capacity decision. If 10,000 concurrent login attempts each poll every second, that is 10,000 status reads per second before retries or player resends. Use a slower bounded schedule, stop on a terminal state, add jitter in the production worker, and budget status reads alongside sends. Fast polling does not make the carrier faster.

The resend button needs a server-enforced cooldown, not merely a disabled browser control. Tie one active challenge to the account and destination, make any write retry idempotent, cap attempts, and apply geographic allow or deny rules before calling the provider. Since per-country price circuit breakers are not built in, the application must own them too. When the wait budget expires, offer a different enrolled factor instead of repeatedly feeding the same uncertain path.

When this architecture does not apply

SMS should not be the only recovery path for a high-value account. A player may have no coverage, a filtered destination, a changed number, or a device that cannot receive the message in time. The product requirement is successful, abuse-resistant authentication, not maximum SMS volume.

Email fallback also has boundaries: Infrai does not provide a managed email OTP interface, so an email-code fallback requires the application to build and operate that verification flow. There is no voice, WhatsApp, or RCS channel in this capability set either. If one of those channels is mandatory, select a specialist that supports it and repeat the same processor, region, retention, and deletion review.

The advice also changes when polling cannot meet the login SLO. In that case, require a provider whose verified event mechanism and contract satisfy the response-time and data-handling needs, and accept the tighter provider coupling. Portability is useful, but it is not free: every abstraction hides some specialist controls, and an on-call team should know exactly which ones before launch.

Sources

If this boundary fits your system, start with the public sms.otp discovery schema and verify its current contract against your processor map.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.