Dev.to Security 🔐 Cybersecurity 👁 0 📖 8 min read

Go SMS OTP vs Email OTP for Two-Factor Login (Storefront Runbook)

Use hosted SMS OTP as the primary second factor for a US/EU SaaS login, and keep email OTP as a backup only when the platform team is prepared to own the email-code lifecycle. TL;DR: SMS is the simpler operational defaul

Use hosted SMS OTP as the primary second factor for a US/EU SaaS login, and keep email OTP as a backup only when the platform team is prepared to own the email-code lifecycle. TL;DR: SMS is the simpler operational default for interactive authentication; email requires the application to generate, store, expire, and validate codes while also absorbing inbox delay. In an e-commerce system, keep both authentication templates separate from the order receipt sent after payment settles, because a receipt is durable business communication while an OTP is short-lived credential material.

This is a template-ownership decision before it is a transport decision. A hosted SMS verification product couples challenge issuance, delivery, and verification. A normal email send API does not. Giving the product team complete control of an email template also gives the engineering team responsibility for every security transition behind it.

Should SMS OTP or email OTP handle two-factor login?

An OTP is state, not six digits in a message. An email fallback needs an unpredictable code, a stored digest, an expiry, a failed-attempt limit, single-use consumption, and a rule that makes a resend supersede the previous challenge. Those transitions must remain atomic when two application replicas receive the same verification attempt. The login SLO includes all of them.

SMS verification narrows that ownership boundary. It usually gives the messaging provider more control over the verification text, which may frustrate a brand team, but the constraint is useful when the alternative is letting a receipt-template editor alter security semantics. The order receipt can remain a rich, versioned email template owned by commerce; the OTP template and its state machine should be owned by identity.

Keep them separate.

Do not infer delivery from API acceptance. Email delivery and inbox placement can delay a login code, while SMS generally fits interactive authentication better for US/EU products. Neither channel is universal, and neither should be described as immune to interception, account compromise, carrier filtering, or user error.

The practical buy-versus-build boundary looks like this:

Option What the vendor owns What your team still owns Best fit and boundary
Twilio Verify A managed verification workflow Account policy, recovery, regional testing, and abuse controls Teams that want a dedicated verification product and accept its workflow constraints
Vonage Verify A managed verification workflow Identity state, recovery policy, destination controls, and SLOs Teams evaluating a second established managed-verification option for their destination mix
Amazon SNS SMS message transport Code generation, expiry, validation, throttling, and resend semantics Teams deliberately building the authentication state machine on lower-level messaging primitives
Infrai Hosted SMS OTP, plus ordinary email sending rather than hosted email OTP The entire email-code lifecycle and polling-based channel orchestration Teams that value a plain REST integration and can accept those ownership limits

This is not a price ranking. Email may have a lower marginal delivery cost as a fallback, but on-call load, abuse handling, delayed logins, and the credential store are part of its cost model. Twilio Verify and Vonage Verify are the clearer candidates when a dedicated managed-verification product is the desired boundary; Amazon SNS fits when owning verification state is intentional.

Infrai is a credible unified option where one plain REST API is preferable to installing and upgrading another client library. Its second relevant advantage is operational consolidation: one key can cover the hosted SMS login path and the email order-receipt path, while one bill reduces credential rotation and invoice attribution work for the platform team. The public discovery surface is self-describing without a key and exposes request JSON Schema, so an adapter can validate its contract before deployment. Its main limitation here is the absence of hosted email OTP and webhook-driven failover, which makes it unsuitable when either is mandatory. There is also no SMTP relay, and voice, WhatsApp, and RCS are outside this channel set; choose Twilio Verify or Vonage Verify when a dedicated managed-verification boundary matters more, or Amazon SNS when the platform team deliberately wants lower-level transport and will own the state machine.

Inspect the contract before issuing a challenge

The expensive mistake is guessing a convenient OTP request body from an old example. Query the discovery document first, confirm that the capability is available, and generate or validate the adapter from the returned params schema. The following runnable Go program performs that read, sets its HTTP method explicitly, checks non-success responses, and handles HTTP 429 with bounded exponential backoff or Retry-After.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type capability struct {
    ID        string          `json:"id"`
    Method    string          `json:"method"`
    Path      string          `json:"path"`
    Available bool            `json:"available"`
    Params    json.RawMessage `json:"params"`
}

func retryDelay(value string, attempt int, now time.Time) time.Duration {
    if seconds, err := strconv.Atoi(strings.TrimSpace(value)); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    if at, err := http.ParseTime(value); err == nil && at.After(now) {
        return at.Sub(now)
    }
    return time.Duration(1<<attempt) * time.Second
}

func discover(ctx context.Context, client *http.Client) (capability, error) {
    endpoint := os.Getenv("INFRAI_DISCOVERY_URL")
    if endpoint == "" {
        return capability{}, fmt.Errorf("INFRAI_DISCOVERY_URL is required")
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return capability{}, fmt.Errorf("INFRAI_API_KEY is required")
    }
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return capability{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Accept", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return capability{}, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return capability{}, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            timer := time.NewTimer(retryDelay(resp.Header.Get("Retry-After"), attempt, time.Now()))
            select {
            case <-ctx.Done():
                timer.Stop()
                return capability{}, ctx.Err()
            case <-timer.C:
                continue
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return capability{}, fmt.Errorf("discovery returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
        }
        var result capability
        if err := json.Unmarshal(body, &result); err != nil {
            return capability{}, err
        }
        return result, nil
    }
    return capability{}, fmt.Errorf("discovery remained rate limited")
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
    defer cancel()

    result, err := discover(ctx, &http.Client{Timeout: 10 * time.Second})
    if err != nil {
        panic(err)
    }
    out, err := json.MarshalIndent(result.Params, "", "  ")
    if err != nil {
        panic(err)
    }
    fmt.Printf("%s %s (available=%t)\n%s\n", result.Method, result.Path, result.Available, out)
}

The discovery read is public, although this example deliberately exercises the same environment-based Bearer credential path as the subsequent authenticated adapter. Set INFRAI_DISCOVERY_URL to the current discovery capability URL and never hardcode INFRAI_API_KEY. For every issue operation, use a client-generated operation identifier as an idempotency key, keep retries bounded, and surface the body of non-success responses to controlled internal logs. A retry must not create two active challenges.

Build the email fallback only after its state machine is reviewed as authentication code. Store a digest rather than the raw code; bind it to account, purpose, issue time, expiry, failed-attempt count, and consumed state; invalidate the prior value on resend. Do not copy arbitrary expiry or attempt-count numbers into production without a threat model.

Size the fallback path against the login SLO

Capacity planning starts with peak authentication arrivals, not average daily messages. If a modeled peak is 120 login attempts per second and 8% request email fallback, the fallback tier receives 9.6 requests per second before resends, automated abuse, or a primary-channel impairment. These are example inputs, not measured vendor performance. Replace them with your own peak, fallback ratio, resend multiplier, and failure budget, then size the transactional store, send workers, polling workers, and provider quota for the resulting burst.

Track challenge acceptance and successful verification as separate SLO indicators. Acceptance measures whether the application and provider took the work. Verification includes the user-visible delivery and entry path. Combining them hides an inbox-placement problem behind a healthy API graph.

Cross-channel failover is not immediate here because both email and SMS events are polling-based. Pick a bounded polling interval, allow the user to request the backup explicitly, invalidate superseded challenges, and rate-limit by account, IP, destination, and relevant device signals. The business layer must also provide geographic fencing and country-level spend circuit breakers. Do not expect tag-aggregated provider cost reports to reconstruct tenant or authentication intent; record stable operation and tenant identifiers in internal telemetry. Consider the concrete failure sequence: the SMS issue call is accepted, the user requests email before the next status poll, and the original SMS arrives after the email code has been generated. If both values remain valid, a convenience feature has doubled the active credential surface. The channel switch must therefore be a state transition, not a second independent send, and its transaction must revoke the superseded challenge before the backup becomes usable.

The overlap matters.

Short outages are not the only planning case. If neither channel can complete, the terminal branch must already be defined as recovery codes, support-assisted recovery, or another separately assessed factor. No voice, WhatsApp, or RCS channel is available from this integration to fill that gap automatically.

For the order receipt, use a different queue and SLO. Payment settlement should create an idempotent receipt job, and a duplicate delivery attempt must not create a second business event. Email domain authentication, including DKIM, belongs in the readiness checklist. Opens do not belong in an authentication decision because mail privacy features can prevent that telemetry from representing a person's action.

Verify behavior and make rollback boring

Verification starts with deterministic clocks and injected provider failures. Exercise code expiry, concurrent consumption, resend invalidation, duplicate issue requests, HTTP 429 handling, delayed email status, and a worker restart between send acceptance and local persistence. Then run synthetic checks to destinations the organization controls in each important region; do not turn aggregate claims into a US/EU delivery SLO without evidence from the actual destination mix.

Before rollout, define three gates: the verification-success indicator must remain inside its error budget, fallback demand must stay within provisioned capacity, and abuse controls must reject a deliberate burst without blocking the test accounts used for regional probes. Start with a small cohort. Keep the previous authentication path available until both the primary and fallback branches have passed a complete expiry window and the polling backlog has drained.

Rollback should stop new challenge issuance on the new path while continuing to verify challenges that were already issued until they expire. Otherwise a rollback converts a provider problem into guaranteed login failures. Preserve the operation identifiers and state-transition audit trail, but never retain raw OTP values in logs.

The decision rule is narrow: choose hosted SMS OTP first when delivery speed and a smaller authentication ownership boundary dominate; add email only as an explicit recovery path when the team can operate the code lifecycle and accept polling-delayed orchestration. Choose a dedicated verification vendor when that managed boundary matters more than consolidating communications. Choose lower-level transport when building and owning the state machine is a deliberate platform capability.

References

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.