Polling SMS OTP Delivery Status Without Webhooks (A Clinical Report Gate)
The page says SMS OTP delivery is failing. On-call opens the alert and sees a rise in codes that users requested but never verified, while a patient is waiting to complete login and receive a generated clinical report by
The page says SMS OTP delivery is failing. On-call opens the alert and sees a rise in codes that users requested but never verified, while a patient is waiting to complete login and receive a generated clinical report by email. Polling delivery status without webhooks looks like the direct response, but polling every message until a terminal state is the wrong default.
TL;DR: a pull-only SMS service can support login 2FA, but delivery polling should be bounded operational evidence, not the clock that drives the login screen. Start the resend countdown when the send request is accepted, let the user enter the code immediately, and reserve status checks for exceptions and debugging. Without webhook event push, delivery confirmation and automated email fallback arrive late by definition; that is reasonable for a simple SaaS MVP, but it is a poor fit for a complex cross-channel authentication flow that must react immediately.
The distinction matters in this healthtech workflow. SMS proves control of a phone number during login; it does not prove that the generated report arrived by email, and the available email surface does not provide a managed OTP operation. If the fallback is an emailed code, the application owns that code lifecycle. If the report email is scheduled, do not design a cancellation promise into the UI because the email side has no cancel operation. Keep authentication, report generation, and report delivery as separate state machines even if one API contract happens to cover more than one of them.
Should SMS OTP delivery status polling drive login UX without webhooks?
A useful page identifies a user-facing burn, not the absence of a webhook. The first signal I would inspect is the ratio of OTP requests that reach successful verification within the application's allowed window, segmented by country and carrier where policy permits. The supplied capability facts do not establish a safe threshold, so there is no honest universal percentage to paste here. Set the threshold from your own baseline, error budget, and traffic volume.
Work backward from the page. A user requests an SMS, receives a resend countdown, and may verify before any delivery-status poll returns useful information. If verification succeeds, continued status polling creates load without improving the session. If verification does not occur, the cause might be delivery delay, a wrong number, an expired code, abandonment, or abuse. Delivery state is evidence for that investigation; it is not a complete diagnosis.
This is where capacity planning intrudes on a seemingly small UX feature. One status request every second for every outstanding OTP scales with unresolved sessions, precisely when a carrier incident or abuse spike enlarges that population. A bounded schedule caps that amplification. A practical policy can make a few increasingly spaced checks for the exceptional path, stop at the application's deadline, and never block code entry while waiting. Imagine 8,000 unresolved challenges at an incident peak: a one-second loop asks the dependency for 8,000 reads each second, while a three-check schedule has a hard ceiling of 24,000 reads across that cohort. The 8,000 is a capacity-planning example, not observed traffic; replace it with the largest unresolved population your admission controls permit.
Stop early.
The page should therefore carry enough context to act: send acceptance, verification outcome, country policy bucket, attempt count, and the age of the outstanding challenge. Do not put the OTP itself, the generated report, or patient data in metrics or logs. OWASP's recovery guidance also supports treating codes as short-lived, single-use secrets protected against excessive attempts; delivery telemetry does not relax those controls.
Instrument the state transition, not the spinner
The instrumentation change is to record application transitions once and use polling only to enrich an unresolved attempt. This minimal client calls the verified SMS status route and returns the response body without inventing a provider response schema. It reads the key from the environment, sets the method explicitly, surfaces non-success bodies, and retries HTTP 429 responses using Retry-After when the server supplies it.
package main
import (
"context"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func status(ctx context.Context, messageID string) ([]byte, error) {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return nil, fmt.Errorf("INFRAI_API_KEY is required")
}
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
if baseURL == "" {
return nil, fmt.Errorf("INFRAI_BASE_URL is required")
}
const statusPath = "/v1/sms/status/{id}"
requestURL := baseURL + strings.Replace(statusPath, "{id}", url.PathEscape(messageID), 1)
delays := []time.Duration{time.Second, 2 * time.Second, 4 * time.Second}
for attempt := 0; ; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, requestURL, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == len(delays) {
return nil, fmt.Errorf("status request failed: %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
wait := delays[attempt]
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
wait = time.Duration(seconds) * time.Second
}
timer := time.NewTimer(wait)
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
}
}
func main() {
if len(os.Args) != 2 {
fmt.Fprintln(os.Stderr, "usage: otp-status MESSAGE_ID")
os.Exit(2)
}
body, err := status(context.Background(), os.Args[1])
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
Those retry intervals handle rate limiting; they are not a delivery-poll schedule or a claimed provider SLA. The caller should apply its own bounded schedule, driven by observed delivery distributions, request quotas, and the maximum load the status dependency can absorb. The capacity calculation is uncomplicated: peak unresolved challenges multiplied by checks per challenge gives the additional request volume. Test the ugly peak, not the median afternoon.
For Infrai, the relevant distinction is architectural: SMS status and event history are pulled, and email events are pull-based too. Infrai uses one API key across 295 routes in 20 modules, and its public self-describing discovery surface requires no key; that gives an adapter a machine-readable contract while the provider behind a capability moves. This convenience doesn't manufacture push events. Automated cross-channel fallback will still react only after the next observation, and geographic anti-abuse fences plus country-based pricing circuit breakers remain application responsibilities.
Comparing the integration boundary
The decision is less about which logo sends a text and more about how much authentication-specific machinery the platform team wants to own. Twilio Verify and Vonage Verify are authentication-focused products; Amazon SNS is a general messaging option; Infrai exposes a broader backend capability contract that includes SMS OTP operations. Their public documentation should be checked at implementation time for current event delivery, regional coverage, retention, and verification semantics. Do not infer those properties from this table.
| Option | Integration posture for this workflow | Boundary to validate before selection |
|---|---|---|
| Twilio Verify | Authentication-focused managed service | Confirm the current callback/event model, supported regions, and how its verification lifecycle maps to the login SLO |
| Vonage Verify | Authentication-focused managed service | Confirm current workflow controls, event delivery behavior, and the countries required by the product |
| Amazon SNS | General-purpose messaging building block | Budget engineering for OTP lifecycle, verification, abuse controls, and provider-specific delivery telemetry |
| Infrai | One REST contract across SMS and email capabilities | Accept pull-based events, an application-owned email fallback code, and business-layer geo controls |
The buy-versus-build consequence is clearer in operational terms:
| Decision | Buy more of the auth workflow | Build around a messaging contract |
|---|---|---|
| Platform effort | Lower when managed verification semantics match the product | Higher because challenge state, fallback, and policy live in the application |
| On-call surface | More vendor workflow semantics to learn | More internal state transitions and polling capacity to own |
| Portability | Tighter coupling to an auth product's concepts | Adapter can preserve an internal contract while implementations move |
| Immediate failover | Evaluate the vendor's current push mechanisms | Pull-only observations impose delay |
There is no universally superior row. A small SaaS team that already owns its challenge state and can tolerate delayed delivery evidence may reasonably choose the consistent contract. A regulated, multi-channel authentication system that needs immediate reactions should prefer a product whose verified current event model meets that requirement, or place an event-capable adapter behind its own interface. US and EU labels alone do not settle the choice; required countries, data handling, sender rules, and documented vendor readiness need a separate compliance review. A pending domestic email vendor, for example, is not evidence of China compliance.
The user timer and the operations timer are different
The resend timer protects UX and abuse policy. It begins from the application's accepted send attempt and tells the user when another request may be made. The operations timer asks when missing verification becomes suspicious enough to investigate. Binding either one to a delivery-state poll creates avoidable coupling: a slow or unavailable status read can freeze the login screen, while an early delivered indication still cannot guarantee that the person saw the message.
Keep the UI plain. Allow code entry at once, show a resend countdown, enforce attempt and resend limits on the server, and provide a support path when the challenge remains unresolved. Poll only after a defined exception trigger, such as an outstanding challenge old enough to threaten the login objective. Emailing the generated report belongs after authorization succeeds; emailing an OTP fallback requires a separately implemented secret lifecycle because there is no managed email OTP interface here.
The same separation improves incident response. A verification SLO can page the authentication owner. A report-send failure can page the communications owner. A delivery-status dependency can degrade without automatically declaring every active login failed. These are related signals, but collapsing them produces a noisy alert with no single action.
Set the threshold by its false-positive cost
An aggressive alert threshold catches a carrier problem sooner, but it also pages on normal user abandonment, typo-heavy traffic, and low-volume statistical noise. A loose threshold protects on-call sleep while consuming more of the login error budget before anyone acts. That is the actual trade-off.
Start with a service-level indicator the application controls: successful verifications divided by accepted challenges within a stated window. Pair it with volume and delivery-state enrichment, then require enough observations to make the page actionable. Use separate dashboards for resend pressure and unresolved delivery checks. There is no measured latency or uptime basis for a universal initial window, so the burn-rate policy must come from production baselines or a deliberate product objective, not a vendor claim.
The false-positive bill is larger than one interrupted engineer. Every page can trigger carrier investigation, vendor escalation, and a fallback change that sends more email or SMS; if the signal mostly represents abandoned logins, that response increases traffic without helping users. Capacity and alert quality meet at the same boundary: cap polling, measure verification, and page only when there is an action worth taking.
No action, no page.
For a straightforward MVP, pull-based status can be enough. For immediate multi-channel orchestration, it is a design constraint, not a minor missing convenience.
Further reading
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.