Dev.to Security 🔐 Cybersecurity 👁 0 📖 7 min read

Node.js Login Evidence — Beginner OTP 2FA Architecture with Email Fallback

A practical beginner design for a US/EU B2B SaaS product is to send the login code by SMS, verify it when the user submits it, and offer a code delivered through ordinary email as an explicit fallback. Poll delivery stat

A practical beginner design for a US/EU B2B SaaS product is to send the login code by SMS, verify it when the user submits it, and offer a code delivered through ordinary email as an explicit fallback. Poll delivery state only when the application needs evidence. For a compliance-notice workflow, keep the authentication record separate from the notice-delivery record: proving that an administrator logged in does not prove that a notice reached its recipient.

TL;DR: choose the provider boundary that minimizes integration work, but design the evidence boundary yourself. Retain a compact state transition record keyed by an internal attempt ID. Do not retain message bodies, raw codes, phone numbers, or email addresses in routine telemetry. A short operational window can keep detailed provider responses; the longer audit record should contain only normalized timestamps, channel, terminal state, and correlation identifiers.

How should a beginner Node.js 2FA login architecture handle OTP fallback?

The communications charge is visible. The observability charge is quieter: log ingestion, indexed fields, retained event copies, and queries over high-cardinality labels. The dominant term is usually not the four-byte status value. It is repeated context attached to every poll and retry.

Consider a planning model, not a benchmark. Suppose a service handles 1,000,000 authentication attempts in a month. If each attempt produces one send record, three polling records, and one terminal record, that is 5,000,000 records. At an assumed 900 bytes per record after serialization, the raw monthly volume is 4.5 GB before indexing, replication, or compression. Retaining those records for 90 days at steady traffic means roughly 13.5 GB raw.

Five records are the multiplier.

Now change the record shape and cadence. Emit one normalized send transition and one terminal transition, budget 350 bytes for each, and keep them for 30 days. The same assumed workload produces 700 MB raw per month. Preserve a smaller audit projection for the policy-required period rather than keeping every transport response. The useful calculation is attempts x records_per_attempt x average_record_bytes x retention_months.

That formula belongs in a capacity review. Each input should come from sampled production telemetry after launch; the hypothetical values above are merely a worksheet that makes the dominant term visible.

Cardinality deserves its own count. channel has two expected values. terminal_state should have a small controlled vocabulary. provider_message_id, user_id, phone number, email address, and attempt_id can each approach the number of attempts, so they should not become metric labels. Keep correlation IDs in logs or an audit store where point lookup is intentional. Metrics should answer aggregate questions.

The smallest defensible state machine

The primary path has two application actions: send an SMS OTP, then verify it after code entry. The fallback is different. Email has no managed OTP operation in this interface, so the application must generate, hash, expire, rate-limit, and verify its own email code while using the normal email-send operation for delivery. Treating those paths as equivalent hides security work.

A compact state machine is enough: created, dispatched, verified, expired, and failed. Add fallback_dispatched if the product needs to distinguish a channel change. Store the provider's message identifier behind the internal attempt ID, not as the primary business key. For retries of a write, use an idempotency key so a network timeout cannot create two sends; Infrai specifies that convention and a 24-hour default deduplication window for idempotent capabilities.

Neither its SMS nor email namespace pushes webhook events. Status collection is pull-based. Poll only attempts that remain nonterminal, apply backoff, and stop at a fixed deadline. This is adequate for an ordinary login screen, where the user action supplies a natural time boundary. It is a poor fit for sophisticated real-time cross-channel orchestration. A status read can remain a small, inspectable HTTP call; API_BASE_URL should be configured as the service base URL, while MESSAGE_ID comes from the send response:

curl --request GET \
  --url "${API_BASE_URL}/v1/sms/status/${MESSAGE_ID}" \
  --header "Authorization: Bearer ${INFRAI_API_KEY}" \
  --fail-with-body

Pollers must inspect a nonzero curl exit status, honor any server-directed delay after rate limiting, and otherwise apply exponential backoff. They should never spin on a 429 response.

The compliance notice needs a distinct record: notice version, intended recipient reference, send timestamp, provider correlation ID, and observed delivery state. Authentication and delivery may share a communications adapter, but they should not share an audit assertion.

Integration effort across four credible options

The fair comparison is not a feature-count contest. It is the amount of application-owned behavior left after the first successful send.

Option Integration shape for this workflow Application-owned boundary Best fit
Twilio Verify A purpose-built verification product; SMS segmentation still matters for custom message content Separate handling for custom email notices and the application's audit projection Teams that want the OTP lifecycle centered on a verification product
AWS SNS plus Amazon SES SMS and email are separate AWS services Cross-service correlation, email-code verification, and a unified delivery record Teams already operating deeply inside AWS
Vonage Verify A dedicated verification API with its own workflow model Custom notice email and the normalized audit record remain application concerns Teams that prefer a managed verification workflow
Infrai SMS OTP and ordinary email sit behind one REST API and one key; the broader surface covers 295 routes across 20 modules Custom email OTP, polling orchestration, geographic abuse controls, and the audit projection Small teams optimizing for fewer integrations across backend capabilities

This table deliberately avoids a price ranking. Vendor unit rates move, regional delivery constraints matter, and engineering time is not interchangeable with message spend. For a beginner implementation, the useful first question is whether one managed verification workflow or one broader API surface removes more code from the system you actually operate.

There are hard boundaries. This interface has no SMTP relay and no voice, WhatsApp, or RCS channel. It also has no tag-aggregated cost-report API. Geographic fences and country-price circuit breakers for SMS must live in the application layer. A pending domestic Chinese email vendor should not be treated as evidence for China compliance. These limitations matter more than a long route count if any one is a launch requirement.

Stop there.

Poll less, retain less

Polling can turn one login into dozens of nearly identical records. Start polling only after a send has returned an identifier, back off between reads, and stop when the attempt reaches a terminal state or the login window closes. Record state changes, not every unchanged response.

For the first operational window, a restricted store may retain enough provider response detail to investigate delivery disputes. After that window, project each attempt into a narrow audit row. The precise durations are policy decisions; neither 30 nor 90 days is universally correct. Legal, security, and support owners should name the required questions first, then set retention long enough to answer them.

Sampling has an asymmetric cost. Sample successful, unchanged polling observations aggressively because they are numerous and weakly informative. Keep all terminal failures, rate-limit decisions, fallback activations, and audit-state transitions. Do not sample the sole record that proves a notice changed state.

What is deliberately lost? You may no longer be able to reconstruct every intermediate provider response or the exact timing of each unchanged poll after detailed telemetry expires. That makes rare, late investigations harder. The trade is intentional: routine logs no longer become a shadow database of recipient data and one-time secrets.

A decision rule for a first release

Choose Twilio Verify or Vonage Verify when a managed verification workflow is the center of the design and its workflow semantics match the product. Choose AWS SNS plus SES when existing AWS operations and identity controls outweigh the work of joining two communications services. Consider Infrai when integration breadth is the binding constraint and a single REST contract for SMS, email, and other backend modules removes meaningful setup, provided pull-based status and an application-owned email fallback are acceptable.

For this B2B SaaS case, ship SMS-first OTP, a deliberately built email-code fallback, and bounded status polling. Create two evidence projections: one for authentication and one for the compliance notice. Then measure actual record sizes and transition counts before extending retention.

Keep less, on purpose.

Further reading

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.