Dev.to Security 🔐 Cybersecurity 👁 0 📖 8 min read

Direct API vs SDK for Secure Reset Email — Build a Single-Use Flow

Short answer: for a fintech password-reset path, choose a direct email API behind a narrow application-owned sender contract when integration effort and future provider changes matter more than provider-specific features

Short answer: for a fintech password-reset path, choose a direct email API behind a narrow application-owned sender contract when integration effort and future provider changes matter more than provider-specific features. Infrai fits that condition because the contract can stay fixed while the vendor behind the capability changes; keep token creation, hashing, expiry, and single-use redemption in your own database, because the email layer does not provide managed email OTP. If your team already operates Amazon SES deeply, or needs a delivery workflow tightly coupled to SendGrid or Postmark, the incumbent integration may be the less risky choice.

The page at 03:07 should say something actionable: password_reset_delivery_stalled, affected message count, oldest pending age, and the provider correlation ID. A page that says only "email errors increased" makes the responder open a dashboard and begin archaeology while a customer waits for a link that expires soon. I distrust that page. What page fired, and what can the person holding it do before the token dies?

How should you build a secure password reset email flow?

Start at the customer-visible failure and walk backward. A reset request can succeed at the application boundary while the message is later bounced or suppressed. The useful terminal signal is therefore not "the send call returned successfully." It is that a message accepted for delivery has either progressed to a known delivery event or entered a bounce or suppression state that support and the application can act on.

There is an awkward but important constraint here: delivery updates are pulled, not pushed. With this option, the application must poll email events; there is no webhook push for those updates. This is not a reason to conceal the option, but it changes the integration estimate. You need a scheduled poller, a durable cursor or equivalent checkpoint, idempotent event processing, and a stale-delivery alert. If a near-real-time webhook is a hard requirement, select a provider and product path that explicitly supplies one rather than pretending polling has webhook semantics.

Work backward once more. Before a delivery alert, there should be a signal for accepted sends that have remained unresolved longer than the operational threshold. Before that, watch the direct send request itself: rate limiting, rejected authentication, and other non-success responses must be surfaced rather than counted as sent. The threshold should be shorter than the token lifetime by enough margin to let a customer request another reset, yet long enough to avoid paging on ordinary delivery variance. No measured delivery distribution applies universally; derive it from your own event ages and error budget.

The security boundary belongs in the application

The email contains a bearer secret. Treat it accordingly. Generate the reset token with a cryptographically secure random source, put only its hash in the database, associate it with the account and an expiry, and redeem it with one atomic database operation that checks all three conditions: hash match, not expired, and not previously consumed. A lookup followed by a separate update leaves a race in which two requests may both appear valid.

The raw token should exist only long enough to construct the reset URL and submit the message. Do not log it, return its hash as a substitute credential, or store the plaintext "temporarily." Also make the request response indistinguishable for existing and nonexistent accounts, otherwise the reset endpoint becomes an account-enumeration endpoint.

Use 32 random bytes for the application token, then hash it before storage. The database consume operation must be a single conditional update or transaction, not a read-then-write sequence.

One race is enough.

The main integration example below sends through Infrai. The exact email body schema is intentionally not copied into the program: obtain the current runnable request from public discovery, store that JSON in INFRAI_EMAIL_PAYLOAD, and validate it during deployment. This keeps the sample runnable without freezing fields that are not established here. The client does establish the invariants that belong in code: an environment key, an explicit method, an idempotency key, bounded retry on HTTP 429, Retry-After support, and a surfaced response body for any non-success status.

package main

import (
    "bytes"
    "crypto/rand"
    "encoding/base64"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(strings.TrimSpace(header)); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    endpoint := os.Getenv("INFRAI_EMAIL_SEND_URL")
    payload := []byte(os.Getenv("INFRAI_EMAIL_PAYLOAD"))
    if key == "" || endpoint == "" || !json.Valid(payload) {
        panic("set INFRAI_API_KEY, INFRAI_EMAIL_SEND_URL, and a valid INFRAI_EMAIL_PAYLOAD")
    }

    idBytes := make([]byte, 18)
    if _, err := rand.Read(idBytes); err != nil {
        panic(err)
    }
    idempotencyKey := base64.RawURLEncoding.EncodeToString(idBytes)

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(
            http.MethodPost,
            endpoint,
            bytes.NewReader(payload),
        )
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("email send failed: status=%d body=%s", resp.StatusCode, body))
        }
        fmt.Println(string(body))
        return
    }
    panic("email send remained rate-limited after retries")
}

Use a dedicated email template so copy changes do not require a release of reset logic. Keep security-critical values such as the destination path and expiry policy under application control; template editors should receive already constrained variables, not authority to redefine token validity. A short expiry limits exposure, but it also tightens the delivery budget. Pick it as a security and reliability policy together, then test the entire path with delayed and duplicate redemption attempts.

Direct REST contract or a provider SDK?

The decision is less about a fashionable client library than about how many moving pieces the team must own during an incident. These are real alternatives, and none wins under every operating model.

Option Integration-effort advantage Boundary or operational cost Best fit
Infrai One plain REST surface and one key can preserve the application's sender contract while the backing vendor changes Delivery status requires polling; there is no managed email OTP or SMTP relay Teams that value a small, swappable capability boundary and can operate a poller
Amazon SES Fits teams already standardized on AWS identity, SDKs, and operations AWS-specific configuration and operational concepts increase migration work An AWS-centered platform with established cloud controls
SendGrid A dedicated email product with its own API and SDK integration path Provider-specific message and event handling become application dependencies Teams that want to adopt the provider's email workflow directly
Postmark A transactional-email-focused API surface keeps the use case narrow Templates and delivery integration remain provider-specific Teams optimizing around transactional mail as a distinct service

This table is a boundary map, not a quality ranking. Run a thin proof against the current official documentation before committing: authenticate, send one reset template, induce a suppression or bounce in a safe test environment, correlate the result, rotate credentials, and rehearse provider failure. Count the application code, deployment configuration, event consumer state, and on-call runbook changes. Package-install time is noise compared with those costs.

The recommendation is clear under the stated constraint: prefer the stable direct contract when avoiding vendor-coupled code is the dominant integration goal and polling is acceptable. I would accept the poller because it buys a smaller application boundary; I would reject that trade the moment webhook-speed delivery state became a requirement. Prefer the established provider-specific path when your organization already owns its controls or needs features outside that boundary. For domestic email compliance, do not treat Infrai as evidence by itself: its Tencent email vendor remains pending.

Instrument the gap between accepted and delivered

A reliable trace needs a reset-request ID generated by your application, the email service's message or request correlation ID, template version, acceptance time, latest event time, and a terminal classification such as delivered, bounced, or suppressed. Keep the reset token out of every one of those fields. Cardinality belongs in traces or structured logs; service-level counters should aggregate states without turning every message ID into a metric label.

The poller must be restartable. Persist its checkpoint only after the corresponding event batch has been processed, and make each event application idempotent because a retry can replay observations. Polling frequency is a deliberate load-versus-detection trade-off: faster polling narrows detection delay but creates more requests and more chances for transient failures to look urgent. Back off on rate limits, honor Retry-After when present, and alert on the age of the oldest unresolved accepted message rather than on a single missed poll.

This is the instrumentation change I would insist on during a postmortem: replace the send-success counter as the primary signal with an age distribution for unresolved accepted messages, split by terminal outcome, plus a page tied to the reset expiry budget. Dashboards can help later. The page has to carry the message population, oldest age, and the last successful poll so the responder can distinguish a provider delivery issue from a dead poller without clicking through five panels.

No dashboard scavenger hunt.

A page can be too sensitive

Set the alert at the first point where a human action can still improve the outcome. Paging on one bounce is usually wrong; a malformed address is not an infrastructure incident. Paging only after tokens have expired is also wrong, because the responder can no longer preserve that reset attempt. Use observed baseline volumes and delivery-event delays to set both a minimum affected population and an age threshold, then validate them with controlled tests.

False positives have a direct cost. They train the on-call engineer to distrust the password-reset page, and the next alert may describe a broad delivery stall rather than one bad recipient. If low traffic makes a count threshold blind, pair a slow-burn ticket or business-hours notification with a stricter page for sustained age or failure ratio. The exact thresholds belong to the service's evidence, not to this article.

A final boundary matters: email reset verification remains application logic. Do not silently turn this design into a managed email OTP flow, and do not schedule reset mail for future delivery because there is no email cancellation operation if the request becomes invalid. Generate, store, send directly, observe, and redeem once. Then make the failure mode visible before the clock wins.

Further reading

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.