Password Reset Email Deliverability: 6 Custom-Domain DKIM and Bounce Controls
Use an application-owned template contract, a dedicated authenticated sending subdomain, and a durable delivery ledger for password resets and marketplace compliance notices. Treat the email API as a replaceable transpor
Use an application-owned template contract, a dedicated authenticated sending subdomain, and a durable delivery ledger for password resets and marketplace compliance notices. Treat the email API as a replaceable transport. Six controls matter: domain alignment, immutable template versions, idempotent requests, classified bounce events, scoped suppression, and region-aware evidence retention. A provider dashboard is useful for diagnosis, but it is not the system of record.
This decision rule protects two different obligations. A password reset must arrive quickly without becoming a replayable credential, while a seller compliance notice needs evidence of what was requested, rendered, accepted, and later rejected or delivered. Neither obligation is satisfied by an API returning 202. SMTP acceptance is a handoff, not proof that a mailbox received or displayed a message.
What should a custom-domain password reset email setup prove?
The application team should own the semantic contract: template key, version, required variables, permitted locales, subject intent, and the policy event that authorized the send. The transport may render a synchronized copy, but a provider-side template identifier must not be the only description retained in the audit trail. Otherwise a console edit can change the message without changing application code, and an old event becomes impossible to reconstruct.
Ownership comes first.
There is a real trade-off. Provider-managed editing can shorten a compliance team's publishing path and may offer previews. Repository-managed source gives review history, deterministic tests, and an easier provider migration. For a marketplace notice, keep reviewed source and a content digest in the application boundary, then deploy an immutable version to the transport. Human approval and machine delivery remain separate steps.
The ledger should record the business event ID, recipient reference, template version, content digest, locale, policy basis, send attempt, transport message ID, timestamps, and normalized outcome. Do not store a reset token or an unnecessary rendered body in that ledger. GDPR Article 5 requires data minimization and storage limitation; define retention by evidence need instead of keeping webhook payloads forever.
Build the six-control path
Use a subdomain such as notify.example.com so operational mail has an explicit identity and changes do not casually affect unrelated corporate mail. Publish SPF for authorized senders, sign with DKIM, and establish DMARC alignment and reporting. Google documents SPF or DKIM requirements for all senders to personal Gmail accounts and stronger SPF, DKIM, and DMARC requirements for bulk senders. Even if reset volume is small, aligned authentication is the sensible baseline because forwarding, delegation, and future volume changes are easier to reason about.
Keep the visible From identity stable. Rotate DKIM selectors by overlapping old and new keys, verify DNS from outside the deployment network, and remove an old key only after its signed traffic has aged out. DNS existence alone is a weak deployment check; verify that a received test message passes authentication and alignment.
DNS is not the receipt.
The enqueue boundary needs idempotency. A retry after an application timeout must refer to the same logical send, not create a second password-reset message or duplicate a legally significant notice. This compact Go shape keeps the business key independent of any transport:
package notice
import (
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
)
type Request struct {
EventID string
RecipientRef string
TemplateKey string
TemplateVersion string
Locale string
}
type Transport interface {
Send(ctx context.Context, idempotencyKey string, req Request) (messageID string, err error)
}
func IdempotencyKey(r Request) string {
material := fmt.Sprintf("%s\x00%s\x00%s\x00%s",
r.EventID, r.RecipientRef, r.TemplateKey, r.TemplateVersion)
sum := sha256.Sum256([]byte(material))
return hex.EncodeToString(sum[:])
}
Persist that key before calling the transport, and enforce a unique constraint around it. Record attempts separately from the logical message. This matters during an ambiguous timeout: the worker can reconcile the attempt using the same key rather than guessing whether another send is harmless. It isn't.
Bounce handling is an event-processing problem, not a boolean webhook. Authenticate webhook delivery using the transport's documented signing scheme, retain the raw event only as long as policy permits, and make event ingestion idempotent because retries and reordering are normal properties to design for. Map provider-specific values into a small internal state model while preserving the original status code for investigation. RFC 3463 distinguishes persistent permanent failures (5.X.X) from persistent transient failures (4.X.X); that distinction should drive policy.
A hard recipient failure can suppress that exact address for future nonessential sends. A temporary mailbox or routing failure belongs in a bounded retry path with jitter and an expiry, not a permanent global block. Security mail needs a deliberate exception policy: repeatedly sending resets to an address known to reject mail damages reputation, but silently dropping the request leaves the user stranded. Surface an alternate recovery path without revealing whether an account exists.
The dangerous sequence is ordinary: a worker sends, times out before recording the transport response, and retries while the first request is still in flight. Later, two webhook deliveries arrive out of order. Without a stable business key, an append-only attempt record, and monotonic outcome rules, the system can send two valid reset links and finish with the older event displayed as current. The same defect is worse for a marketplace compliance notice because the audit view can show the wrong template version or timestamp. Idempotency at enqueue, transport, and event ingestion addresses three separate duplication boundaries; one idempotency header cannot cover all three.
How do US and EU paths change the design?
They change data handling more than SMTP mechanics. DKIM, SPF, DMARC, MIME, and SMTP status semantics do not acquire regional variants merely because the recipient is in the EU. The questions are where recipient data, event payloads, logs, backups, and support access are processed; which entities act as processors; how long evidence is retained; and how deletion or legal-hold rules interact with the audit ledger.
Make region a routing attribute resolved from an authoritative account policy, not an email suffix. Keep the same template contract and outcome vocabulary in both regions so operators can compare failure rates, while configuring region-specific transport endpoints and storage where the organization's legal assessment requires them. GDPR Article 32 calls for security appropriate to risk, including measures concerning confidentiality, integrity, availability, resilience, restoration, and regular testing. Encryption and access control are necessary, but the runbook also needs restoration tests and evidence that webhook processing can recover.
Be precise about the audit claim. A transport acceptance event proves that a transport accepted a request. A delivery event usually reflects downstream SMTP acceptance. It does not prove that a human read the notice. Store each state with its source and timestamp rather than collapsing all of them into sent=true.
Acceptance isn't delivery.
Verify before shifting traffic
Start with DNS and message-level checks. Query the published records from independent resolvers, send to controlled mailboxes in the regions in scope, and inspect the received authentication results. Test the actual immutable template version with long names, missing optional fields, every supported locale, plain-text fallback, and an expired reset link. The reset URL should carry a single-use, short-lived secret; logs and ledger records should contain a redacted correlation identifier instead.
Then exercise failures. Use controlled test recipients or a standards-based test harness to produce transient and permanent outcomes. Confirm that duplicate events do not duplicate state transitions, an out-of-order delivery event cannot erase a later permanent failure, and the suppression scope matches the reason. Measure queue age, oldest unprocessed webhook, attempts by normalized outcome, unknown-event rate, and the gap between accepted requests and terminal outcomes. Alert on missing signals as well as explicit failures.
Quiet can mean stuck.
A deploy gate should compare the content digest and template version in the application manifest with the version installed in each region. Send a small canary cohort first, then widen traffic while watching authentication failures, bounce classes, queue age, and event lag. Compliance reviewers should approve content before promotion; operators should approve transport health.
Roll back without losing evidence
Rollback should switch new jobs to the last approved template version while leaving existing ledger rows immutable. Never rewrite an earlier record to make it look as though the previous content was never sent. Append the rollback decision, actor, reason, and replacement version.
If authentication fails after a DNS or key change, pause new sends for the affected identity, keep jobs durably queued, and restore the last verified signing configuration. Do not fail over to an unrelated From domain just to drain the queue; that changes identity, alignment, reputation, and the evidence a recipient sees. Resume with a canary after message-level verification.
Transport failover deserves the same caution. A second transport must already have aligned authentication, synchronized immutable templates, equivalent webhook verification, tested suppression import, and a region-approved data path. Otherwise failover converts one visible outage into duplicate messages or an audit gap. The safer rollback can be a queue pause with explicit recovery objectives.
The selection decision follows from this runbook. Choose a transport whose boundaries let the team prove these six controls, export event data, verify callbacks, isolate regions as required, and preserve application ownership of message meaning. Evaluate that contract with failure drills, not a feature-grid score.
References
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.