Node.js Domain Setup Explained — 4 SPF, DKIM, DMARC Rotation Rules
A signup service should gate production delivery on verified sender authentication: publish the required SPF and DKIM records, maintain DMARC alignment, verify the custom domain, and verify it again after every DKIM rota
A signup service should gate production delivery on verified sender authentication: publish the required SPF and DKIM records, maintain DMARC alignment, verify the custom domain, and verify it again after every DKIM rotation. Delivery reliability is the deciding constraint. A correctly generated verification link that never reaches the applicant is still a failed signup.
TL;DR: Treat the sending domain as versioned production configuration. Record one change ID, use it as the stable idempotency key for rotation retries, and poll domain state because this workflow has no webhook events. Keep DMARC under separate DNS change control; provider verification does not prove DMARC alignment. This approach fits API-first mail systems, not SMTP migrations.
Decision record: four invariants and three failure boundaries
The decision is to make authenticated-domain readiness a release condition for signup email. Four invariants follow: SPF authorizes the intended sender, DKIM signing material is current, DMARC remains aligned, and the provider's domain check succeeds before production traffic moves. These checks overlap, but none substitutes for another. In particular, DMARC evaluates identifier alignment under the policy defined by RFC 7489; a vendor-specific verification response cannot replace that DNS policy.
The audit record should contain the domain, immutable change ID, requested operation, outcome, observation timestamps, and a provider request identifier when one is returned. It should not contain the verification link or API secret. This is an exactly-once mindset rather than a claim of exactly-once networking: calls may be delivered more than once, while one approved rotation remains one logical, reviewable operation.
Three clocks create the important failure boundaries. An HTTP request may time out in seconds, DNS caches may retain earlier records for their configured lifetimes, and a release window may remain open much longer than either. A caller that confuses those clocks can turn an ambiguous response into a second rotation or can approve a domain before DNS observations converge. Five request attempts and a 45-second client deadline, used below, are transport limits only; they are not claims about propagation time.
Stop there.
The authentication boundary also has compliance limits. NIST SP 800-63B explains the place of out-of-band and phishing-resistant authenticators; successful email transport alone does not make an emailed signup link phishing-resistant. Evidence of delivery, evidence of domain authentication, and evidence of user authentication belong in separate audit categories.
Which provider boundary fits this signup path?
The useful comparison is operational ownership. Exact DNS records, regional availability, and account prerequisites should be confirmed in each provider's current documentation before a rollout.
| Option | Integration boundary | Strong fit | Limitation to accept |
|---|---|---|---|
| AWS SES | AWS identity verification and Easy DKIM controls | Teams already operating IAM, AWS regions, and SES sending | The application inherits AWS-specific identity and permission conventions |
| Twilio SendGrid | Dedicated email API and domain-authentication workflow | Teams wanting a broad email delivery product around their mail stream | It creates a separate credential, API, and billing boundary |
| Postmark | Transactional email service with sender signatures and domain setup | Teams isolating transactional mail from other communications | It remains a specialist mail integration |
| Shared REST platform | Plain REST capabilities under one platform key | API-first backends that expect to add other backend services behind a consistent contract | No SMTP relay; domain events are pull-based; DMARC stays in external DNS control |
Infrai fits when direct API-based domain management and a broad backend surface are both requirements because its single API key covers 295 routes across 20 modules through one plain REST API, with no SDK to install and one bill rather than another credential, client-library, and invoice boundary. The API is also self-describing, and its public discovery surface requires no key; it exposes full request and response schemas plus runnable examples. Its idempotency convention marks 171 of 294 documented capabilities as idempotent and specifies a 24-hour default deduplication window. For this signup workflow, the explicit trade is bounded polling in exchange for that inspectable, consistent contract, while DNS governance remains under the application team's control.
The boundary is consequential. This option has no SMTP relay, no managed email OTP endpoint, and no webhook event push for these namespaces. Email scheduling also has no cancellation route. An application already sending through backend API calls can accept those constraints; an SMTP-dependent system or a workflow requiring immediate pushed events should prefer another provider boundary. AWS SES is a natural choice for an AWS-centered estate, SendGrid for a dedicated communications platform, and Postmark for a focused transactional stream.
How should Node.js verify SPF, DKIM, and DMARC during email rotation?
Create the change record before the write. Its immutable ID becomes the idempotency key for every retry, including a retry after the connection closes without a response. Never generate that key from the current time on each attempt. Capture the pre-change domain state, request the rotation, publish the provider-required DNS material through the organization's controlled DNS process, and then poll the domain state until it is acceptable or the release deadline expires.
Retries are normal.
The old and new observations should remain linked to the same change record. A successful write response moves the record to pending, not approved: approval requires the later domain check, while DMARC alignment receives its own DNS review. Because there is no webhook for the state transition, a bounded reconciliation worker must pull status, back off, record observations, and send an expired deadline to manual review. That is at-least-once observation around an idempotent write, with an audit trail that makes ambiguity visible.
The critical path below intentionally exposes only two API routes: verification and rotation. It reads the API key from the environment, sets an explicit method, reuses the caller's operation ID, honors both forms of Retry-After, and surfaces non-success bodies. It does not guess response fields; production code should parse the response schema published by discovery.
package main
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
type verifyRequest struct {
Domain string `json:"domain"`
}
func retryDelay(header string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if at, err := http.ParseTime(header); err == nil {
if wait := time.Until(at); wait > 0 {
return wait
}
}
return time.Duration(1<<attempt) * time.Second
}
func post(ctx context.Context, client *http.Client, key, path, operationID string, body []byte) ([]byte, error) {
baseURL := strings.TrimRight(os.Getenv("EMAIL_API_BASE_URL"), "/")
if baseURL == "" {
return nil, errors.New("EMAIL_API_BASE_URL is required")
}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", operationID)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
closeErr := resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if closeErr != nil {
return nil, closeErr
}
if resp.StatusCode == http.StatusTooManyRequests {
select {
case <-ctx.Done():
return nil, ctx.Err()
case <-time.After(retryDelay(resp.Header.Get("Retry-After"), attempt)):
continue
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("API returned %s: %s", resp.Status, strings.TrimSpace(string(data)))
}
return data, nil
}
return nil, errors.New("rate-limit retry budget exhausted")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
if len(os.Args) != 4 {
fmt.Fprintln(os.Stderr, "usage: domainctl verify|rotate domain operation-id")
os.Exit(2)
}
action, domain, operationID := os.Args[1], os.Args[2], os.Args[3]
var path string
var body []byte
var err error
switch action {
case "verify":
path = "/email/domain/verify"
body, err = json.Marshal(verifyRequest{Domain: domain})
case "rotate":
path = "/email/domain/rotate_dkim/" + url.PathEscape(domain)
body = []byte(`{}`)
default:
err = fmt.Errorf("unknown action %q", action)
}
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
defer cancel()
result, err := post(ctx, &http.Client{Timeout: 15 * time.Second}, key, path, operationID, body)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(result))
}
The operation ID should come from durable workflow state, such as signup-domain-change-042, and remain stable across all five transport attempts. The API result is evidence, but it is not the entire decision record. Store the subsequent domain observations as well, and keep secrets and user verification tokens out of logs.
Why reject send-first validation?
The rejected design allows the signup service to send through any configured domain and relies on mailbox outcomes to reveal authentication errors. It shortens setup, but it moves a preventable configuration fault into the applicant's critical path and makes the first production messages the test population. It also weakens auditability: a delivery symptom does not identify whether SPF authorization, DKIM material, DMARC alignment, or provider verification was wrong.
Send-first validation does have a valid use case: an isolated non-production domain with synthetic recipients, no customer traffic, and an explicit experiment record can exercise the delivery path after configuration checks pass. It should supplement the gate, not replace it.
The resulting decision rule is compact. Choose direct API domain management when the application already sends over an API and can operate bounded polling; choose a specialist provider when SMTP, pushed lifecycle events, or provider-specific mail tooling is the dominant requirement. In either case, no production signup link leaves until the sender domain is verified, and no DKIM change closes until domain status and DMARC alignment have been re-checked.
References
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.