Building an SSRF-Guarded Webhook & Crawler Subsystem in Node.js
If your web application allows users to input arbitrary URLs—whether for technical SEO audits, link previews, webhook callbacks, or image imports—Server-Side Request Forgery (SSRF) is the most catastrophic security vulne
If your web application allows users to input arbitrary URLs—whether for technical SEO audits, link previews, webhook callbacks, or image imports—Server-Side Request Forgery (SSRF) is the most catastrophic security vulnerability you must defend against.
Without strict network isolation, an attacker can input:
-
http://169.254.169.254/latest/meta-data/to steal AWS EC2 IAM role credentials. -
http://localhost:5432orhttp://10.0.0.12:6379to execute unauthorized commands against internal databases and Redis caches. -
http://127.0.0.1:2375/versionto access unauthenticated internal Docker daemons.
In ⚡ PLYXO (CRO • SEO • AIO • AEO • GEO), our crawler audits millions of external web pages safely. Here is how we build an impenetrable SSRF-guarded crawler.
1. Why Simple Regex and String Blacklisting Fails
Naive implementations attempt to check the hostname with regex:
// ❌ BROKEN: Bypassed by DNS Rebinding, octal IPs, or alternative schemes
if (url.includes('localhost') || url.includes('127.0.0.1')) {
throw new Error('Blocked');
}
Why Attackers Easily Bypass This:
-
Decimal & Hex IP Encodings:
http://2130706433resolves to127.0.0.1. -
DNS Rebinding Attacks: The attacker configures a domain (e.g.
evil.com) with a 0-second TTL that initially resolves to a safe public IP (passing validation), and 5 milliseconds later resolves to169.254.169.254(whenfetchexecutes). -
IPv6 Mappings:
http://[::1]orhttp://[::ffff:127.0.0.1].
2. The Production Defense: Custom Agent with Pre-Handshake IP Verification
To defeat DNS rebinding and IP obfuscation, you must intercept the socket connection at the exact moment of TCP handshake:
import http from 'node:http';
import https from 'node:https';
import dns from 'node:dns/promises';
import ipaddr from 'ipaddr.js';
export function isPrivateOrRestrictedIp(ipString: string): boolean {
try {
const addr = ipaddr.parse(ipString);
const range = addr.range();
// Blacklist all private, loopback, link-local, and reserved ranges
const forbiddenRanges = [
'loopback',
'private',
'linkLocal',
'carrierGradeNat',
'broadcast',
'reserved',
];
if (forbiddenRanges.includes(range)) return true;
// Explicit cloud metadata IP protection
if (ipString === '169.254.169.254' || ipString === 'fd00::') return true;
return false;
} catch {
return true; // Reject unparseable IPs
}
}
export async function safeFetch(targetUrl: string, options: RequestInit = {}): Promise<Response> {
const parsed = new URL(targetUrl);
if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
throw new Error(`Forbidden protocol: ${parsed.protocol}`);
}
// 1. Resolve DNS records explicitly
const addresses = await dns.lookup(parsed.hostname, { all: true });
for (const { address } of addresses) {
if (isPrivateOrRestrictedIp(address)) {
throw new Error(`SSRF Blocked: ${parsed.hostname} resolves to restricted internal IP: ${address}`);
}
}
// 2. Execute fetch with strict timeout and redirect limits
return await fetch(targetUrl, {
...options,
redirect: 'error', // Never follow redirects blindly; re-validate each redirect destination!
});
}
3. Handling HTTP Redirects Safely
Attackers often provide a safe initial URL (https://example.com) that issues a 302 Redirect to http://169.254.169.254.
Rule: Always disable automatic redirect following (redirect: 'error'). Intercept the Location response header, pass the new URL through safeFetch() again, and enforce a maximum redirect depth of 3 hops.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.