Dev.to Security 🔐 Cybersecurity 👁 1 📖 7 min read

Every URL our server fetches was typed by somebody else, so fetch() was never an option

Nakodo reads websites it does not own. Three different places in the app do it: the brand's own site, pasted into onboarding, to write the search brief a creator's linked site, looking for the business email they publi

Nakodo reads websites it does not own. Three different places in the app do it:

  • the brand's own site, pasted into onboarding, to write the search brief
  • a creator's linked site, looking for the business email they published
  • public TikTok profile pages

Every one of those URLs comes from outside. A brand types its domain. A creator puts a link in a bio. That makes all of them the same class of input as a form field, and it makes the plain global fetch() the wrong tool, for a reason that is not obvious until you go looking for it.

The hole in check-then-fetch

The naive version of SSRF protection is: resolve the hostname, check the address is public, then fetch.

const { address } = await dns.promises.lookup(host);
if (isPrivate(address)) throw new Error("no");
await fetch(url); // resolves again, independently

Two lookups. The attacker controls the DNS, so the second one can answer differently from the first. Point a record at 1.1.1.1 with a one second TTL, let the check pass, and have the real connection go to 169.254.169.254 or 127.0.0.1. The window is small and entirely free to attack, because the request can be retried as often as you like.

fetch() gives you nowhere to stand here. It does not let you say which address to connect to, and it does not let you inspect the one it chose. Node's http.request does, through the lookup option:

const publicLookup: LookupFunction = (hostname, options, callback) => {
  dnsLookup(hostname, { ...options, all: true }, (err, addresses) => {
    if (err) return callback(err, "", 0);
    if (addresses.length === 0 || addresses.some((a) => isBlockedAddress(a.address))) {
      const blocked = new BlockedUrlError(`Refusing to fetch ${hostname}: it resolves to a private address`);
      return callback(blocked, "", 0);
    }
    if (options.all) return callback(null, addresses);
    callback(null, addresses[0].address, addresses[0].family);
  });
};

That is the whole point of writing the fetcher on http/https rather than fetch. There is one resolution, the hook sees every address it returned, and the socket connects to the addresses the hook handed back. There is no second lookup to poison. Note all: true and some(): a host with one public and one private address is refused outright, because otherwise which one you connect to is luck.

Two block lists, not one

const BLOCKED_V4 = new BlockList();
const BLOCKED_V6 = new BlockList();

This looked like duplication until it did not. Node's BlockList matches an IPv4 address against IPv6 rules in its mapped form. The v6 list blocks ::ffff:0:0/96, the entire IPv4-mapped range. Put both sets of rules in one list and every IPv4 address on earth is refused, including 1.1.1.1. Two lists, and isBlockedAddress picks by family:

export function isBlockedAddress(ip: string): boolean {
  const family = isIP(ip);
  if (family === 4) return BLOCKED_V4.check(ip, "ipv4");
  if (family === 6) return BLOCKED_V6.check(ip, "ipv6");
  return true;
}

The last line is the one I would like to recommend hardest. Anything that is not an IP address is blocked. Not "passed through for someone else to handle". A function that answers "is this address safe" and is handed something that is not an address has only one honest answer.

The v4 list is the ranges you would expect, plus carrier-grade NAT (100.64.0.0/10), the documentation and benchmark ranges, and 192.0.0.0/24. The v6 list is more interesting, because the attacks there are about carrying a v4 address inside a v6 one:

["::", 96],            // unspecified, loopback, IPv4-compatible
["::ffff:0:0", 96],    // IPv4-mapped
["64:ff9b::", 96],     // NAT64
["64:ff9b:1::", 48],
["100::", 64],         // discard
["2001::", 32],        // Teredo
["2001:10::", 28],     // ORCHID
["2001:db8::", 32],    // documentation
["2002::", 16],        // 6to4
["fc00::", 7],         // unique local
["fe80::", 10],        // link-local
["fec0::", 10],        // site-local, deprecated
["ff00::", 8],         // multicast

The tests are a list of addresses rather than a description of behaviour, which is the right shape for this:

for (const ip of [
  "169.254.169.254", "127.0.0.1", "100.64.0.1", "198.19.255.255",
  "::1", "::ffff:127.0.0.1", "::ffff:7f00:1",
  "64:ff9b::a00:1", "2002:7f00:1::", "2001:0:4136:e378::1",
]) assert.equal(isBlockedAddress(ip), true, ip);

for (const ip of ["1.1.1.1", "8.8.8.8", "172.32.0.1", "198.20.0.1", "2606:4700:4700::1111"]) {
  assert.equal(isBlockedAddress(ip), false, ip);
}

::ffff:127.0.0.1 and ::ffff:7f00:1 are the same address written two ways. 2002:7f00:1:: is 6to4 for 127.0.0.1. 2001:0:4136:e378::1 is Teredo. 172.32.0.1 and 198.20.0.1 are in the allow list on purpose: they are one step outside 172.16.0.0/12 and 198.18.0.0/15, and they are how you find out whether your prefix lengths are right.

Three refusals before DNS is involved

function assertFetchable(u: URL): void {
  if (u.protocol !== "https:" && u.protocol !== "http:") throw new FetchError(`Unsupported URL ${u}`);
  if (u.port && u.port !== "80" && u.port !== "443") throw new BlockedUrlError(`Refusing to fetch port ${u.port}`);
  const host = u.hostname.replace(/^\[|\]$/g, "");
  if (isIP(host)) throw new BlockedUrlError(`Refusing to fetch IP address ${host}`);
}

No file:, no gopher:, no data:. No ports other than 80 and 443, which closes the "scan the internal network by URL" game even for hosts that resolve publicly. And a URL whose host is a literal IP is refused whether or not the IP is public, because no business and no creator publishes their website as an IP address, so the only people who would hand us one are not customers.

Redirects are each a new URL

export async function fetchUrl(url: string): Promise<Fetched> {
  let current = url;
  for (let hop = 0; hop <= MAX_REDIRECTS; hop++) {
    const u = new URL(current);
    assertFetchable(u);
    const res = await get(u);
    const location = res.headers.location;
    if (res.status >= 300 && res.status < 400 && location) {
      current = new URL(location, current).toString();
      continue;
    }
    // ...
  }
  throw new FetchError(`Too many redirects for ${url}`);
}

redirect: "follow" validates the URL you passed and then goes wherever it is sent. A public page that 302s to http://127.0.0.1:6379 defeats every check you did on the first address. So the loop is manual, five hops maximum, and every hop goes back through assertFetchable and the lookup hook. The get() helper deliberately does not follow anything: a 3xx resolves with its headers and an empty body.

Two caps and a thing Postgres cannot store

A 10 second timeout through AbortSignal.timeout, and a 2 MB body cap counted after decompression:

const stream = decompress(res);
stream.on("data", (chunk: Buffer) => {
  chunks.push(chunk);
  size += chunk.length;
  if (size >= MAX_BYTES) {
    finish();
    req.destroy();
  }
});

Counting compressed bytes would be the bug worth having a laugh about later: a one kilobyte gzip response can expand to gigabytes. We decompress through zlib ourselves, count what comes out, and resolve with what we have as soon as the cap is reached rather than erroring, because 2 MB of someone's home page is plenty.

The last one is not a security problem, it is a data problem that can take out a whole batch of writes:

const dropNul = (s: string) => s.replaceAll("\0", "");
const noNul = (_key: string, v: unknown) => (typeof v === "string" ? dropNul(v) : v);
export const parseJson = (text: string): unknown => JSON.parse(text, noNul);

Postgres text and jsonb cannot hold a NUL character. Real pages and real API fields do contain them, usually where someone has pasted UTF-16 text into a profile description. One NUL anywhere in a row fails the insert, and since these fetches are written in batches, one creator's bio can take the whole batch with it. Stripping on the way in, both from decoded HTML and through a JSON.parse reviver, is the cheapest possible fix:

assert.deepEqual(parseJson('{"description":"T\\u0000e\\u0000s\\u0000t","n":1,"tags":["a\\u0000b"]}'), {
  description: "Test", n: 1, tags: ["ab"],
});

Being a guest

The fetcher identifies itself, and the identification is assembled from the brand constant so it cannot drift from the product name:

const BOT_NAME = `${BRAND.name}Bot`;
const USER_AGENT = `Mozilla/5.0 (compatible; ${BOT_NAME}/0.1; +site research for influencer campaigns)`;

Before reading a page that belongs to a creator or a business, robots.txt is consulted for that bot name, with the answer cached per origin for the life of the process:

export async function allowedByRobots(url: string): Promise<boolean> {
  const origin = new URL(url).origin;
  let robots = robotsCache.get(origin);
  if (!robots) {
    if (robotsCache.size > 1000) robotsCache.clear();
    const robotsUrl = `${origin}/robots.txt`;
    robots = tryFetch(robotsUrl).then((r) => (r?.status === 200 ? robotsParser(robotsUrl, r.body) : null));
    robotsCache.set(origin, robots);
  }
  return (await robots)?.isAllowed(url, BOT_NAME) ?? true;
}

The cache stores the promise, not the result, so ten pages from one origin fetch robots.txt once rather than ten times. No robots.txt means allowed, which is what the standard says. And the robots fetch itself goes through the same guarded path as everything else, because robots.txt is a URL somebody else controls too.

See it work

The brief end of this is visible in the product. How Nakodo works describes what happens to a website address after you paste it, and what is never collected. Finding a YouTuber's email is the rule the contact reader implements: read what the creator published, on their own pages, and never guess an address.

If you want to watch the first half happen to a site you own, sign in and paste it into a new campaign. The brief that comes back is built from pages this fetcher was allowed to read.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.