Dev.to Security 🔐 Cybersecurity 👁 0 📖 9 min read

A hundred phones in one pub share one exit IP, so half our rate limiters are not keyed on it

Try this: curl -s https://pub-trivia.app/robots.txt | grep too-many # Disallow: /too-many-requests curl -s https://pub-trivia.app/too-many-requests | grep -o '<meta name="robots"[^>]*>' # <meta name="robots" content=

Try this:

curl -s https://pub-trivia.app/robots.txt | grep too-many
# Disallow: /too-many-requests

curl -s https://pub-trivia.app/too-many-requests | grep -o '<meta name="robots"[^>]*>'
# <meta name="robots" content="noindex, nofollow"/>

A rate limit page that is disallowed in robots.txt and carries a noindex tag. Both, for reasons I will come to, and the second one matters more than the first.

We build quiz software for pubs. The rate limiting story in a normal web app is "key on the IP, pick a number, move on". That advice is wrong for us in a way that took a production incident to see properly, because the defining fact about our traffic is that a hundred people in one room are a hundred clients behind one NAT address.

This post is about the keying, not the game. The keying is the part that generalises.

The limiter everyone writes first is a limit on the room

An IP-keyed limiter treats the IP as a stand-in for "one client". In a venue it is a stand-in for "everyone here". So the limiter does not protect the endpoint from a bad client, it protects the endpoint from the venue, and it fires hardest exactly when the venue is busiest, which is the one moment the product has to work.

We have ten limiters. They use three different keys, and the choice is made per endpoint:

Keyed on Used for
IP the global fallback, joining, the first page load, auth, password reset, the public pricing endpoint
Participant the two endpoints a player's own device calls repeatedly
User ID host actions and checkout session creation

The rule that produces that table: key on the IP when the identity does not exist yet, and key on the identity the moment it does. Joining has to be IP-keyed, because before you join there is nothing else to key on. Everything a joined client does afterwards has an identifier attached, and using the IP there instead is throwing away a better key you already hold.

/**
 * submitAnswer: 10 per 60s per participantId (not IP).
 *
 * Keyed on the actual participant so one player cannot flood the endpoint,
 * while players sharing a WiFi IP are not penalised for each other.
 */
submitAnswer: sliding(10, 60),

Ten per minute per participant is a tight limit, and it can be tight precisely because it is not keyed on something shared. On an IP key the same protection would have to be a hundred times looser to be safe, which means it would protect nothing.

1200 a minute is arithmetic, not generosity

The global limiter in the proxy is the one that has to be IP-keyed, because it runs before anything is known about the request. Its number is the most interesting one in the file:

/**
 * global: 1200 requests per minute per IP.
 *
 * Sized for the worst legitimate case: an entire pub behind one WiFi exit IP
 * with the WebSocket server unreachable, so every phone is on the polling
 * fallback. 100 players x 6 polls/min = 600, plus joins, page loads and
 * answer submissions. The previous 300 was sized for the join burst alone
 * and so 429'd the whole venue precisely when the fallback kicked in.
 */
global: sliding(1200, 60),

That last sentence is the incident. The original 300 was chosen by thinking about the worst case we had imagined, which was everybody joining at once. It was not sized for the worst case that actually happens, which is everybody joining at once and the realtime transport being down, so every client is also on its slower fallback path.

The failure mode compounds in the worst possible direction. The fallback exists so that a venue survives the realtime server being unreachable. An undersized IP-keyed global limit turns that fallback into a denial of service against the venue, so the mitigation fires only in the situation the mitigation was built for.

Two lessons I would take anywhere. Size a shared-key limiter against the sum of the legitimate clients behind the key, not against one of them. And size it for the degraded path, because the degraded path is the expensive one and it is the one that is live when you are already having a bad day.

Tighter where the attempt is cheap

Not everything gets a venue-sized number. The auth limiters are deliberately mean:

/** login / signup / Google OAuth: 10 per 15 minutes per IP. */
auth: sliding(10, 900),

/** forgotPassword / resetPassword: 5 per 15 minutes per IP. */
passwordReset: sliding(5, 900),

Ten login attempts per quarter hour is nothing like 1200 a minute, and the asymmetry is the point: the legitimate ceiling is low, because nobody signs in eleven times in fifteen minutes, while the attacker ceiling is whatever you allow. Password reset is tighter again, because what is being rationed there is not just attempts, it is emails sent to an address somebody else owns.

Checkout is the other tight one, five per ten minutes, keyed on the authenticated user ID. The resource being protected is not our CPU, it is our Stripe account: a spammed endpoint creates junk objects in somebody else's system, and you cannot delete your way out of that quickly.

analytics: false, because it doubled the writes on the hottest path

return new Ratelimit({
    redis,
    limiter: Ratelimit.slidingWindow(requests, `${windowSeconds} s`),
    // No analytics. It costs an extra Upstash write on every limit() call,
    // doubling the command count on the hottest path in the app, for data
    // nothing in this repo ever reads. proxy.ts also never awaits the
    // returned `pending` promise, so on Vercel the write was liable to be
    // torn down mid-flight anyway.
    analytics: false,
    prefix: 'pub_trivia',
})

Two separate reasons, and the second is the one worth passing on. @upstash/ratelimit returns a pending promise for its background work, and you are meant to hand it to waitUntil or await it. The proxy did neither, so on a serverless platform the function could be torn down before the analytics write completed. The result was paying for a write, on the busiest code path in the app, that might not land, to produce a dataset nobody had ever opened.

Default-on telemetry inside a hot-path library is worth a look in any dependency. The cost is per call, and the per-call cost is invisible in review.

No Redis means no limiting, on purpose

function makeNoopLimiter() {
    return { limit: async (_id: string) => ({ success: true as const }) }
}

const redis = makeRedis()

function sliding(requests: number, windowSeconds: number) {
    if (!redis) return makeNoopLimiter()
    // ...
}

Without the Upstash env vars every limit() call returns success, so a fresh clone runs with no Redis account. I know the argument against it: fail-open is how a misconfigured production deploy ends up with no rate limiting at all. I still think it is right, with one condition, which is that the absence has to be loud somewhere else. The limiters are a protection on traffic rather than a feature of the product, and a dev setup that will not boot without a Redis account is a setup people work around in worse ways.

The same shape appears one level up:

export async function currentRequestIp(): Promise<string> {
    try {
        return getRequestIp(await headers())
    } catch {
        return 'no-request-scope'
    }
}

headers() throws when there is no request scope, which is what happens when a server action is called straight from a test. Losing the key degrades to one shared bucket named no-request-scope rather than throwing out of the action being protected. A limiter should never be the reason the thing it guards fails.

That helper lives in its own module rather than beside the limiters, because the limiter file is imported by the proxy, which runs in the middleware runtime where next/headers does not exist. The proxy reads the IP off the request object instead. Two call sites, two runtimes, one helper each.

The limit page had to be exempted from the limit

When a browser trips the global limiter it is redirected to /too-many-requests. Everything else gets a real status code:

const isHtmlRequest = request.headers.get('accept')?.includes('text/html')
if (isHtmlRequest) {
    const url = request.nextUrl.clone()
    url.pathname = RATE_LIMITED_PATH
    url.search = ''
    return NextResponse.redirect(url)
}
return NextResponse.json(
    { error: 'Too many requests. Please slow down.' },
    { status: 429, headers: { 'Retry-After': '60' } }
)

Branching on Accept rather than picking one answer is worth the four lines. A fetch call wants a 429 and a Retry-After it can act on, and a person wants a page that explains itself. Sending the human a bare JSON body is unhelpful, and sending the script a 307 to an HTML page is worse than unhelpful, because a 307 looks like success and the script will happily parse the limit page as data.

The bug in this code is the one I would least have predicted:

The error page itself is exempt. It used to be limited like any other
path, so a browser following the redirect was limited again, redirected
to the same URL, and looped: the user saw ERR_TOO_MANY_REDIRECTS instead
of the page, and every hop burned another token so the window never drained.

A rate limit page served through the rate limiter is a redirect loop, and the loop is self-sustaining. Each hop spends a token, so the window that the user is waiting to drain is being topped up by their own browser retrying. The symptom is not "limited", it is a Chrome error page, and the fix is the one-line exemption that now guards the check:

if (request.nextUrl.pathname !== RATE_LIMITED_PATH) {

Generalising: any time a protection redirects to a page, check whether the protection applies to that page. The same shape shows up with auth gates on sign-in pages and with maintenance pages behind a health check.

The catch block around the whole thing is deliberate too:

} catch {
    // Redis unavailable, fail open so the app stays up
}

Fail open, for the same reason the missing-Redis case returns success. If Upstash is having an outage, an unprotected app is a much better outcome than an app that is down.

A 200 page whose body says 429

The page answers with an HTTP 200. The body says 429, the status line does not, and that gap is why it carries a noindex tag:

export const metadata = pageMetadata({
    path: "/too-many-requests",
    title: "Too many requests",
    // An HTTP 200 page whose body says 429. Left indexable it can be
    // recorded as the content of whatever real URL the crawler asked for.
    noIndex: true,
})

It is in the robots.txt disallow list as well, and those two are not redundant in the way they look. A disallowed page is never fetched, so the noindex tag on it is never read, which means on the path robots.txt covers the tag does nothing. What the tag guards is the case where this content arrives somewhere else: a crawler mid-crawl of a real page, redirected here, left with "Slow down" as the best available answer for the URL it asked about.

The honest version, then, is that the noindex is cheap insurance against a class of mistake rather than a necessary partner to the robots rule. It becomes load bearing the moment that redirect is changed to a rewrite, because a rewrite throws the status code away and the tag is the only thing left.

Check it

pub-trivia.app/too-many-requests is the page, and it tells you HTTP 429 in a monospace line while answering with a 200:

curl -s -o /dev/null -w "%{http_code}\n" https://pub-trivia.app/too-many-requests
# 200

View source and the noindex tag is in the head, which is the belt-and-braces argument above made checkable.

The limiter you are most likely to meet by accident is the public pricing endpoint at 20 a minute, behind pub-trivia.app/pricing. Reloading it twenty-odd times in a minute is a rude thing to do to somebody's site, so do not, but it is the one place the numbers in this post are reachable from outside.

The venue-sized ones are reachable only by filling a room with phones, which is what the free tier is for. It needs no card, and it is the arrangement all of these numbers were chosen around.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.