Dev.to Security 🔐 Cybersecurity 👁 0 📖 4 min read

Node.js Rate Limiting Techniques: Protect APIs from Abuse

Rate limiting is one of the simplest ways to protect a Node.js API from accidental overload, abusive clients, brute-force attempts, and traffic spikes. Instead of allowing every request to reach your application, a rate

Rate limiting is one of the simplest ways to protect a Node.js API from accidental overload, abusive clients, brute-force attempts, and traffic spikes. Instead of allowing every request to reach your application, a rate limiter controls how frequently a client can access a resource.

For production systems, rate limiting is not just a security feature. It also helps keep latency predictable, protects database connections, and prevents one client from consuming resources that should be available to everyone.

Understanding Rate Limiting Techniques in Node.js

The most common rate limiting algorithms are fixed window, sliding window, token bucket, and leaky bucket. A fixed window is easy to implement but can create traffic bursts around window boundaries, while sliding-window approaches provide more accurate control over request frequency.

For many APIs, a token bucket is a practical choice because it allows controlled bursts while maintaining an average request rate. The example below implements a lightweight in-memory token bucket middleware without external dependencies, making the behavior easy to understand and test.

In a real distributed application, an in-memory limiter should usually be replaced or coordinated through Redis or another shared store. Otherwise, requests distributed across multiple Node.js instances can bypass the intended global limit because each process maintains its own state.

The example also demonstrates important operational details: returning HTTP 429, sending Retry-After information, logging accepted and rejected requests, and periodically removing inactive clients. These details make rate limiting easier to observe and maintain in a real service.

const http = require("http");

// Token bucket configuration.
// Each client gets its own bucket with a limited number of tokens.
const MAX_TOKENS = 5;
const REFILL_RATE = 1; // Tokens added every second.
const CLEANUP_INTERVAL = 60 * 1000;

// Store client buckets in memory for this example.
const clients = new Map();

function getClientKey(req) {
  // In production, consider authentication identity instead of IP alone.
  return req.headers["x-client-id"] || req.socket.remoteAddress || "unknown";
}

function createBucket() {
  return {
    tokens: MAX_TOKENS,
    lastRefill: Date.now(),
    lastSeen: Date.now()
  };
}

function refillBucket(bucket) {
  const now = Date.now();
  const elapsedSeconds = (now - bucket.lastRefill) / 1000;
  const tokensToAdd = elapsedSeconds * REFILL_RATE;

  // Never allow the bucket to exceed its maximum capacity.
  bucket.tokens = Math.min(MAX_TOKENS, bucket.tokens + tokensToAdd);
  bucket.lastRefill = now;
}

function rateLimit(req, res) {
  const clientKey = getClientKey(req);
  let bucket = clients.get(clientKey);

  // Create a new bucket when this client appears for the first time.
  if (!bucket) {
    bucket = createBucket();
    clients.set(clientKey, bucket);
    console.log(`[RATE LIMIT] Created bucket for ${clientKey}`);
  }

  bucket.lastSeen = Date.now();
  refillBucket(bucket);

  // Each request consumes exactly one token.
  if (bucket.tokens < 1) {
    const secondsUntilNextToken = Math.ceil((1 - bucket.tokens) / REFILL_RATE);

    console.log(`[RATE LIMIT] REJECTED ${clientKey} - retry in ${secondsUntilNextToken}s`);

    res.statusCode = 429;
    res.setHeader("Content-Type", "application/json");
    res.setHeader("Retry-After", String(secondsUntilNextToken));
    res.end(JSON.stringify({
      error: "Too Many Requests",
      message: "Rate limit exceeded. Please try again later.",
      retryAfter: secondsUntilNextToken
    }));
    return false;
  }

  bucket.tokens -= 1;

  console.log(`[RATE LIMIT] ACCEPTED ${clientKey} - tokens left: ${bucket.tokens.toFixed(2)}`);
  return true;
}

const server = http.createServer((req, res) => {
  console.log(`\n[SERVER] ${req.method} ${req.url}`);

  // Apply rate limiting before executing the expensive API logic.
  if (!rateLimit(req, res)) {
    return;
  }

  // Simulate an API endpoint that performs application work.
  if (req.url === "/api/data") {
    res.statusCode = 200;
    res.setHeader("Content-Type", "application/json");
    res.setHeader("X-RateLimit-Limit", String(MAX_TOKENS));
    res.end(JSON.stringify({
      success: true,
      message: "Request processed successfully",
      timestamp: new Date().toISOString()
    }));
    return;
  }

  res.statusCode = 404;
  res.setHeader("Content-Type", "application/json");
  res.end(JSON.stringify({ error: "Not Found" }));
});

// Remove inactive clients so memory usage does not grow forever.
setInterval(() => {
  const expirationTime = Date.now() - CLEANUP_INTERVAL;

  for (const [clientKey, bucket] of clients.entries()) {
    if (bucket.lastSeen < expirationTime) {
      clients.delete(clientKey);
      console.log(`[CLEANUP] Removed inactive client: ${clientKey}`);
    }
  }
}, CLEANUP_INTERVAL);

const PORT = 3000;

server.listen(PORT, () => {
  console.log(`[SERVER] API running at http://localhost:${PORT}`);
  console.log(`[SERVER] Limit: ${MAX_TOKENS} burst requests`);
  console.log(`[SERVER] Refill rate: ${REFILL_RATE} token per second`);
  console.log("[SERVER] Try requesting /api/data repeatedly to observe 429 responses.");
});

// Gracefully shut down the server when the process receives Ctrl+C.
process.on("SIGINT", () => {
  console.log("\n[SERVER] Shutting down gracefully...");
  server.close(() => {
    console.log("[SERVER] Server stopped.");
    process.exit(0);
  });
});

Conclusion

Choosing the right algorithm depends on the traffic pattern and consistency requirements of your application. Fixed windows are simple, sliding windows provide smoother enforcement, and token buckets are useful when controlled bursts are acceptable.

For production Node.js APIs running on multiple instances, centralized state is usually essential. Redis-backed rate limiting is a common approach because every application instance can enforce the same limits regardless of which server receives the request.

Rate limiting should also be combined with authentication, request validation, monitoring, logging, and sensible endpoint-specific limits. Login, password-reset, payment, and public API endpoints often need different policies rather than one global request limit.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.