Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 2 min read

I slowed my scraper down 10x and it finally stopped getting banned

Faster is not better. I spent two weeks squeezing every millisecond out of my Playwright scraper, and every speedup bought me a new IP ban. The fix was throttling it down to a tenth of the speed β€” bans dropped to zero.

Faster is not better. I spent two weeks squeezing every millisecond out of my Playwright scraper, and every speedup bought me a new IP ban. The fix was throttling it down to a tenth of the speed β€” bans dropped to zero.

The number that made me stop and re-read my logs

I was running a headless Playwright loop over a public catalog, ~120 requests a minute, no delay, no jitter. The result was predictable: a fresh datacenter IP died in under 40 minutes, and I was burning residential proxy traffic on a loop that wasn't even finishing.

The log line that broke the illusion: the scraper wasn't failing on the target site at all. It was failing on the rate-limiter page β€” a 429 that my retry logic treated as a transient error and hammered even harder. I had built a feedback loop that made the ban faster every time it retried.

What I changed, in order

  1. Paced requests instead of bursting them. I capped the loop at 8 requests per minute with a random 2–6 second sleep between hits. Not a fixed sleep(5), but sleep(random.uniform(2, 6)) β€” fixed delays are a fingerprint too.

  2. Stopped retrying on 429. A 429 is not "try again in a second", it's "back off or leave". I added an exponential backoff that starts at 60 seconds and doubles, capped at 10 minutes, and it stops the loop entirely after three consecutive 429s.

  3. Spread the load across sessions, not threads. I had four concurrent browser contexts sharing one IP. Now one context per IP, one page at a time. Throughput per minute went down, total pages completed per hour went up, because nothing was getting killed mid-run anymore.

  4. Dropped the "faster" micro-optimizations. I had parallelized page loads, disabled images, disabled CSS, and used raw HTTP where I could. All of it made the traffic pattern look less human. I reverted to a normal page load with default timeouts.

The hot-take

The bottleneck was never my code's speed. It was my traffic's shape. A scraper that finishes 100 pages and gets banned collected 100 pages. A scraper that finishes 50 pages an hour and stays alive collects 500 by morning. Rate limits are not a bug to work around β€” they're the only signal a site gives you before it silently shadow-bans your IP.

What I'd tell past me

Measure pages completed per day, not requests per second. If your success rate is under 100%, you don't have a speed problem β€” you have a survival problem. Slow the loop down until bans stop, then inch it back up. You'll find the real ceiling is far lower than your code's ceiling, and that's fine.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.