How to Build a Proxy Rotator in Python (That Actually Works)
If you have ever run a web scraper for more than a few hours, you know the pattern. It works great at first. Then the requests start failing. Then your IP gets blocked entirely. The site was not even angry at you, you ju
If you have ever run a web scraper for more than a few hours, you know the pattern. It works great at first. Then the requests start failing. Then your IP gets blocked entirely. The site was not even angry at you, you just looked like a bot because every request came from the same address.
The fix is rotating proxies: send each request (or every few requests) through a different IP address. It sounds complicated, but the core idea fits in about 40 lines of Python. Here is how I built mine.
What you need
- Python 3.8+
-
requestslibrary (pip install requests) - A list of proxies (I will show you where to get free ones)
Step 1: Get a proxy list
You need a pool of proxies to rotate through. Free proxy lists work fine for learning and light jobs. I maintain ProxyNest, a free proxy list that retests its entries continuously, and it has a simple JSON API:
GET https://proxynest.live/api/proxies?limit=50
That returns a list of live proxies with their IP, port, protocol (http/socks4/socks5), and country. For this tutorial you can also paste proxies from any list into a text file, one per line in ip:port format.
Step 2: Build the rotator
The idea is simple. Keep a list of proxies, pick one at random (or round-robin) for each request, and drop the ones that fail so they do not poison your pool.
import random
import requests
import time
PROXY_API = "https://proxynest.live/api/proxies?limit=100"
def fetch_proxies():
resp = requests.get(PROXY_API, timeout=10)
resp.raise_for_status()
proxies = []
for p in resp.json():
url = f"{p['protocol']}://{p['ip']}:{p['port']}"
proxies.append({"http": url, "https": url})
return proxies
class ProxyRotator:
def __init__(self):
self.proxies = fetch_proxies()
self.bad = set()
def get(self):
good = [p for p in self.proxies
if p["http"] not in self.bad]
if not good:
raise RuntimeError("Proxy pool exhausted, refetching...")
return random.choice(good)
def mark_bad(self, proxy):
self.bad.add(proxy["http"])
rotator = ProxyRotator()
That is the whole rotator. Ten lines of real logic.
Step 3: Use it in your scraper
Now wrap your requests so a dead proxy does not kill the run:
def fetch(url, max_retries=5):
for attempt in range(max_retries):
proxy = rotator.get()
try:
resp = requests.get(
url,
proxies=proxy,
timeout=10,
headers={"User-Agent": "Mozilla/5.0 (compatible; my-scraper/1.0)"},
)
resp.raise_for_status()
return resp.text
except requests.RequestException:
rotator.mark_bad(proxy)
time.sleep(1)
raise RuntimeError(f"Failed to fetch {url} after {max_retries} retries")
html = fetch("https://example.com/products")
print(len(html), "characters fetched")
A few things to notice:
-
Dead proxies get removed.
mark_badtakes failing proxies out of the pool instead of retrying them forever. - Random selection beats round-robin for small pools, because it spreads load unevenly and looks less mechanical.
-
The User-Agent matters. Rotating proxies but sending the same default
python-requestsuser agent is like wearing a disguise and then shouting your name. Set a realistic one.
Step 4: Do not be rude
Rotation is not a license to hammer a site. Add delays between requests (time.sleep(1) at minimum), respect robots.txt, and keep your concurrency sane. Most blocks happen because of request rate, not because of proxies. The proxy just buys you a second identity; it does not make aggressive scraping polite.
If you need serious throughput, free proxies will eventually frustrate you. They die fast and they are slow. But for learning, prototypes, and light scheduled jobs, this setup works surprisingly well.
The takeaway
Proxy rotation is one of those things that sounds like infrastructure but is really just a list and a random choice. Fetch a pool, pick one per request, drop the failures, and be polite about your request rate. That is 90 percent of it.
Muhammad Imran works on network tooling and runs ProxyNest, a free directory of public proxies with a live API and a proxy checker that re-verifies any proxy on demand.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.