How I Built an Autonomous B2B Directory Scraping & Enrichment Pipeline with n8n and Python
How I Built an Autonomous B2B Directory Scraping & Enrichment Pipeline with n8n and Python Modern web scraping for B2B intelligence has shifted. Three years ago, a static requests session combined with BeautifulSoup wa
How I Built an Autonomous B2B Directory Scraping & Enrichment Pipeline with n8n and Python
Modern web scraping for B2B intelligence has shifted. Three years ago, a static requests session combined with BeautifulSoup was sufficient for extracting structured company profiles from most business directories. Today, client-side hydration (Next.js, Nuxt), dynamic fingerprinting, and complex pagination patterns cause naive scrapers to fail silently or stall out.
At the same time, commercial data vendors charge thousands of dollars a month for stale directory dumps that lack real-time enrichment. In this guide, we break down an enterprise-ready, decoupled architecture: a headless Python extraction engine orchestrated by n8n, complete with recursive pagination, schema validation via Pydantic, and multi-destination ingestion.
1. The Bottleneck: Why Traditional Pipelines Break
Scraping production directories typically breaks across three operational vectors:
- DOM Hydration & Anti-Automation Fingerprints: Modern directory listings defer rendering contact cards and metadata until specific micro-interactions occur (scroll events, viewport visibility). Standard HTTP clients pull empty HTML shells.
- Orchestration Coupling: Embedding proxy rotation, retry backoff, browser pools, and CRM export logic inside a single monolithic Python script creates unmaintainable code. A single unexpected DOM variation halts the entire batch.
- Schema Drift & Data Quality: Unsanitized scraped fields (inconsistent international dialing codes, obfuscated mailto links, malformed URLs) corrupt downstream analytics and CRM pipelines.
The Solution
Decouple extraction from orchestration:
- Headless Worker (Python / Playwright): Runs in an isolated container, handles stealth browser contexts, executes JavaScript micro-scrolls, and extracts raw data.
- Orchestration & State Management (n8n): Handles rate-limiting queues, recursive pagination states, failure retries with exponential backoff, and webhook dispatching.
- Data Integrity Layer (Pydantic): Enforces strict types, normalizes phone numbers to E.164, extracts naked domains, and validates emails before storage.
2. System Architecture
+-------------------------------------------------------------+
| n8n Orchestrator |
| [Trigger] -> [Queue / Pagination Loop] -> [Webhook Dispatch]|
+--------------+-------------------------------+--------------+
| HTTP POST ^
v | Clean JSON
+-------------------------------+ |
| Python FastAPI + Playwright | |
| - Stealth context init | |
| - Scroll hydration | |
| - DOM query extraction | |
+--------------+----------------+ |
| Raw Objects |
v |
+-------------------------------+ |
| Pydantic Validation Layer |--------------+
| - RFC Email validation |
| - E.164 Phone formatting |
| - Domain sanitization |
+-------------------------------+
- Queue Trigger: n8n pulls an unvisited directory page URL from a PostgreSQL queue or a configured array.
-
Browser Context Execution: n8n sends a payload to the internal Python FastAPI worker (
/scrapeendpoint). - Extraction & Normalization: Playwright runs headlessly, extracts targets, runs records through Pydantic validators, and returns a verified JSON array.
- Database Routing & Sync: n8n deduplicates records against existing primary keys and concurrently writes to PostgreSQL, Airtable, or your target CRM.
3. The Code & Core Logic
Step A: The Playwright Stealth Scraper (Python)
This worker exposes a lightweight FastAPI endpoint. It leverages Playwright with realistic user-agent headers, evasive navigator properties, and automatic waiting for selector hydration.
# worker.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, HttpUrl
from playwright.async_api import async_playwright
import asyncio
app = FastAPI()
class ScrapeRequest(BaseModel):
url: HttpUrl
selector_container: str
@app.post("/extract")
async def extract_directory_page(payload: ScrapeRequest):
async with async_playwright() as p:
browser = await p.chromium.launch(
headless=True,
args=["--disable-blink-features=AutomationControlled", "--no-sandbox"]
)
context = await browser.new_context(
user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36",
viewport={"width": 1920, "height": 1080}
)
page = await context.new_page()
try:
await page.goto(str(payload.url), wait_until="networkidle", timeout=30000)
# Trigger lazy loading by scrolling
await page.evaluate("window.scrollTo(0, document.body.scrollHeight / 2);")
await asyncio.sleep(1.2)
await page.evaluate("window.scrollTo(0, document.body.scrollHeight);")
await asyncio.sleep(1.0)
cards = await page.query_selector_all(payload.selector_container)
extracted = []
for card in cards:
name_el = await card.query_selector("h3.company-name")
website_el = await card.query_selector("a.website-link")
phone_el = await card.query_selector(".phone-display")
extracted.append({
"raw_name": await name_el.inner_text() if name_el else None,
"raw_website": await website_el.get_attribute("href") if website_el else None,
"raw_phone": await phone_el.inner_text() if phone_el else None,
})
return {"status": "success", "total": len(extracted), "records": extracted}
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
finally:
await context.close()
await browser.close()
Step B: Pydantic Validation & Normalization
Raw scraped data is notoriously messy. We enforce types and transform inputs into canonical shapes before the data reaches n8n:
# schemas.py
from pydantic import BaseModel, field_validator, HttpUrl
import re
from urllib.parse import urlparse
class CorporateRecord(BaseModel):
company_name: str
domain: str
phone: str | None = None
@field_validator("domain", mode="before")
@classmethod
def clean_domain(cls, v: str) -> str:
if not v:
raise ValueError("Empty website URL")
v = v.strip().lower()
if not v.startswith(("http://", "https://")):
v = f"https://{v}"
netloc = urlparse(v).netloc
return netloc.replace("www.", "")
@field_validator("phone", mode="before")
@classmethod
def normalize_phone(cls, v: str | None) -> str | None:
if not v:
return None
digits = re.sub(r"[^\d+]", "", v)
if len(digits) == 10 and not digits.startswith("+"):
return f"+1{digits}" # Standardize US numbers
return digits if len(digits) >= 8 else None
Step C: n8n Dynamic Pagination & Deduplication Logic
In n8n, use a Code Node to manage state, calculate the next page offset, and filter out records that are already tracked in the database:
// n8n Code Node: Process Page Results & Generate Next Target
const response = $input.first().json;
const existingDomains = $('Check Existing Domains').all().map(item => item.json.domain);
const currentOffset = $node["Pagination Loop"].json.offset || 0;
const pageSize = 20;
// Deduplicate incoming batch against DB results
const validRecords = response.records.filter(record => {
return record.domain && !existingDomains.includes(record.domain);
});
const hasMorePages = response.records.length === pageSize;
return {
json: {
validRecords: validRecords,
nextOffset: currentOffset + pageSize,
shouldContinue: hasMorePages,
processedAt: new Date().toISOString()
}
};
4. Deployment, Rate Limits, and Fault Tolerance
When running this architecture continuously, keep the following infrastructure practices in mind:
1. Handling Deadlocks & Stalled Instances
Playwright instances can occasionally freeze due to memory leaks on unhandled single-page apps. In Docker, configure strict container limits:
# docker-compose.yml snippet
services:
playwright-worker:
build: ./worker
restart: unless-stopped
deploy:
resources:
limits:
cpus: '2.0'
memory: 2048M
environment:
- PYTHONUNBUFFERED=1
2. Exponential Backoff in n8n
Inside the n8n HTTP Request Node that calls the extraction worker:
- Set Retry on Fail to
true. - Set Max Tries to
3. - Set Wait Between Tries to
5000ms with dynamic jitter to avoid burst patterns against the target directory.
3. CRM Rate-Limiting
CRMs like HubSpot and Salesforce enforce strict rate limits (e.g., 100 requests per 10 seconds). In n8n, insert a Split In Batches node set to batch sizes of 10, combined with a Wait node set to 1.5 seconds, ensuring pipeline compliance without dropping records.
5. Conclusion & Ready-to-Use Workflow
By decoupling the raw Playwright browser execution from the higher-level workflow orchestration in n8n, you get a clean, fault-tolerant B2B extraction engine that scales without requiring commercial SaaS subscriptions.
You can implement this architecture manually using the snippets above, or deploy our production-tested, turnkey package that includes pre-configured n8n JSON templates, Docker Compose files, stealth driver scripts, and Pydantic validation models:
- Instant Access on Whop: Get the Pipeline Template
-
Direct Download on Gumroad: Download on Gumroad ā use promo code
EARLYBIRDfor 20% off.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes ā full credit and traffic to the original publisher.