Dev.to WebDev šŸ›  Dev šŸ‘ 0 šŸ“– 5 min read

How I Built an Autonomous B2B Directory Scraping & Enrichment Pipeline with n8n and Python

How I Built an Autonomous B2B Directory Scraping & Enrichment Pipeline with n8n and Python Modern web scraping for B2B intelligence has shifted. Three years ago, a static requests session combined with BeautifulSoup wa

How I Built an Autonomous B2B Directory Scraping & Enrichment Pipeline with n8n and Python

Modern web scraping for B2B intelligence has shifted. Three years ago, a static requests session combined with BeautifulSoup was sufficient for extracting structured company profiles from most business directories. Today, client-side hydration (Next.js, Nuxt), dynamic fingerprinting, and complex pagination patterns cause naive scrapers to fail silently or stall out.

At the same time, commercial data vendors charge thousands of dollars a month for stale directory dumps that lack real-time enrichment. In this guide, we break down an enterprise-ready, decoupled architecture: a headless Python extraction engine orchestrated by n8n, complete with recursive pagination, schema validation via Pydantic, and multi-destination ingestion.

1. The Bottleneck: Why Traditional Pipelines Break

Scraping production directories typically breaks across three operational vectors:

  1. DOM Hydration & Anti-Automation Fingerprints: Modern directory listings defer rendering contact cards and metadata until specific micro-interactions occur (scroll events, viewport visibility). Standard HTTP clients pull empty HTML shells.
  2. Orchestration Coupling: Embedding proxy rotation, retry backoff, browser pools, and CRM export logic inside a single monolithic Python script creates unmaintainable code. A single unexpected DOM variation halts the entire batch.
  3. Schema Drift & Data Quality: Unsanitized scraped fields (inconsistent international dialing codes, obfuscated mailto links, malformed URLs) corrupt downstream analytics and CRM pipelines.

The Solution

Decouple extraction from orchestration:

  • Headless Worker (Python / Playwright): Runs in an isolated container, handles stealth browser contexts, executes JavaScript micro-scrolls, and extracts raw data.
  • Orchestration & State Management (n8n): Handles rate-limiting queues, recursive pagination states, failure retries with exponential backoff, and webhook dispatching.
  • Data Integrity Layer (Pydantic): Enforces strict types, normalizes phone numbers to E.164, extracts naked domains, and validates emails before storage.

2. System Architecture

   +-------------------------------------------------------------+
   |                      n8n Orchestrator                       |
   |  [Trigger] -> [Queue / Pagination Loop] -> [Webhook Dispatch]|
   +--------------+-------------------------------+--------------+
                  | HTTP POST                     ^
                  v                               | Clean JSON
   +-------------------------------+              |
   |  Python FastAPI + Playwright  |              |
   |  - Stealth context init       |              |
   |  - Scroll hydration           |              |
   |  - DOM query extraction       |              |
   +--------------+----------------+              |
                  | Raw Objects                   |
                  v                               |
   +-------------------------------+              |
   |   Pydantic Validation Layer   |--------------+
   |  - RFC Email validation       |
   |  - E.164 Phone formatting     |
   |  - Domain sanitization        |
   +-------------------------------+
  1. Queue Trigger: n8n pulls an unvisited directory page URL from a PostgreSQL queue or a configured array.
  2. Browser Context Execution: n8n sends a payload to the internal Python FastAPI worker (/scrape endpoint).
  3. Extraction & Normalization: Playwright runs headlessly, extracts targets, runs records through Pydantic validators, and returns a verified JSON array.
  4. Database Routing & Sync: n8n deduplicates records against existing primary keys and concurrently writes to PostgreSQL, Airtable, or your target CRM.

3. The Code & Core Logic

Step A: The Playwright Stealth Scraper (Python)

This worker exposes a lightweight FastAPI endpoint. It leverages Playwright with realistic user-agent headers, evasive navigator properties, and automatic waiting for selector hydration.

# worker.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, HttpUrl
from playwright.async_api import async_playwright
import asyncio

app = FastAPI()

class ScrapeRequest(BaseModel):
    url: HttpUrl
    selector_container: str

@app.post("/extract")
async def extract_directory_page(payload: ScrapeRequest):
    async with async_playwright() as p:
        browser = await p.chromium.launch(
            headless=True,
            args=["--disable-blink-features=AutomationControlled", "--no-sandbox"]
        )
        context = await browser.new_context(
            user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36",
            viewport={"width": 1920, "height": 1080}
        )
        page = await context.new_page()

        try:
            await page.goto(str(payload.url), wait_until="networkidle", timeout=30000)

            # Trigger lazy loading by scrolling
            await page.evaluate("window.scrollTo(0, document.body.scrollHeight / 2);")
            await asyncio.sleep(1.2)
            await page.evaluate("window.scrollTo(0, document.body.scrollHeight);")
            await asyncio.sleep(1.0)

            cards = await page.query_selector_all(payload.selector_container)
            extracted = []

            for card in cards:
                name_el = await card.query_selector("h3.company-name")
                website_el = await card.query_selector("a.website-link")
                phone_el = await card.query_selector(".phone-display")

                extracted.append({
                    "raw_name": await name_el.inner_text() if name_el else None,
                    "raw_website": await website_el.get_attribute("href") if website_el else None,
                    "raw_phone": await phone_el.inner_text() if phone_el else None,
                })

            return {"status": "success", "total": len(extracted), "records": extracted}
        except Exception as e:
            raise HTTPException(status_code=500, detail=str(e))
        finally:
            await context.close()
            await browser.close()

Step B: Pydantic Validation & Normalization

Raw scraped data is notoriously messy. We enforce types and transform inputs into canonical shapes before the data reaches n8n:

# schemas.py
from pydantic import BaseModel, field_validator, HttpUrl
import re
from urllib.parse import urlparse

class CorporateRecord(BaseModel):
    company_name: str
    domain: str
    phone: str | None = None

    @field_validator("domain", mode="before")
    @classmethod
    def clean_domain(cls, v: str) -> str:
        if not v:
            raise ValueError("Empty website URL")
        v = v.strip().lower()
        if not v.startswith(("http://", "https://")):
            v = f"https://{v}"
        netloc = urlparse(v).netloc
        return netloc.replace("www.", "")

    @field_validator("phone", mode="before")
    @classmethod
    def normalize_phone(cls, v: str | None) -> str | None:
        if not v:
            return None
        digits = re.sub(r"[^\d+]", "", v)
        if len(digits) == 10 and not digits.startswith("+"):
            return f"+1{digits}"  # Standardize US numbers
        return digits if len(digits) >= 8 else None

Step C: n8n Dynamic Pagination & Deduplication Logic

In n8n, use a Code Node to manage state, calculate the next page offset, and filter out records that are already tracked in the database:

// n8n Code Node: Process Page Results & Generate Next Target
const response = $input.first().json;
const existingDomains = $('Check Existing Domains').all().map(item => item.json.domain);
const currentOffset = $node["Pagination Loop"].json.offset || 0;
const pageSize = 20;

// Deduplicate incoming batch against DB results
const validRecords = response.records.filter(record => {
  return record.domain && !existingDomains.includes(record.domain);
});

const hasMorePages = response.records.length === pageSize;

return {
  json: {
    validRecords: validRecords,
    nextOffset: currentOffset + pageSize,
    shouldContinue: hasMorePages,
    processedAt: new Date().toISOString()
  }
};

4. Deployment, Rate Limits, and Fault Tolerance

When running this architecture continuously, keep the following infrastructure practices in mind:

1. Handling Deadlocks & Stalled Instances

Playwright instances can occasionally freeze due to memory leaks on unhandled single-page apps. In Docker, configure strict container limits:

# docker-compose.yml snippet
services:
  playwright-worker:
    build: ./worker
    restart: unless-stopped
    deploy:
      resources:
        limits:
          cpus: '2.0'
          memory: 2048M
    environment:
      - PYTHONUNBUFFERED=1

2. Exponential Backoff in n8n

Inside the n8n HTTP Request Node that calls the extraction worker:

  • Set Retry on Fail to true.
  • Set Max Tries to 3.
  • Set Wait Between Tries to 5000 ms with dynamic jitter to avoid burst patterns against the target directory.

3. CRM Rate-Limiting

CRMs like HubSpot and Salesforce enforce strict rate limits (e.g., 100 requests per 10 seconds). In n8n, insert a Split In Batches node set to batch sizes of 10, combined with a Wait node set to 1.5 seconds, ensuring pipeline compliance without dropping records.

5. Conclusion & Ready-to-Use Workflow

By decoupling the raw Playwright browser execution from the higher-level workflow orchestration in n8n, you get a clean, fault-tolerant B2B extraction engine that scales without requiring commercial SaaS subscriptions.

You can implement this architecture manually using the snippets above, or deploy our production-tested, turnkey package that includes pre-configured n8n JSON templates, Docker Compose files, stealth driver scripts, and Pydantic validation models:

šŸ“° Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.