Filtering by marketCap and peFilter is Cheaper Than Filtering After the Run
Ingesting raw equity screening data into a downstream analytical warehouse frequently creates an expensive pipeline bottleneck. When engineers write scrapers to pull broad lists of public companies, they often collect ev
Ingesting raw equity screening data into a downstream analytical warehouse frequently creates an expensive pipeline bottleneck. When engineers write scrapers to pull broad lists of public companies, they often collect every listed security across major exchanges and defer filteringβsuch as valuation multiples or capitalization thresholdsβto subsequent transformation layers in dbt, DuckDB, or Snowflake. Pulling thousands of records daily when your ingestion model only targets profitable mid-cap companies inflates network payloads, fills object stores with discarded records, and increases extraction costs.
The Finviz Stock Screener Scraper resolves this by exposing Finviz's native screening engine directly through an extraction API. Rather than writing custom DOM scrapers that break whenever Finviz updates its table layouts, you can apply technical, fundamental, and descriptive filters upfront.
Why Scraping Unfiltered Financial Tables Breaks Down
Scraping Finviz directly through headless browser sessions or fragile HTML parsing scripts presents several operational headaches:
- Pagination and Table Volatility: Finviz organizes screener results into paginated HTML tables containing 20 rows per view. Navigating through hundreds of pages requires continuous DOM parsing, session state handling, and dynamic waiting to avoid throttling.
- Missing Field Inconsistencies: Financial tables have variable schema density. Growth tech stocks lack price-to-earnings metrics, debt-free firms omit debt-to-equity ratios, and newly listed equities lack 5-year averages. Custom scripts frequently throw parsing errors or output null values that misalign columnar arrays.
- Redundant Data Payloads: Extracting all 8,000+ US-listed equities just to extract the 150 that meet a specific valuation profile wastes pipeline bandwidth.
By structuring the extraction request directly around Finviz's server-side filtering parameters, you restrict the payload upstream. If you only care about profitable companies with low debt and positive earnings growth, you can pass those conditions directly to the scraper configuration.
Query Construction with Native Schema Parameters
The actor accepts an extensive list of configuration fields that match Finviz's native screener filters. Instead of extracting everything and querying downstream, you can construct a targeted JSON payload.
Key Input Fields
-
mode: Sets the operational mode. The default is"screener", but you can also set it to"stockOverview","newsForTicker","insiderTrading", or"groupPerformance". -
marketCap: Filters by capitalization tier. Options include"mega","large","mid","small","micro", and"nano". -
peFilter: Constrains trailing price-to-earnings ratios using values such as"profitable","low","u15","u20", or"o25". -
sector: Limits the universe to one of 11 sectors (for example,"Technology"or"Healthcare"). -
debtEquityFilter: Sets balance-sheet solvency bounds (such as"u0.5"or"low"). -
includeExtendedMetrics: A boolean flag. When set totrue, the standard screener row is enriched with ~45 additional fundamental and technical data points, includingforwardPE,peg,ps,pb,roe,roa,debtEq,sma20,sma50, andrsi.
Finviz omits empty fields from its output, meaning a metric is only present in a returned JSON record if Finviz published a value for it. This preserves sparse financial data without forcing you to handle synthetic null rows.
Step-by-Step Implementation
Setting up an automated run requires defining the input object and executing the actor through an HTTP client or SDK.
1. Define the Run Configuration
Construct a JSON object selecting mid-cap companies with a price-to-earnings ratio under 20 and a debt-to-equity ratio under 0.5 within the Technology sector, returning extended valuation metrics:
{
"mode": "screener",
"sector": "Technology",
"marketCap": "mid",
"peFilter": "u20",
"debtEquityFilter": "u0.5",
"includeExtendedMetrics": true,
"sortBy": "marketcap",
"sortDescending": true
}
2. Execute via the Python SDK
You can trigger this job and fetch results directly into your pipeline using the official client library:
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_API_TOKEN")
run_input = {
"mode": "screener",
"sector": "Technology",
"marketCap": "mid",
"peFilter": "u20",
"debtEquityFilter": "u0.5",
"includeExtendedMetrics": True,
"sortBy": "marketcap",
"sortDescending": True
}
run = client.actor("crawlerbros/finviz-scraper").call(run_input=run_input)
dataset_items = client.dataset(run["defaultDatasetId"]).list_items().items
for item in dataset_items:
print(
item.get("ticker"),
item.get("company"),
item.get("marketCap"),
item.get("peRatio"),
item.get("roe")
)
Each record contains the base fields (ticker, company, sector, industry, country, marketCap, peRatio, price, change, volume) along with the extended metrics requested via includeExtendedMetrics.
Output Structure and Schema Details
When running in screener mode with extended metrics enabled, records contain both basic trading data and deep fundamentals:
{
"ticker": "EXAMPLE",
"company": "Example Semiconductor Inc.",
"sector": "Technology",
"industry": "semiconductors",
"country": "USA",
"marketCap": "8.45B",
"peRatio": "16.42",
"forwardPE": "14.10",
"peg": "1.25",
"ps": "3.12",
"pb": "2.85",
"price": "45.20",
"change": "1.25%",
"volume": "1,245,800",
"avgVolume": "1.10M",
"debtEq": "0.32",
"roe": "18.40%",
"roa": "11.20%",
"rsi": "54.12",
"sma50": "43.10",
"finvizUrl": "https://finviz.com/quote.ashx?t=EXAMPLE",
"recordType": "stock",
"scrapedAt": "2025-05-15T14:30:00.000Z"
}
If you switch mode to "stockOverview" and supply a specific "ticker", the output pivots to full company fundamentals, adding fields like enterpriseValue, bookPerShare, shortRatio, and historical performance metrics.
Pricing and Cost Dynamics
The actor operates under a PAY_PER_EVENT model:
- Each result saved to the default dataset costs $0.005 on the FREE tier ($0.00433 on BRONZE, $0.00367 on SILVER, and $0.003 on GOLD, PLATINUM, and DIAMOND).
- An Actor Start event costs $0.005 per GB of memory allocated to the run.
- Platform usage for the run is billed separately at your Apify plan's rates.
Applying targeted parameters like marketCap, peFilter, and sector upfront keeps the output focused strictly on qualifying securities. Because each output item incurs a per-event result fee, narrowing the search space on Finviz's servers directly lowers the number of billed result events compared to scraping all listed equities and cleaning them downstream.
What This Tool Does Not Solve
This tool does not provide real-time, low-latency tick data or intraday order book depth; it reflects Finviz's web snapshot and screener updates, making it unsuitable for high-frequency trading execution pipelines.
For batch workflows, daily quantitative screens, and automated market monitoring, configuring filters natively before extraction provides a structured data stream without the overhead of maintaining local parsing scripts. An open question remains whether to run separate focused screens per sector or to merge them into a single cross-market pipeline.
Finviz Stock Screener Scraper is the Actor behind these examples. If a selector in your own version breaks, compare your output against the fields listed in its README first.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-09. Check the Actor page for the current rates.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.