Dev.to WebDev 🛠 Dev 👁 0 📖 21 min read

Why Preventing Duplicate SEMrush Scraper Output Needs Named Storage

How can I prevent duplicate processing of SEMrush Scraper output? To prevent duplicate output from the SEMrush Free Website Stats Scraper, your data pipeline needs a robust mechanism for tracking what has already been

How can I prevent duplicate processing of SEMrush Scraper output?

To prevent duplicate output from the SEMrush Free Website Stats Scraper, your data pipeline needs a robust mechanism for tracking what has already been processed. Maintain a list of all domains for which you have successfully retrieved and processed SEMrush stats, storing this information in persistent storage. Before initiating a new run, filter your input list of domains against this stored state.

While Apify offers named Key-Value Stores as a suitable solution for this, similar results can be achieved using other persistent storage solutions like Redis, a relational database, or cloud-based object storage. Named Key-Value Stores, unlike unnamed storages which expire, are persistent and provide a reliable place to store your processed domain identifiers and their scrapedAt timestamps. This approach, leveraging persistent storages, is crucial for long-running or scheduled pipelines, as free-tier accounts only retain the 10 most recent run storages for four months.

Here's how a Python application might manage this state before feeding a list of domains to the semrush-scraper Actor:

from apify_client import ApifyClient
import os

# Initialize ApifyClient with your API token
# For local testing, ensure APIFY_TOKEN is set in environment variables
apify_client = ApifyClient(os.environ["APIFY_TOKEN"])

# Define the name for your state storage
PROCESSED_DOMAINS_KV_STORE = "semrush-processed-domains"

def load_processed_domains(client: ApifyClient) -> set[str]:
    """Loads the set of domains previously processed from a named Key-Value Store."""
    try:
        kv_store = client.get_or_create_key_value_store(store_name=PROCESSED_DOMAINS_KV_STORE)
        processed_data = kv_store.get_record(key="processed_domains")
        if processed_data and processed_data["value"]:
            return set(processed_data["value"])
        return set()
    except Exception as e:
        print(f"Error loading processed domains: {e}")
        return set()

def save_processed_domains(client: ApifyClient, domains: set[str]):
    """Saves the current set of processed domains to a named Key-Value Store."""
    try:
        kv_store = client.get_or_create_key_value_store(store_name=PROCESSED_DOMAINS_KV_STORE)
        kv_store.set_record(key="processed_domains", value=list(domains), content_type="application/json")
    except Exception as e:
        print(f"Error saving processed domains: {e}")

# Example usage:
# all_domains_to_check = ["wikipedia.org", "github.com", "example.com", "apify.com"]
#
# processed = load_processed_domains(apify_client)
# print(f"Previously processed domains: {processed}")
#
# domains_for_new_run = [d for d in all_domains_to_check if d not in processed]
# print(f"Domains for current run: {domains_for_new_run}")
#
# # In a real scenario, you'd now pass domains_for_new_run to the Actor.
# # After the Actor runs and data is consumed, you'd update 'processed' and save it.

After a scraper run completes, fetch its dataset items, process them, and then persist the identifiers of successfully handled items into a named Apify Key-Value Store or an external database to prevent reprocessing in future runs. This post-run update of your state ensures that only newly acquired or modified data is considered for downstream actions. Once the SEMrush Free Website Stats Scraper Actor run finishes, you need to retrieve its results from the default dataset and then update your state.

Here's how you might fetch the dataset items and update your set of processed domains, including handling domains for which SEMrush reported no data (notFound):

from apify_client import ApifyClient
import os

apify_client = ApifyClient(os.environ["APIFY_TOKEN"])
PROCESSED_DOMAINS_KV_STORE = "semrush-processed-domains"

# load_processed_domains and save_processed_domains functions are defined as above
# ...

def process_and_update_state(client: ApifyClient, run_id: str, current_processed_domains: set[str]):
    """
    Fetches dataset items from a run, processes them, and updates the set of processed domains.
    """
    run = client.get_run(run_id=run_id)
    dataset = client.get_dataset(dataset_id=run["defaultDatasetId"])

    newly_processed_domains = set()
    for item in dataset.iterate_items():
        domain = item.get("domain")
        if domain:
            # Here you would typically send 'item' to your downstream system
            # For demonstration, we just mark it as processed
            print(f"Processing data for domain: {domain}. Not found: {item.get('notFound', False)}")
            newly_processed_domains.add(domain)
        else:
            print(f"Warning: Dataset item missing 'domain' field: {item}")

    updated_processed_domains = current_processed_domains.union(newly_processed_domains)
    save_processed_domains(client, updated_processed_domains)
    print(f"State updated. Total processed domains: {len(updated_processed_domains)}")

# Example of how you would integrate this:
#
# # 1. Load existing state
# existing_processed = load_processed_domains(apify_client)
# all_input_domains = ["apify.com", "example.com", "nodata.org", "wikipedia.org"] # Your full list
#
# # 2. Filter input domains
# domains_for_run = [d for d in all_input_domains if d not in existing_processed]
#
# if domains_for_run:
#     print(f"Starting run for {len(domains_for_run)} new domains.")
#     run_input = {
#         "domains": domains_for_run,
#         "mode": "full",
#         "proxyConfiguration": {"useApifyProxy": True, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "US"}
#     }
#     actor_run = apify_client.actor("crawlerbros/semrush-scraper").call(run_input=run_input)
#     print(f"Actor run started with ID: {actor_run['id']}")
#
#     # 3. After the run completes (or after fetching its status), process output and update state
#     # In a real pipeline, you'd wait for run completion or use a webhook.
#     # For this example, let's assume 'actor_run' is already a completed run object.
#     # process_and_update_state(apify_client, actor_run["id"], existing_processed)
# else:
#     print("No new domains to process.")

By storing the set of domain values for every item successfully pulled from the dataset, you create a robust processing history. Even if a run fails midway or a downstream system experiences an outage, subsequent runs can always pick up where the last successful processing left off, preventing redundant work and ensuring data consistency. The notFound status, when present in an item, indicates that SEMrush has no data for that domain. This is still a processed state; you've checked the domain, and the Actor has reported its status. Therefore, such domains should also be added to your processed_domains set to avoid repeatedly checking them in future runs.

What unique identifier works best for SEMrush Website Stats?

The domain field from the scraper's output, once normalized, serves as the most effective unique identifier for tracking individual records and preventing duplicates across pipeline runs. Each record represents the stats for a single domain, making domain a natural candidate for an ID.

The semrush-scraper output for each domain includes a domain field that has already been normalized (e.g., https://www.custify.com/ becomes custify.com). This ensures consistency and makes it suitable for use as a primary key in a database or a unique identifier in your state management logic.

Consider the typical structure of a single output item from the semrush-scraper:

{
  "type": "semrush_website_stats",
  "domain": "wikipedia.org",
  "authorityScore": 92,
  "visits": 1100000000,
  "visitsText": "1.1B",
  "organicSearchTraffic": 485000000,
  "organicSearchTrafficText": "485M",
  "referringDomains": 3900000,
  "referringDomainsText": "3.9M",
  "referringDomainsChange": 1.2,
  "backlinks": 2960000000,
  "backlinksText": "2.96B",
  "backlinksChange": 0.5,
  "asOf": "July 2026",
  "sourceUrl": "https://www.semrush.com/website/wikipedia.org",
  "scrapedAt": "2026-09-28T10:30:00.000Z"
}

The domain field is consistently present and represents the specific website whose statistics were scraped. For tracking historical data, domain can be combined with asOf (the data month) or scrapedAt (the UTC ISO timestamp of the scrape) to create a composite key, allowing you to store multiple data points for the same domain over time without collisions. When a domain is not found by SEMrush, the output includes notFound: true and a reason, but crucially, the domain field is still present, allowing you to record that specific domain was checked and no data was available.

How do I orchestrate runs beyond the 300-second synchronous limit?

For semrush-scraper runs that might exceed five minutes due to enableBrowserFallback for many domains, use the asynchronous run API endpoint (POST /v2/acts/<actor>/runs) and implement a webhook or polling mechanism to receive results. This avoids the hard-capped synchronous endpoint which returns HTTP 408 past 300 seconds.

The semrush-scraper README explicitly states that some domains, particularly those where SEMrush doesn't have a public overview page, trigger a "browser fallback" which involves launching a real Chrome instance, warming Google cookies, navigating to SEMrush, and rotating through residential proxy IPs. This process can easily take "60+ seconds for some domains." If your input domains list contains many such cases, a batch run is very likely to exceed the 300-second (5-minute) synchronous API timeout.

To reliably run the scraper for larger batches or domains requiring browser fallback, you must initiate the Actor asynchronously and then either poll its status or configure a webhook to be notified upon completion. This kind of orchestration could also be handled by dedicated workflow management systems like Airflow, Prefect, or custom-built solutions using cloud functions and message queues.

from apify_client import ApifyClient
import os

apify_client = ApifyClient(os.environ["APIFY_TOKEN"])
ACTOR_ID = "crawlerbros/semrush-scraper"

def run_semrush_scraper_async(domains_list: list[str], webhook_url: str = None):
    """
    Initiates an asynchronous run of the SEMrush scraper and optionally sets up a webhook.
    """
    run_input = {
        "domains": domains_list,
        "mode": "full",
        "enableBrowserFallback": True,
        "proxyConfiguration": {"useApifyProxy": True, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "US"}
    }

    # Define the webhook if a URL is provided
    webhooks = []
    if webhook_url:
        webhooks.append({
            "eventTypes": ["ACTOR.RUN.SUCCEEDED", "ACTOR.RUN.FAILED"],
            "requestUrl": webhook_url,
            "payloadTemplate": "{\"runId\": {{run.id}}, \"datasetId\": {{run.defaultDatasetId}}, \"status\": \"{{run.status}}\", \"actorId\": \"{{run.actorId}}\"}"
        })

    # Call the Actor asynchronously
    print(f"Initiating asynchronous run for {len(domains_list)} domains...")
    actor_run = apify_client.actor(ACTOR_ID).call(
        run_input=run_input,
        webhooks=webhooks,
        timeout_secs=0, # Crucial: setting timeout_secs to 0 ensures asynchronous execution
        wait_for_finish=0 # Also ensures async, does not wait for finish
    )
    print(f"Actor run initiated. Run ID: {actor_run['id']}")
    print(f"View run details at: https://console.apify.com/actors/{ACTOR_ID}/runs/{actor_run['id']}")
    return actor_run['id']

# Example usage:
# domains_to_scrape = ["github.com", "youtube.com", "example.com", "nodata.org", "wikipedia.org", "apify.com"]
# my_webhook_endpoint = "https://your-n8n-or-service-webhook-url.com/semrush-callback"
#
# if domains_to_scrape:
#    run_id = run_semrush_scraper_async(domains_to_scrape, my_webhook_endpoint)
#    print(f"Scraper run {run_id} started asynchronously. Webhook configured to {my_webhook_endpoint}.")
# else:
#    print("No domains to scrape.")

When you call an Actor and set timeout_secs=0 (or omit it and rely on the default behavior for POST requests), the API immediately returns the runId, and the Actor execution continues in the background. You can then use this runId to poll the run status or, more efficiently, rely on the webhook to trigger your downstream processing once the run completes, fails, or is aborted. This asynchronous pattern is fundamental for building robust pipelines that aren't constrained by short API timeouts.

How does the mode input parameter affect the SEMrush Scraper's output?

The mode input parameter controls the scope of data returned by the semrush-scraper, allowing you to retrieve either a full set of metrics or narrow the output to authority_only, backlinks_only, or traffic_only. This enables efficient data retrieval by only fetching the specific information relevant to your use case.

The semrush-scraper offers four distinct mode options, which directly influence which fields are populated in the output for each domain. By default, the mode is set to full, providing a comprehensive set of metrics. However, if your application only requires a specific subset of data, selecting a more granular mode can potentially optimize the scraping process by reducing the amount of data that needs to be extracted and processed from the raw SEMrush page. While the Actor's pricing is per result item, fetching less data might reduce downstream processing load and storage.

Here's how each mode impacts the output fields:

  • full (default): This mode returns all available metrics for a domain. This includes authorityScore, visits (and visitsText), organicSearchTraffic (and organicSearchTrafficText), referringDomains (and referringDomainsText, referringDomainsChange), and backlinks (and backlinksText, backlinksChange). It also always includes type, domain, asOf, sourceUrl, and scrapedAt.
    {
      "domains": ["wikipedia.org"],
      "mode": "full"
    }
    ```
{% endraw %}


    The resulting output item would contain all fields, such as:
{% raw %}


```json
    {
      "type": "semrush_website_stats",
      "domain": "wikipedia.org",
      "authorityScore": 92,
      "visits": 1100000000,
      "visitsText": "1.1B",
      "organicSearchTraffic": 485000000,
      "organicSearchTrafficText": "485M",
      "referringDomains": 3900000,
      "referringDomainsText": "3.9M",
      "referringDomainsChange": 1.2,
      "backlinks": 2960000000,
      "backlinksText": "2.96B",
      "backlinksChange": 0.5,
      "asOf": "July 2026",
      "sourceUrl": "https://www.semrush.com/website/wikipedia.org",
      "scrapedAt": "2026-09-28T10:30:00.000Z"
    }
    ```
{% endraw %}


*   **{% raw %}`authority_only`{% endraw %}**: This mode focuses solely on retrieving the {% raw %}`authorityScore`{% endraw %}. Other metric-related fields will be omitted from the output.
{% raw %}


```json
    {
      "domains": ["google.com"],
      "mode": "authority_only"
    }
    ```
{% endraw %}


    An example output for {% raw %}`authority_only`{% endraw %}:
{% raw %}


```json
    {
      "type": "semrush_website_stats",
      "domain": "google.com",
      "authorityScore": 100,
      "asOf": "July 2026",
      "sourceUrl": "https://www.semrush.com/website/google.com",
      "scrapedAt": "2026-09-28T10:30:00.000Z"
    }
    ```
{% endraw %}


*   **{% raw %}`backlinks_only`{% endraw %}**: This mode retrieves {% raw %}`backlinks`{% endraw %}, {% raw %}`backlinksText`{% endraw %}, {% raw %}`backlinksChange`{% endraw %}, {% raw %}`referringDomains`{% endraw %}, {% raw %}`referringDomainsText`{% endraw %}, and {% raw %}`referringDomainsChange`{% endraw %}. Traffic and authority score data are excluded.
{% raw %}


```json
    {
      "domains": ["apify.com"],
      "mode": "backlinks_only"
    }
    ```
{% endraw %}


    An example output for {% raw %}`backlinks_only`{% endraw %}:
{% raw %}


```json
    {
      "type": "semrush_website_stats",
      "domain": "apify.com",
      "referringDomains": 10000,
      "referringDomainsText": "10K",
      "referringDomainsChange": 0.8,
      "backlinks": 500000,
      "backlinksText": "500K",
      "backlinksChange": 1.1,
      "asOf": "July 2026",
      "sourceUrl": "https://www.semrush.com/website/apify.com",
      "scrapedAt": "2026-09-28T10:30:00.000Z"
    }
    ```
{% endraw %}


*   **{% raw %}`traffic_only`{% endraw %}**: This mode exclusively fetches {% raw %}`visits`{% endraw %}, {% raw %}`visitsText`{% endraw %}, {% raw %}`organicSearchTraffic`{% endraw %}, and {% raw %}`organicSearchTrafficText`{% endraw %}. Backlink and authority score data are omitted.
{% raw %}


```json
    {
      "domains": ["github.com"],
      "mode": "traffic_only"
    }
    ```
{% endraw %}


    An example output for {% raw %}`traffic_only`{% endraw %}:
{% raw %}


```json
    {
      "type": "semrush_website_stats",
      "domain": "github.com",
      "visits": 450000000,
      "visitsText": "450M",
      "organicSearchTraffic": 200000000,
      "organicSearchTrafficText": "200M",
      "asOf": "July 2026",
      "sourceUrl": "https://www.semrush.com/website/github.com",
      "scrapedAt": "2026-09-28T10:30:00.000Z"
    }
    ```
{% endraw %}


Choosing the appropriate {% raw %}`mode`{% endraw %} is a key consideration for optimizing your data pipeline, especially when dealing with large volumes of domains and specific data requirements. It ensures that you only retrieve the data you truly need, potentially reducing processing overhead and clarifying your data analysis focus.

## How can I specify the proxy configuration for SEMrush Scraper runs?

You can specify the proxy configuration for {% raw %}`semrush-scraper`{% endraw %} runs using the {% raw %}`proxyConfiguration`{% endraw %} input field, which allows you to define whether to use Apify Proxy, which proxy groups to use, and even target specific countries. This is essential because SEMrush's reCAPTCHA v3 often rejects datacenter IPs, necessitating the use of residential proxies for successful scraping. Other proxy solutions could be integrated via custom Actor development, but the {% raw %}`semrush-scraper`{% endraw %} is designed around Apify Proxy.

The {% raw %}`proxyConfiguration`{% endraw %} input object is critical for ensuring the {% raw %}`semrush-scraper`{% endraw %} can bypass reCAPTCHA v3, which SEMrush uses to guard its free website checker. The Actor's README explicitly states that reCAPTCHA v3 "rejects datacenter IPs" by scoring them as 0.0 (bot). Therefore, using residential proxies is a prerequisite for reliable scraping, particularly for domains that trigger the browser fallback.

The {% raw %}`proxyConfiguration`{% endraw %} object has several key properties:

*   **{% raw %}`useApifyProxy`{% endraw %} (boolean)**: A flag to indicate whether to use the Apify Proxy. For this Actor, it should almost always be {% raw %}`true`{% endraw %}.
*   **{% raw %}`apifyProxyGroups`{% endraw %} (array of strings)**: Specifies which proxy groups from the Apify Proxy to use. The Actor's default and recommendation is {% raw %}`["RESIDENTIAL"]`{% endraw %}. Residential proxies simulate real user traffic from real residential IP addresses, making them much less likely to be blocked by services like reCAPTCHA v3.
*   **{% raw %}`apifyProxyCountry`{% endraw %} (string)**: Allows you to specify a particular country for the residential proxy IPs. For example, {% raw %}`US`{% endraw %} for United States. Apify Proxy supports country targeting, and even US state granularity (e.g., {% raw %}`US_CA`{% endraw %} for California). This can be useful if your target audience or the specific SEMrush instance you are hitting is geographically sensitive, or if you want to ensure the most stable connection from a certain region. If left unspecified, Apify Proxy will select from available residential IPs globally.

Here's an example of how you might set the {% raw %}`proxyConfiguration`{% endraw %} in your input:
{% raw %}


```json
{
  "domains": ["example.com", "apify.com"],
  "mode": "full",
  "enableBrowserFallback": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": ["RESIDENTIAL"],
    "apifyProxyCountry": "US"
  }
}

If you omit the proxyConfiguration field entirely, the Actor defaults to using {"useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "US"}, which is a sensible default for most use cases. However, explicitly defining it ensures clarity and allows for customization if, for instance, you need to scrape from a non-US region.

Keep in mind that using residential proxies contributes to your Apify platform usage costs, which are separate from the Actor's event-based pricing. Residential proxy sessions typically persist for around 30 minutes, meaning the Actor might rotate through several IPs during a longer run, especially for domains requiring extensive browser fallback. This automatic rotation is managed by the Apify Proxy and the Actor's internal logic, abstracting away the complexity of proxy management from your pipeline.

How can I connect to downstream systems like n8n using webhooks?

Integrate the semrush-scraper output directly into downstream systems like n8n by configuring a webhook on Actor completion to push dataset item URLs or directly trigger a workflow for data extraction and transformation. This eliminates the need for manual polling or complex custom API integrations on your end.

Webhooks provide a real-time, event-driven mechanism to connect Apify Actor runs with other services. When an Actor run finishes (or any other specified event occurs), Apify sends an HTTP POST request to a URL you define. This payload can contain details about the run, including the ID of the default dataset where the results are stored. This event-driven approach could also be replicated with custom orchestration logic using message queues, but webhooks provide a direct integration.

For n8n, a powerful workflow automation tool, this integration is particularly seamless. n8n has a dedicated Apify Trigger node that can listen for ACTOR.RUN.SUCCEEDED or ACTOR.RUN.FAILED events. When an Actor run completes, the Apify platform sends a notification to your n8n webhook URL, providing the run ID and default dataset ID. Your n8n workflow can then use the Apify node to fetch the dataset items and proceed with transformation, storage, or further processing.

Here's an example of how you might create a webhook for an Actor run using the Apify API (though the apify-client in Python simplifies this, as shown in the previous section):

curl -X POST "https://api.apify.com/v2/webhooks?token=YOUR_API_TOKEN" \
     -H "Content-Type: application/json" \
     -d '{
           "eventTypes": ["ACTOR.RUN.SUCCEEDED", "ACTOR.RUN.FAILED"],
           "requestUrl": "https://your-n8n-webhook-url.com/semrush-data-processor",
           "payloadTemplate": "{\"runId\": {{run.id}}, \"datasetId\": {{run.defaultDatasetId}}, \"status\": \"{{run.status}}\", \"actorId\": \"{{run.actorId}}\"}"
         }'

The payloadTemplate is highly customizable. You can include any details from the run object, allowing your downstream system to receive precisely the information it needs to identify the source of the data and fetch it. This event-driven architecture is far more efficient than polling, reducing idle waiting times and immediately kicking off processing as soon as data is available.

For setting up webhooks programmatically, ensure your webhooks array in the call method of apify_client.actor() is correctly structured, specifying the eventTypes, requestUrl, and payloadTemplate.

Designing input for scalability and partial failures

Structure the semrush-scraper input by pre-filtering domains based on prior processing state and by considering smaller batch sizes, which can isolate failures and simplify recovery compared to monolithic runs. This strategy enhances resilience and makes debugging easier when issues arise within a large input set.

The semrush-scraper accepts an array of domains as its primary input. While it's tempting to throw a massive list of domains into a single run, this has several drawbacks for robust pipelines:

  1. Single Point of Failure: If a single run fails (e.g., due to a platform issue, network instability, or hitting maxTotalChargeUsd), you might lose context on which domains were successfully processed and which were not, complicating recovery.
  2. Debugging Complexity: Diagnosing issues within a run that processed thousands of domains is much harder than examining a run with a few hundred.
  3. Timeout Risk: As discussed, domains requiring browser fallback can be slow. A large batch amplifies the risk of exceeding platform timeouts, even with asynchronous calls.
  4. Cost Control: While maxTotalChargeUsd can terminate a run, splitting into smaller batches gives you more granular control over spend and allows for pause-and-review points.

Instead, combine your state management (filtering already processed domains) with batching. This means taking your domains_for_new_run (after filtering) and splitting them into smaller, manageable chunks (e.g., 500-1000 domains per chunk). You can then initiate separate asynchronous Actor runs for each batch. A critical platform constraint to remember here is that a request queue can only be processed by one Actor or task run at a time. So, if your original domain list was in a request queue, you'd need a different strategy, like the one shown, where the orchestrator directly passes domains arrays as input. Batching and orchestration could also be managed by external systems like AWS Step Functions, Azure Logic Apps, or Google Cloud Workflows, rather than solely within Apify.

from apify_client import ApifyClient
import os

apify_client = ApifyClient(os.environ["APIFY_TOKEN"])
ACTOR_ID = "crawlerbros/semrush-scraper"
PROCESSED_DOMAINS_KV_STORE = "semrush-processed-domains"

# load_processed_domains function is defined as above
# ...

def chunk_list(lst: list, n: int):
    """Yield successive n-sized chunks from lst."""
    for i in range(0, len(lst), n):
        yield lst[i:i + n]

def orchestrate_batched_semrush_runs(all_domains_to_check: list[str], batch_size: int = 500, webhook_url: str = None):
    """
    Orchestrates batched, asynchronous runs of the SEMrush scraper with state management.
    """
    existing_processed = load_processed_domains(apify_client)
    domains_for_new_runs = [d for d in all_domains_to_check if d not in existing_processed]

    if not domains_for_new_runs:
        print("No new domains to process.")
        return []

    print(f"Found {len(domains_for_new_runs)} new domains for processing.")
    run_ids = []
    for i, domain_batch in enumerate(chunk_list(domains_for_new_runs, batch_size)):
        print(f"Starting batch {i+1} with {len(domain_batch)} domains...")
        run_input = {
            "domains": domain_batch,
            "mode": "full",
            "enableBrowserFallback": True,
            "proxyConfiguration": {"useApifyProxy": True, "apifyProxyGroups": ["RESIDENTIAL"], "apifyProxyCountry": "US"}
        }

        webhooks = []
        if webhook_url:
            webhooks.append({
                "eventTypes": ["ACTOR.RUN.SUCCEEDED", "ACTOR.RUN.FAILED"],
                "requestUrl": webhook_url,
                "payloadTemplate": "{\"runId\": {{run.id}}, \"datasetId\": {{run.defaultDatasetId}}, \"status\": \"{{run.status}}\", \"actorId\": \"{{run.actorId}}\", \"batchNumber\": " + str(i+1) + "}"
            })

        actor_run = apify_client.actor(ACTOR_ID).call(
            run_input=run_input,
            webhooks=webhooks,
            timeout_secs=0,
            wait_for_finish=0
        )
        run_ids.append(actor_run['id'])
        print(f"Batch {i+1} started with Run ID: {actor_run['id']}")
    return run_ids

# Example of full pipeline orchestration:
#
# full_list_of_all_domains = ["domain1.com", "domain2.org", "domainN.net"] # Your complete list
# my_callback_webhook = "https://your-n8n-or-service-webhook-url.com/semrush-batch-callback"
#
# initiated_run_ids = orchestrate_batched_semrush_runs(full_list_of_all_domains, batch_size=200, webhook_url=my_callback_webhook)
# print(f"All batches initiated with run IDs: {initiated_run_ids}")
#
# # Your webhook handler would then need to process each batch's dataset and update the master state.

This batched approach allows for parallel processing of smaller, more manageable units. If one batch run fails, others can still succeed. Your state management logic (updating processed_domains) would then need to aggregate results from all completed batches.

Understanding the semrush-scraper event-based pricing

The semrush-scraper charges $0.002 for each "result" item emitted, with potential discounts based on platform usage, plus a $0.005 "Actor Start" charge per GB of memory allocated once per run, in addition to standard Apify platform usage. These event prices are not subject to subscription plan rates.

Understanding the cost model is crucial for any pipeline. The semrush-scraper uses a PAY_PER_EVENT pricing model, meaning you are charged for specific actions or outputs generated by the Actor, on top of the general Apify platform usage (which covers things like compute, proxy, and storage).

Here are the specific charged events and their prices:

  • "result" (apify-default-dataset-item): This event is charged for every single result record that the Actor pushes to its default dataset.
    • Base Price: $0.002 per event (per item). Discounts may apply based on platform usage.
    • Impact of input: The number of "result" events scales directly with the number of domains you provide in the input. If you request stats for 1,000 domains and all yield a result, you will incur 1,000 "result" charges.
  • "Actor Start" (apify-actor-start): This event is charged when the Actor begins its execution.
    • Price: $0.005 per GB of memory allocated to the run. Even if the Actor uses less than 1GB, the minimum charge is for 1GB.
    • Impact of input: This is a flat charge per run based on memory allocation and does not directly scale with the number of domains input, but rather with the resources required for the Actor to initiate and operate.

It's important to remember that these event charges are separate from your general Apify platform usage, which is billed based on your specific Apify plan's rates for compute, proxy, and storage. The event prices themselves do not change based on your Apify subscription plan, though your overall platform usage costs will. When considering the total cost of a run, you must factor in both the event charges and the platform usage. For example, using residential proxies (which this Actor requires for many domains) will contribute to your platform proxy usage.

What are the limitations and considerations for SEMrush data pipelines?

The semrush-scraper is limited to free-tier SEMrush data, requires residential proxies which can increase run duration, and relies on a dynamic anti-detection browser fallback that makes run times variable and potentially long for "notFound" domains. These factors necessitate careful pipeline design to manage cost and performance expectations.

While powerful, the semrush-scraper (and any scraping tool) operates within specific boundaries:

  1. Free-Tier Data Only: The Actor only scrapes SEMrush's publicly available "Website Traffic Checker" data. It does not provide access to paid SEMrush features like full keyword research, competitor analysis, or detailed backlink profiles. If your use case requires deeper SEMrush insights, this Actor will not suffice, and you'll need to consider a paid SEMrush API subscription.
  2. Residential Proxy Requirement: SEMrush employs reCAPTCHA v3, which heavily penalizes datacenter IP addresses. As such, the Actor explicitly requires Apify Proxy Residential IPs. This is prefilled and necessary for the Actor to function effectively, particularly for domains triggering the browser fallback. While Apify Proxy handles the IP rotation and management, it adds to your platform usage cost. Other proxy solutions exist, but integrating them would require modifying the Actor's code or building a custom scraper. Residential proxy sessions typically persist for around 30 minutes, which the Actor accounts for by rotating IPs, but it can still influence the consistency of very long-running browser fallback operations. You can target specific countries (e.g., US or even US_TX for Texas) via apifyProxyCountry if your use case demands it.
  3. Variable and Potentially Long Run Times: As noted in the Actor's FAQ, domains that SEMrush doesn't have a public overview page for (the "false 404" cases) require the "browser fallback" mechanism. This involves launching a real Chrome browser, warming Google cookies, navigating to SEMrush, and rotating through up to 8 residential proxy IPs to find one with a sufficient reCAPTCHA score. This significantly extends run duration for those specific domains (potentially 60+ seconds per domain), impacting the overall run time for batches with many such domains. Fast-path domains, conversely, complete in under 5 seconds.
  4. No Nulls/Sentinels: The output schema explicitly states, "Empty fields are omitted, no nulls, no sentinels." This means your downstream parsing logic should anticipate the absence of certain fields rather than expecting null values. For example, if a domain has no organic search traffic, the organicSearchTraffic field might simply not be present in the JSON record.
  5. Data Freshness: The asOf field provides the data month (e.g., July 2026), indicating the freshness of the SEMrush data. This is useful for tracking changes over time but means you're scraping data that SEMrush itself last updated, which might not be real-time.

These limitations aren't showstoppers, but they are crucial for designing a realistic and efficient data pipeline. Account for the variable run times, the specific data available, and the proxy requirements when planning your infrastructure, budget, and data processing logic.

Checked against the Actor's input schema and Apify docs on 2026-10-08.

The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email [email protected]

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.