Dev.to WebDev 🛠 Dev 👁 0 📖 4 min read

How to check HTTP status codes and redirect chains for a list of URLs in bulk

You just migrated a website, or inherited a spreadsheet of 5,000 backlinks, and you need to know which URLs still work, which ones redirect, and where the redirects end up. Opening them one by one in a browser hides the

You just migrated a website, or inherited a spreadsheet of 5,000 backlinks, and you need to know which URLs still work, which ones redirect, and where the redirects end up. Opening them one by one in a browser hides the redirect chain, and most single URL checkers online stop at a handful of links. You want one row per URL with the status code, the full chain of hops and a clear "broken or not" flag.

The Bulk URL Status Code and Redirect Chain Checker, published by Hay Equipos on the Apify Store, does exactly that. It is useful for SEO migrations, broken link audits, backlink and affiliate link checks, and cleaning a URL list before you hand it to another tool.

What you get back

One row per URL. Example (illustrative values):

{
  "url": "http://example.com/old-page",
  "ok": true,
  "isBroken": false,
  "statusCode": 301,
  "finalStatusCode": 200,
  "finalUrl": "https://www.example.com/new-page",
  "redirectCount": 2,
  "redirectChain": [
    { "url": "http://example.com/old-page", "statusCode": 301, "method": "HEAD", "location": "https://example.com/old-page" },
    { "url": "https://example.com/old-page", "statusCode": 301, "method": "HEAD", "location": "https://www.example.com/new-page" },
    { "url": "https://www.example.com/new-page", "statusCode": 200, "method": "HEAD", "location": null }
  ],
  "redirectsToHttps": true,
  "changedDomain": true,
  "redirectLoop": false,
  "tooManyRedirects": false,
  "contentType": "text/html; charset=utf-8",
  "server": "nginx",
  "xRobotsTag": null,
  "hsts": true,
  "responseTimeMs": 840
}

As a table:

url statusCode finalStatusCode isBroken redirectCount finalUrl
http://example.com/old-page 301 200 false 2 https://www.example.com/new-page
https://example.com/missing 404 404 true 0 https://example.com/missing
https://no-such-host.example/x null null true 0 (error: ENOTFOUND)

isBroken is true for a final 4xx or 5xx, a redirect loop, too many hops, or no answer at all. ok is true only for a final 2xx. Each row also carries contentLength, lastModified and cacheControl from the final answer. URLs that get no HTTP answer carry an error such as ENOTFOUND, ECONNREFUSED or Timed out.

Step by step in the Apify Console

  1. Open the actor on the Apify Store (link at the end) and click Try for free.
  2. Add your addresses to URLs, one per line, or paste a whole column into Or paste a list (spaces, commas or new lines all work). Duplicates are removed and addresses without a scheme get https:// added.
  3. Leave Follow redirects on to record every hop, and set Maximum redirect hops (default 10, up to 20).
  4. Keep Retry with GET when HEAD is refused on. Some servers reject the lightweight HEAD request but serve the page normally, and the GET retry reads at most 64 KB.
  5. Adjust Requests per second per site (default 3) and Timeout per request (default 20 seconds) if needed.
  6. Leave Respect robots.txt on unless you are checking a site you own or have permission to check.
  7. Click Start and export the results from the Output tab as CSV, Excel or JSON.

Calling it from code

The actor id is pistachio_implementation/url-status-redirect-checker. Keep your token in APIFY_TOKEN.

curl:

curl -X POST \
  "https://api.apify.com/v2/acts/pistachio_implementation~url-status-redirect-checker/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["http://apify.com", "https://httpbin.org/status/404"], "followRedirects": true}'

Python with apify-client, checking a migration redirect map:

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])

old_urls = [
    "https://example.com/blog/post-1",
    "https://example.com/blog/post-2",
]

run = client.actor("pistachio_implementation/url-status-redirect-checker").call(
    run_input={"urls": old_urls, "followRedirects": True, "maxRedirects": 10}
)

for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    if row.get("isBroken") or row.get("redirectCount", 0) > 1:
        print(row["url"], row.get("finalStatusCode"), row.get("finalUrl"), row.get("error"))

After a migration, a healthy old URL usually shows finalStatusCode 200, redirectCount 1, and the finalUrl you expect.

Pricing

Pay per event: $0.0008 per URL checked ($0.80 per 1,000). A URL is charged when a server answered it with any HTTP status, including 404 and 500, because that answer is the result you asked for. No start fee and no platform usage on top. These rows are free:

  • invalid URLs
  • URLs with no HTTP answer (the domain does not exist, the connection is refused, or it times out)
  • URLs skipped because robots.txt closes them to automated tools

Limits and what it does not do

  • No JavaScript. Redirects done by JavaScript or by a meta refresh tag are not followed. The page's own status code is reported.
  • It does not crawl. It checks exactly the URLs you give it and does not discover links on a page. To get every URL of a site, extract its sitemap first and pass that list in.
  • Some sites answer automated requests from cloud servers with 403 or a challenge page. That status is reported as received. The actor does not try to get around such blocks.
  • It identifies itself as ApifyLinkChecker and respects robots.txt by default, including rules written for Apify crawlers.
  • Up to 50,000 URLs per run. Many hosts are checked in parallel (25 at a time), but requests to the same host are spaced out to your per site limit, so a long list on one domain runs slower than a list spread across many.
  • 429 and 5xx answers are retried with backoff before they are reported, and a host that times out twice is treated as not answering for the rest of the run.

Try it on the Apify Store: https://apify.com/pistachio_implementation/url-status-redirect-checker

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.