Dev.to WebDev 🛠 Dev 👁 0 📖 4 min read

How to get all page URLs of a website from its sitemap

You need a list of every page on a website. Maybe you are auditing your own sitemap, listing a competitor's blog posts with dates, planning a migration, or feeding a crawler or a RAG pipeline. Crawling the whole site lin

You need a list of every page on a website. Maybe you are auditing your own sitemap, listing a competitor's blog posts with dates, planning a migration, or feeding a crawler or a RAG pipeline. Crawling the whole site link by link is slow and expensive. The site already publishes the list in its sitemap, but large sites split it across dozens of files, nested indexes and gzip archives.

Sitemap URL Extractor by Hay Equipos finds those sitemaps, follows every index, opens the gzip files and returns one clean row per URL.

How it finds the sitemaps

Give it a domain and it reads every Sitemap: line in robots.txt first. If there are none, it tries /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /wp-sitemap.xml and /sitemap.txt. You can also give sitemap links directly: .xml, .xml.gz, .txt, sitemap indexes, robots.txt, or an RSS or Atom feed. Sitemap indexes are followed breadth first.

The pages themselves are never crawled, which is why listing is fast and cheap.

What you get back

One row per unique URL. Example with illustrative values:

{
  "url": "https://www.example.com/blog/how-we-cut-load-time",
  "lastmod": "2026-08-14T09:30:00.000Z",
  "changefreq": "monthly",
  "priority": 0.7,
  "alternates": [
    { "hreflang": "en-us", "url": "https://www.example.com/blog/how-we-cut-load-time" },
    { "hreflang": "de-de", "url": "https://www.example.com/de/blog/how-we-cut-load-time" }
  ],
  "imageCount": 2,
  "videoCount": 0,
  "newsTitle": null,
  "newsPublishedAt": null,
  "sitemapUrl": "https://www.example.com/sitemap-posts.xml",
  "site": "https://www.example.com",
  "statusCode": 301,
  "finalUrl": "https://www.example.com/blog/how-we-cut-load-time/",
  "redirected": true,
  "scrapedAt": "2026-09-27T06:14:24.933Z"
}

statusCode, finalUrl and redirected appear only when the status check is on. A RUN_SUMMARY record in the run's key value store shows, for each site, how the sitemap was found, how many sitemap files were read, any sitemap errors and how many URLs were saved.

Step by step in the Apify Console

  1. Open the actor on the Apify Store (link at the end) and click Try for free.
  2. In Websites or sitemap URLs, add domains or sitemap links, one per line.
  3. Filter if you like: Keep URLs matching any of and Drop URLs matching any of accept plain words or regular expressions (for example /blog/ to keep, /tag/ to drop). Changed on or after keeps only URLs with a lastmod on or after a date. Same host only drops URLs on other hosts.
  4. Turn on Check HTTP status of each URL to find pages that are broken or redirected but still listed in the sitemap.
  5. Set the caps: Maximum sitemap files per site (default 500), Maximum URLs per site (default 50,000) and Maximum URLs in total (default 200,000).
  6. Click Start and export as CSV, Excel or JSON. With the status check on, filter for statusCode 404 or redirected true.

Calling it from code

With curl:

curl -X POST \
  "https://api.apify.com/v2/acts/pistachio_implementation~sitemap-url-extractor/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "startUrls": ["https://www.example.com"],
    "includeUrlPatterns": ["/blog/"],
    "lastmodAfter": "2026-01-01",
    "maxUrlsPerSite": 5000
  }'

With Python and the apify-client package, checking a sitemap for broken pages:

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])

run = client.actor("pistachio_implementation/sitemap-url-extractor").call(
    run_input={
        "startUrls": ["https://www.example.com"],
        "sameHostOnly": True,
        "checkStatus": True,
        "maxUrlsPerSite": 2000,
    }
)

for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    status = row.get("statusCode")
    if status is None or status >= 400 or row.get("redirected"):
        print(status, row["url"], "->", row.get("finalUrl"))

summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("RUN_SUMMARY")
for site in summary["value"]["sites"]:
    print(site["input"], site["discoveredVia"], site["sitemapsRead"], site["urlsSaved"])

Pricing

Pay per event, with no subscription and no platform usage charge on top:

Event Price
URL saved $0.00025 per URL ($0.25 per 1,000 URLs)
URL status checked $0.001 per URL ($1.00 per 1,000 checks), only with the status check on

There is no start fee. Filtered URLs, duplicates and sites without a sitemap cost nothing. For example, listing 20,000 URLs costs $5.00, and adding the status check to those 20,000 adds $20.00. You can set a maximum charge per run and the actor stops when it is reached.

Limits and what it does not do

  • It lists what the sitemaps say. Pages missing from the sitemap are not found, because pages are not crawled.
  • Sites that block automated requests to their sitemap return an error in RUN_SUMMARY.
  • lastmod is whatever the site writes. Some sites set every URL to today's date.
  • With a date filter set, URLs without a lastmod are dropped, and sitemap files whose own lastmod is older than the date are skipped.
  • The status check is much slower than listing. It sends a HEAD request (a GET if HEAD is refused) at about three requests a second per site.

Sitemaps exist so that machines can read them, but the content they point to belongs to the site owner. Use the data in line with each source site's terms.

Try it on the Apify Store: https://apify.com/pistachio_implementation/sitemap-url-extractor

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.