Sitemap URL Extractor: Auto-Discovery, Filters & Health Report avatar

Sitemap URL Extractor: Auto-Discovery, Filters & Health Report

Pricing

$0.50 / 1,000 results

Go to Apify Store
Sitemap URL Extractor: Auto-Discovery, Filters & Health Report

Sitemap URL Extractor: Auto-Discovery, Filters & Health Report

Get every URL from a website's sitemaps: just enter the domain. Handles sitemap indexes and .xml.gz, adds lastmod, images, hreflang and news data, filters by date or URL pattern, and reports sitemap problems.

Pricing

$0.50 / 1,000 results

Rating

0.0

(0)

Developer

CreativeFour LLC

CreativeFour LLC

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

What does the Sitemap URL Extractor do?

Enter just a domain and get every URL from the site's sitemaps, with each page's last-modified date, change frequency, and priority, plus any images, language versions (hreflang), and news data the sitemap declares.

It finds sitemaps automatically (from robots.txt, or the usual paths like /sitemap.xml), follows sitemap indexes to any depth, reads gzipped .xml.gz files, and writes a sitemap health report: duplicates, off-domain URLs, invalid dates, and broken sitemap files.

On Apify you can schedule it, call it from the API, or send the list straight into Google Sheets, a crawler, or an SEO tool.

Why use it?

  • Get a site's full page list in seconds, without crawling it.
  • Watch what changes. Use Only URLs modified on or after with a schedule to see what a competitor published or updated this week.
  • Audit your own sitemaps. Catch duplicates, URLs pointing at another domain, broken child sitemaps, and missing or invalid <lastmod> dates.
  • Feed other tools. Pass the URLs to a status checker, crawler, Lighthouse audit, or RAG pipeline.
  • Check international SEO. See each page's hreflang alternates side by side.

How to use it

  1. Open the Input tab and type a domain (for example example.com), or paste sitemap URLs.
  2. Optional: add a modified since date, or URL patterns to include or skip.
  3. Click Start.
  4. Open the Output tab for the URL list, and the SUMMARY record for the health report.

Input

FieldWhat it does
WebsitesDomains. Sitemaps are discovered from robots.txt or the usual paths.
Sitemap URLsOr give sitemap or index URLs directly (.xml or .xml.gz).
Only URLs modified on or afterKeeps URLs whose <lastmod> is on or after this date.
Only URLs matching / Skip URLs matchingRegular expressions or plain text, such as /blog/ or /tag/.
Max URLs / Max sitemap filesSafety limits.
{
"websites": ["example.com"],
"modifiedSince": "2026-09-01",
"includePatterns": ["/blog/"]
}

Output

One row per URL. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

{
"site": "example.com",
"url": "https://example.com/blog/new-post",
"lastmod": "2026-09-24T10:00:00Z",
"changefreq": "weekly",
"priority": 0.8,
"images": ["https://example.com/img/hero.jpg"],
"alternates": [{ "hreflang": "es", "href": "https://example.com/es/blog/new-post" }],
"newsTitle": null,
"newsPublishedAt": null,
"videos": 0,
"sitemap": "https://example.com/post-sitemap.xml.gz",
"offDomain": false,
"lastmodValid": true
}

SUMMARY (key-value store) gives, for each site: the number of sitemaps read, URLs found and saved, duplicates, off-domain URLs, invalid and missing lastmod dates, and every sitemap file that failed to load, with the reason.

Data fields

FieldDescription
url, lastmod, changefreq, priorityStandard sitemap fields
images, videosImage URLs and video count, from image and video sitemap extensions
alternateshreflang language versions
newsTitle, newsPublishedAtNews sitemap fields
sitemapWhich sitemap file listed the URL
offDomain, lastmodValidHealth flags

How much does it cost to extract sitemap URLs?

You pay per URL saved. Filters (modified-since and URL patterns) are applied first, so you only pay for the URLs you keep. Set a maximum charge per run in the run options, and the Actor stops cleanly at that limit.

Tips

Use it from AI agents (MCP)

AI agents can find and run this Actor through the Apify MCP server.

  • Claude, ChatGPT, or any MCP client: add https://mcp.apify.com?tools=creativefour/sitemap-url-extractor as a custom connector, and sign in to Apify when prompted.
  • Claude Code, Cursor, VS Code, or Codex: run apify mcp install claude-code (swap in your client's name), then ask your agent to "list every URL on example.com with creativefour/sitemap-url-extractor".

FAQ and support

What if a site has no sitemap? The summary reports "No readable sitemap found". Sitemaps are optional, and some sites don't publish one.

Does it visit every page? No. It reads only the sitemap files, which makes it fast and light on the site.

Found a bug or need a feature? Open an issue on the Issues tab. Custom versions are available on request.