Sitemap Extractor — Every URL with Dates, Images & hreflang
Pricing
$0.20 / 1,000 url delivereds
Sitemap Extractor — Every URL with Dates, Images & hreflang
Every URL in a website's sitemaps — found through robots.txt, nested indexes and .gz files followed — with last-modified date, change frequency, priority, images, hreflang and news tags.
Pricing
$0.20 / 1,000 url delivereds
Rating
0.0
(0)
Developer
yestrue
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
9 hours ago
Last modified
Categories
Share
Get every URL listed in a website's sitemaps — with last-modified date, change frequency, priority, images, hreflang language versions and Google News tags. Give a website and its sitemaps are found through robots.txt; nested sitemap indexes and gzipped .xml.gz files are followed.
Why this one
- 🧭 Finds the sitemaps for you. From robots.txt (the BBC lists 39, the New York Times 25), or
/sitemap.xmland/sitemap_index.xmlwhen robots.txt names none. - 🪆 Follows everything. Sitemap indexes inside indexes, gzipped sitemaps (recognised by their bytes, whatever the server calls them), and plain-text sitemaps.
- 🏷️ All the tags.
lastmod,changefreq,priority, image links,hreflangalternates, and news titles and dates. - 🧾 Honest results. A sitemap that fails is reported; the others are still delivered. A website without a sitemap comes back as a free record saying so. In a live test run, 40,770 URLs from five websites were delivered in one run.
- 💸 You pay only for URLs delivered. Filters — URL contains, URL excludes, changed since — skip URLs for free.
What you get
For every URL:
| Field | What it is |
|---|---|
url | the page |
lastmod, changefreq, priority | as the sitemap gives them |
images | image links listed for the page |
alternates | language versions: hreflang and url |
news | for news sitemaps: title, publicationDate, publication |
sitemap, site, input | which sitemap file and website it came from |
scrapedAt, status, error | when it was read, and why a website gave nothing |
Example
A real record from a live run, 24 September 2026:
{"site": "https://www.allbirds.com","sitemap": "https://www.allbirds.com/sitemap_products_1.xml?from=1878194389061&to=7369944137808","url": "https://www.allbirds.com/products/mens-wool-runners-natural-white","lastmod": "2026-09-24T09:26:42-07:00","changefreq": "daily","images": ["https://cdn.shopify.com/s/files/1/1104/4168/files/Allbirds_WL_RN_SF_PDP_Natural_White_LAT_5a1e9e6b-9fa1-47d5-a34e-0ab02a0d2e2e.png?v=1751143945"],"alternates": [],"news": null,"status": "OK"}
How to use it
- Add Websites or sitemaps, one per line:
example.com, or a sitemap link. - Optionally keep only URLs containing some text (
/products/,/blog/), skip others, or keep only URLs changed since a date. - Set Max URLs per website.
- Press Start, and download the results from the Output tab as JSON, CSV or Excel.
To catch new and changed pages, schedule it daily with Changed since set to yesterday.
You can also call it from the Apify API, from Make, Zapier or n8n, or from an AI agent through the Apify MCP server.
Pricing
You are charged per URL delivered — see the price on this page. Websites without sitemaps and URLs your filters skip are free.
FAQ
Does it crawl the website? No. It reads what the sitemaps list, which is fast and gentle on the site. For pages missing from the sitemaps, use a crawler.
Why do some URLs have no lastmod? The sitemap does not give one. They are kept when you filter by Changed since.
Like it? A short review on the Store page helps other people find this Actor. Something missing or broken? Tell us on the Issues tab — we read every one.