Sitemap Extractor — Every URL with Dates, Images & hreflang avatar

Sitemap Extractor — Every URL with Dates, Images & hreflang

Pricing

$0.20 / 1,000 url delivereds

Go to Apify Store
Sitemap Extractor — Every URL with Dates, Images & hreflang

Sitemap Extractor — Every URL with Dates, Images & hreflang

Every URL in a website's sitemaps — found through robots.txt, nested indexes and .gz files followed — with last-modified date, change frequency, priority, images, hreflang and news tags.

Pricing

$0.20 / 1,000 url delivereds

Rating

0.0

(0)

Developer

yestrue

yestrue

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 hours ago

Last modified

Share

Get every URL listed in a website's sitemaps — with last-modified date, change frequency, priority, images, hreflang language versions and Google News tags. Give a website and its sitemaps are found through robots.txt; nested sitemap indexes and gzipped .xml.gz files are followed.

Why this one

  • 🧭 Finds the sitemaps for you. From robots.txt (the BBC lists 39, the New York Times 25), or /sitemap.xml and /sitemap_index.xml when robots.txt names none.
  • 🪆 Follows everything. Sitemap indexes inside indexes, gzipped sitemaps (recognised by their bytes, whatever the server calls them), and plain-text sitemaps.
  • 🏷️ All the tags. lastmod, changefreq, priority, image links, hreflang alternates, and news titles and dates.
  • 🧾 Honest results. A sitemap that fails is reported; the others are still delivered. A website without a sitemap comes back as a free record saying so. In a live test run, 40,770 URLs from five websites were delivered in one run.
  • 💸 You pay only for URLs delivered. Filters — URL contains, URL excludes, changed since — skip URLs for free.

What you get

For every URL:

FieldWhat it is
urlthe page
lastmod, changefreq, priorityas the sitemap gives them
imagesimage links listed for the page
alternateslanguage versions: hreflang and url
newsfor news sitemaps: title, publicationDate, publication
sitemap, site, inputwhich sitemap file and website it came from
scrapedAt, status, errorwhen it was read, and why a website gave nothing

Example

A real record from a live run, 24 September 2026:

{
"site": "https://www.allbirds.com",
"sitemap": "https://www.allbirds.com/sitemap_products_1.xml?from=1878194389061&to=7369944137808",
"url": "https://www.allbirds.com/products/mens-wool-runners-natural-white",
"lastmod": "2026-09-24T09:26:42-07:00",
"changefreq": "daily",
"images": ["https://cdn.shopify.com/s/files/1/1104/4168/files/Allbirds_WL_RN_SF_PDP_Natural_White_LAT_5a1e9e6b-9fa1-47d5-a34e-0ab02a0d2e2e.png?v=1751143945"],
"alternates": [],
"news": null,
"status": "OK"
}

How to use it

  1. Add Websites or sitemaps, one per line: example.com, or a sitemap link.
  2. Optionally keep only URLs containing some text (/products/, /blog/), skip others, or keep only URLs changed since a date.
  3. Set Max URLs per website.
  4. Press Start, and download the results from the Output tab as JSON, CSV or Excel.

To catch new and changed pages, schedule it daily with Changed since set to yesterday.

You can also call it from the Apify API, from Make, Zapier or n8n, or from an AI agent through the Apify MCP server.

Pricing

You are charged per URL delivered — see the price on this page. Websites without sitemaps and URLs your filters skip are free.

FAQ

Does it crawl the website? No. It reads what the sitemaps list, which is fast and gentle on the site. For pages missing from the sitemaps, use a crawler.

Why do some URLs have no lastmod? The sitemap does not give one. They are kept when you filter by Changed since.

Like it? A short review on the Store page helps other people find this Actor. Something missing or broken? Tell us on the Issues tab — we read every one.