Sitemap Extractor Pro: URLs, lastmod, images, diff
Pricing
from $0.15 / 1,000 sitemap urls
Sitemap Extractor Pro: URLs, lastmod, images, diff
Extract every URL from XML, gzipped, index and text sitemaps with lastmod, priority, images, news and hreflang; filter by glob or date and diff against the previous run.
Pricing
from $0.15 / 1,000 sitemap urls
Rating
0.0
(0)
Developer
Spongy Frame Tools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
19 hours ago
Last modified
Categories
Share
What does Sitemap Extractor Pro do?
Sitemap Extractor Pro reads XML sitemaps and returns every URL they list as clean, structured data. Give it a sitemap, a sitemap index, a gzipped sitemap, a robots.txt, or just a homepage, and it will:
- Follow sitemap indexes recursively (nested indexes up to a configurable depth, with cycle detection).
- Decompress gzipped sitemaps, detected by content rather than by the
.gzextension. - Read plain-text sitemaps (one URL per line).
- Discover sitemaps from a homepage by checking
/robots.txtSitemap:lines and the common paths/sitemap.xml,/sitemap_index.xmland/sitemap-index.xml. - Extract
lastmod,changefreq,priority, image URLs, video count, Google News title and publication date, andhreflangalternates from each<url>entry. - Filter URLs by glob pattern and by last-modified date range.
- Remove duplicates across sitemaps (the first occurrence wins).
- Optionally compare the URL set with your previous run and report which URLs were added or removed.
It never visits the pages themselves. It only downloads sitemap files, so runs are fast and cheap even for sites with hundreds of thousands of URLs.
Why use this one?
Many sitemap tools either stop at the first index level, choke on gzipped or malformed files, or return a flat list of URLs with no metadata. This Actor is built for the messy reality of real sitemaps:
- Streaming XML parsing keeps memory flat on very large files, with a hard 50 MB per-file limit so a single giant sitemap cannot crash the run.
- Malformed XML is parsed in recovery mode; entries that can be salvaged are kept and the file is listed in the run summary as partially parsed.
- A bad URL never fails the run. Fetch errors, 404s, unknown formats and oversized files are recorded under
failuresin theRUN_SUMMARYkey-value record and the crawl continues. - Pricing is per URL, so you pay for what you get, and
maxItemscaps the spend up front.
What it does not do: it does not crawl pages, does not check whether URLs return 200, and does not render JavaScript. If a site has no sitemap at all, the result is empty.
How to use it
- Enter one or more start URLs. A sitemap URL is best, but a homepage or robots.txt URL works too.
- Optionally set
maxItems, glob or date filters, and turn ondiffAgainstPreviousRunif you want change detection. - Run the Actor and download the dataset as JSON, CSV or Excel, or read it through the API.
Input
{"startUrls": ["https://www.apify.com/sitemap.xml"],"maxItems": 500,"outputFields": "full","includeGlobs": ["https://www.apify.com/*"],"excludeGlobs": ["*.pdf"],"modifiedAfter": "2026-01-01","maxDepth": 5,"followSitemapsFromRobots": true,"respectRobotsDisallow": false,"diffAgainstPreviousRun": false}
Only startUrls is required. Globs use shell-style wildcards (*, ?, [abc]) matched against the whole URL. Date filters accept ISO 8601 dates or datetimes; URLs without a lastmod are dropped when a date filter is active because their modification date is unknown.
Output
With outputFields: "full" every dataset item looks like this:
{"url": "https://example.com/news/story","lastmod": "2026-09-20T08:30:00+00:00","changefreq": "daily","priority": 0.8,"sitemap_url": "https://example.com/sitemap-news.xml","source_url": "https://example.com/sitemap_index.xml","images": ["https://cdn.example.com/img/1.jpg"],"videos_count": 1,"news": {"title": "Big story", "publication_date": "2026-09-20T08:00:00+00:00"},"alternates": [{"hreflang": "de", "href": "https://example.com/de/news/story"}],"depth": 1}
With outputFields: "urlsOnly" items contain just url and sitemap_url. Dates are normalised to ISO 8601 with a timezone; missing values are null. depth is 0 for URLs found directly in a start sitemap and increases by one for each index level.
When diffAgainstPreviousRun is on, one extra item with record_type: "diff_report" is appended containing added, removed, added_count, removed_count, unchanged_count and previous_run_found. A summary of every run (counts, failures, warnings, elapsed time) is saved as RUN_SUMMARY in the default key-value store.
How much does it cost?
The Actor uses pay-per-event pricing. There is no subscription and no charge for failed fetches.
| Event | When it is charged | Price |
|---|---|---|
sitemap-url | Once per URL written to the dataset | $0.0002 |
diff-report | Once per diff report, only when a previous run existed to compare against | $0.005 |
Examples:
- A 2,000-URL site with
maxItems: 1000produces 1,000 items: 1,000 × $0.0002 = $0.20. - Weekly change monitoring of a 5,000-URL site with
diffAgainstPreviousRun: true: 5,000 × $0.0002 + $0.005 = $1.005 per run. The very first run stores the baseline and costs $1.00.
Platform compute is included in the event prices. If your account's spending limit is lower than the cost of maxItems, the Actor lowers maxItems to fit and says so in the log.
Limits and fair use
- Sitemap files larger than 50 MB (after decompression) are skipped and reported. Split such sitemaps or point the Actor at the child sitemaps directly.
maxDepthdefaults to 5 nested index levels, which covers every real-world site we have seen.- Requests use a descriptive User-Agent, a 30-second timeout, three retries with exponential backoff, and honour
Retry-Afteron 429 and 503 responses. Please do not hammer small sites with repeated large runs. - Duplicate URLs are removed per run. Cross-run deduplication is available through the diff feature.
Legal note
The Actor only reads publicly served sitemap files, which site owners publish specifically so that they can be read by automated tools. It does not visit pages, does not collect personal data, and outputs only URLs and the metadata declared in the sitemap. You are responsible for how you use the resulting URL lists.
FAQ
Why did my run stop early? Either maxItems was reached, or your Apify pay-per-event spending limit for this run was hit. Both cases are logged and recorded in RUN_SUMMARY under limit_reason. Raise the limit and run again; the diff feature will pick up where the previous set ended.
The dataset is empty. What happened? Check RUN_SUMMARY.failures. Common causes: the site has no sitemap at the usual locations, the sitemap URL returned an HTML page (some sites redirect unknown paths to the homepage), or a date filter removed every URL because the sitemap has no lastmod values.
Can I get only new URLs since last time? Yes. Turn on diffAgainstPreviousRun and use the same startUrls each time. The diff_report record lists added and removed URLs. Note that if a run is truncated by maxItems, the comparison is against a partial set, which the report flags with current_run_truncated.
Does it respect robots.txt Disallow rules? It reads robots.txt to find sitemaps. With respectRobotsDisallow enabled it adds a robots_disallowed flag to each item, but nothing is filtered, because the Actor never requests the pages themselves.
Can I filter by path or file type? Use includeGlobs and excludeGlobs, for example https://example.com/blog/* or *.pdf. Patterns match the full URL.
Support
Found a bug or missing a feature? Open a ticket in the Issues tab of this Actor. We respond within 24 hours on working days.