Sitemap Extractor: All URLs from Any Website, $0.50/1K
Pricing
from $0.35 / 1,000 sitemap urls
Sitemap Extractor: All URLs from Any Website, $0.50/1K
Get every page URL of a website from its XML sitemaps: sitemaps are found through robots.txt, sitemap indexes are followed, and each URL comes with last modified date, change frequency and priority. Filter by text. $0.50 per 1,000 URLs.
Pricing
from $0.35 / 1,000 sitemap urls
Rating
0.0
(0)
Developer
Don Mangu
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
This sitemap extractor gets every page URL of a website from its XML sitemaps. Enter a bare domain such as gov.uk or a sitemap address, and it finds the sitemaps through robots.txt, falls back to the usual addresses (/sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml), follows sitemap indexes, unpacks .gz files and also reads plain text sitemaps. Each URL comes with last modified date, change frequency and priority.
Cost: $0.0005 per URL ($0.50 per 1,000). Websites without a sitemap, unreadable sitemaps and invalid entries are free. No login or API key.
How to extract all URLs from a sitemap
- Enter websites or sitemap addresses under Websites or sitemaps.
- Optionally set URL contains to keep one section such as /blog/ or /docs/, and URLs per website.
- Click Start, then open the Sitemap URLs table or download the list.
The form opens with GOV.UK and one NASA post sitemap, 50 URLs each. It takes a few seconds and costs about 5 cents. Typical uses are SEO audits (every indexable URL with its last modified date), content inventories and migration checks (compare the URL lists of an old and a new site), and URL lists for other Actors such as a scraper, a screenshot or a Markdown Actor.
How much does the sitemap URL extractor cost?
| Event | Price |
|---|---|
Sitemap URL (sitemap-url) | $0.0005 ($0.50 per 1,000) |
| Website without a sitemap, unreadable sitemap or invalid entry | Free |
Apify adds only its small standard fee per run start. You can set a spending limit on any run, and the Actor stops cleanly when it is reached.
Example: a full inventory of a 40,000-page store costs 40,000 x $0.0005 = $20.
Input
{"startUrls": ["https://www.gov.uk","https://www.nasa.gov/wp-sitemap-posts-post-1.xml"],"maxUrlsPerSite": 50}
Output
One row per result:
{"input": "https://www.gov.uk","url": "https://www.gov.uk/browse/benefits","lastModified": "2026-09-20T10:15:00.000Z","changeFrequency": null,"priority": null,"sitemapUrl": "https://www.gov.uk/sitemaps/sitemap_1.xml","status": "ok","charged": true}
status is ok for URLs, which are charged. A website where no sitemap could be read gives one free row with no_sitemap, no_urls or robots_disallowed and the reason. A STATS record counts URLs, free rows and requests.
Related Actors
- Broken Link Checker: Use it to find broken internal and external links on a website.
- Web Page to Markdown for AI: Use it to turn web pages into clean Markdown for LLMs, RAG and AI agents.
- AI Crawler Access Audit: Use it to check which AI crawlers a website allows or blocks in robots.txt.
- Website Screenshot API: Use it to capture full-page screenshots or PDFs of a list of URLs.
FAQ
Does it crawl pages? No. It reads sitemap files only, which is fast and cheap. Pages that are missing from the sitemap are not found.
Does it respect robots.txt? Yes. Sitemaps are read only where robots.txt allows the token DonMangu-SitemapExtractor or all crawlers.
Which formats work? XML sitemaps and sitemap indexes, gzipped files, image sitemap entries (counted) and plain text sitemaps.
Can I limit it to one section? Set URL contains to a path such as /blog/ or /docs/.