Sitemap URL Extractor — PDF/DOCX Tags + robots.txt
Pricing
from $0.20 / 1,000 url discovereds
Sitemap URL Extractor — PDF/DOCX Tags + robots.txt
Discover every URL in XML sitemaps via robots.txt and indexes; tag PDF, DOCX, XLSX and more for SEO audits and RAG ingest. Auto-finds nested indexes and .xml.gz; exports DOC_TO_MARKDOWN_INPUT for the PDF/DOCX Actor.
Pricing
from $0.20 / 1,000 url discovereds
Rating
0.0
(0)
Developer
新世紀書僮
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Sitemap URL Extractor — Find PDF & Document Links
Get every URL a website publishes in its XML sitemaps — with lastmod, source sitemap and a file-type tag, so PDFs, Word, Excel and PowerPoint files stand out.
Enter a homepage and the Actor finds the sitemaps for you (robots.txt Sitemap: lines plus /sitemap.xml and /sitemap_index.xml), follows nested sitemap indexes, reads gzipped .xml.gz and plain-text sitemaps, removes duplicates and saves one row per URL. Use it for SEO audits, content inventories, change monitoring, or to collect a site's PDF and DOCX documents for RAG and LLM pipelines.
What you get
- 🔎 Automatic sitemap discovery — robots.txt
Sitemap:lines + common paths; or paste sitemap URLs directly - 🌳 Sitemap index recursion — nested indexes up to a depth you choose, loops and repeats skipped
- 🗜️ All common formats — XML
urlsetandsitemapindex,.xml.gz(gzip file or gzip transfer encoding), plain-text sitemaps, RSS/Atom feeds - 🏷️ File-type tag per URL —
pdf,docx,doc,xlsx,xls,pptx,ppt,csv,html,image,other(from the extension; optional HTTP HEAD check confirms the real Content-Type) - 📄 Ready-made input for PDF/DOCX conversion — all PDF and DOCX URLs are saved as
DOC_TO_MARKDOWN_INPUT, formatted as input for the PDF & DOCX to Markdown Actor - 🧯 Broken sitemaps don't break the run — every sitemap gets its own status line (
ok,not_found,error,skipped) with the reason inSITEMAP_REPORT; one bad sitemap or site never fails the whole run - 🐢 Polite and robust — low per-host concurrency, timeouts, retries with backoff for timeouts / 429 / 5xx (honours
Retry-After), size caps, tolerant parsing of malformed XML - 🧮 Filters — file types, lastmod date, include/exclude regex, same domain only, max URLs (total and per site)
- 💾 Light — 256 MB memory by default; HTTP only, no browser
Measured results
Tested by us on 2026-09-29 with small limits (local runs plus 3 runs on the Apify platform):
| Test | Sitemap setup | Result |
|---|---|---|
| fec.gov (US government) | robots.txt lists 3 sitemaps incl. a PDF-only sitemap | 69 PDF URLs tagged pdf; HEAD check on all 69: HTTP 200 application/pdf (69/69) |
| gov.uk | sitemap index → 5.6 MB child sitemap | 20,000 URLs in 11 s on the Apify platform at 256 MB (peak memory 107 MB) |
| nist.gov | Drupal sitemap index with 56 child sitemaps | index followed, URLs saved up to the limit |
| developer.mozilla.org | index of 10 .xml.gz sitemaps | gzip sitemaps read |
| yoast.com | Yoast SEO index with 21 child sitemaps; /sitemap.xml is a second index to the same files | children read once (duplicate sitemaps and URLs skipped) |
| Non-existent host in the same run | – | recorded as error (DNS failure, not retried) in SITEMAP_REPORT; the run still succeeded with the other sites' URLs |
| Local test server | raw gzip file, plain-text sitemap, malformed XML, missing (404) child, index loop, 2-level nesting | all read; malformed XML recovered with the tolerant scan; 404 child reported; loop not followed twice |
These are small tests on a handful of sites, not a benchmark: sites with unusual setups can still fail — when they do, the reason is in SITEMAP_REPORT.
Use cases
- Find all PDFs on a website — set Only these file types to
pdf(anddocx) to list a site's documents: reports, forms, policies, manuals. - Documents for RAG / LLMs — feed the PDF and DOCX list straight into a PDF-to-Markdown converter (see Chaining below).
- SEO audits and content inventories — every indexed URL with
lastmod,changefreq,priorityand the sitemap it came from. - Change monitoring — schedule the Actor with Changed on or after to get only pages updated recently.
- Crawl seeding — use the URL list as start URLs for a scraper or crawler instead of crawling links.
How to use
- Add one or more websites (
https://www.example.gov) or sitemap URLs in Websites or sitemap URLs. - Optional: set Max URLs, Only these file types, a lastmod date or URL patterns.
- Click Start. URLs appear in the Dataset; the per-sitemap report and the document list are in the Key-value store.
Input example
{"startUrls": [{ "url": "https://www.fec.gov" },{ "url": "https://www.example.com/sitemap_index.xml" }],"maxUrls": 5000,"fileTypes": ["pdf", "docx"],"checkContentType": "off"}
Output example (one dataset item per URL)
{"url": "https://www.fec.gov/resources/cms-content/documents/policy-guidance/fecfrm1.pdf","fileType": "pdf","extension": "pdf","lastmod": "2023-08-17T10:00:00+00:00","changefreq": null,"priority": null,"sitemapUrl": "https://www.fec.gov/resources/cms-content/documents/sitemap_pdf.xml","inputUrl": "https://www.fec.gov","host": "www.fec.gov"}
With Confirm file type with HTTP HEAD enabled, items also get httpStatus, contentType, contentLength, finalUrl, contentTypeFileType and checkError.
Key-value store records
| Key | Content |
|---|---|
OUTPUT | Run summary: URLs saved, counts per file type, sitemaps read / failed / skipped, duplicates, per-input status |
SITEMAP_REPORT | One entry per robots.txt and sitemap: status (ok, not_found, error, skipped), HTTP status, type, gzip, URLs found/new, child sitemaps, attempts, error message |
DOC_TO_MARKDOWN_INPUT | {"urls": [{"url": "...pdf"}, ...]} — every PDF and DOCX URL saved in this run (URLs that failed the HEAD check are left out) |
Chaining: sitemap → PDF/DOCX → Markdown
DOC_TO_MARKDOWN_INPUT uses the exact input format of the PDF & DOCX to Markdown Actor (urls list). Example with the Apify Python client:
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")# 1) list a site's PDFs and DOCX filesrun = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input={"startUrls": [{"url": "https://www.fec.gov"}],"fileTypes": ["pdf", "docx"],"maxUrls": 20,})doc_input = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("DOC_TO_MARKDOWN_INPUT")["value"]# 2) convert them to Markdown / RAG chunksrun2 = client.actor("ingenious_quip_bxq/pdf-docx-to-markdown").call(run_input={**doc_input, "outputFormat": "chunks"})for item in client.dataset(run2["defaultDatasetId"]).iterate_items():print(item.get("fileName"), item.get("pageStart"), item.get("text", "")[:80])
No code? In Apify Console, add an Integration → Actor on this Actor's run and pass the key-value store record as input, or copy the DOC_TO_MARKDOWN_INPUT record into the other Actor's JSON input. Mind the second Actor's pricing (per document and per page) before sending hundreds of files.
Pricing
Pay per event:
| Event | Price |
|---|---|
| URL saved to the dataset (primary) | $0.0002 per URL (= $0.20 per 1,000 URLs) |
| Optional HEAD content-type check (off by default) | $0.0003 per URL checked |
| Actor start | Apify's default synthetic start event |
You pay only for unique URLs saved after filters. Duplicates, filtered URLs, sitemaps that fail to load and the per-sitemap report are free. Set Max URLs (or a spending limit on the run) to cap cost — the Actor stops at the limit.
Known limits
- Only URLs listed in sitemaps (or feeds) are found — this Actor does not crawl links. Sites without a sitemap return no URLs (reported as
no_sitemap_found). - File type comes from the URL extension unless the HEAD check is on; download handlers such as
/download?id=123are taggedhtmlwithout it. - Image and video URLs inside
<image:image>/<video:video>sitemap extensions are not extracted as separate rows. - Sites that block data-center IPs may need a proxy (Proxy setting).
- Sitemaps are read up to 100 MB compressed / 200 MB uncompressed each.
FAQ
Does it respect robots.txt? It reads robots.txt to find sitemaps. Sitemaps exist to be read by bots, so it fetches them; it doesn't fetch pages unless you enable the HEAD check.
Why do I see not_found entries? /sitemap.xml and /sitemap_index.xml are probed on every site; when a site doesn't have them, that is recorded as not_found, not as an error.
Is the output compatible with other tools? Yes — download the dataset as CSV, JSON, Excel or XML, or read it via the API.
License & source code
This Actor is open source under the GNU Affero General Public License v3.0 (AGPL-3.0). The full source code is public: https://github.com/xbox002000/sitemap-url-discovery
Built with the Apify Python SDK, httpx and defusedxml.
See the CHANGELOG.md for version history.