Sitemap URL Extractor — PDF/DOCX Tags + robots.txt avatar

Sitemap URL Extractor — PDF/DOCX Tags + robots.txt

Pricing

from $0.20 / 1,000 url discovereds

Go to Apify Store
Sitemap URL Extractor — PDF/DOCX Tags + robots.txt

Sitemap URL Extractor — PDF/DOCX Tags + robots.txt

Discover every URL in XML sitemaps via robots.txt and indexes; tag PDF, DOCX, XLSX and more for SEO audits and RAG ingest. Auto-finds nested indexes and .xml.gz; exports DOC_TO_MARKDOWN_INPUT for the PDF/DOCX Actor.

Pricing

from $0.20 / 1,000 url discovereds

Rating

0.0

(0)

Developer

新世紀書僮

新世紀書僮

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Sitemap URL Extractor — Find PDF & Document Links

Get every URL a website publishes in its XML sitemaps — with lastmod, source sitemap and a file-type tag, so PDFs, Word, Excel and PowerPoint files stand out. Enter a homepage and the Actor finds the sitemaps for you (robots.txt Sitemap: lines plus /sitemap.xml and /sitemap_index.xml), follows nested sitemap indexes, reads gzipped .xml.gz and plain-text sitemaps, removes duplicates and saves one row per URL. Use it for SEO audits, content inventories, change monitoring, or to collect a site's PDF and DOCX documents for RAG and LLM pipelines.

What you get

  • 🔎 Automatic sitemap discovery — robots.txt Sitemap: lines + common paths; or paste sitemap URLs directly
  • 🌳 Sitemap index recursion — nested indexes up to a depth you choose, loops and repeats skipped
  • 🗜️ All common formats — XML urlset and sitemapindex, .xml.gz (gzip file or gzip transfer encoding), plain-text sitemaps, RSS/Atom feeds
  • 🏷️ File-type tag per URL — pdf, docx, doc, xlsx, xls, pptx, ppt, csv, html, image, other (from the extension; optional HTTP HEAD check confirms the real Content-Type)
  • 📄 Ready-made input for PDF/DOCX conversion — all PDF and DOCX URLs are saved as DOC_TO_MARKDOWN_INPUT, formatted as input for the PDF & DOCX to Markdown Actor
  • 🧯 Broken sitemaps don't break the run — every sitemap gets its own status line (ok, not_found, error, skipped) with the reason in SITEMAP_REPORT; one bad sitemap or site never fails the whole run
  • 🐢 Polite and robust — low per-host concurrency, timeouts, retries with backoff for timeouts / 429 / 5xx (honours Retry-After), size caps, tolerant parsing of malformed XML
  • 🧮 Filters — file types, lastmod date, include/exclude regex, same domain only, max URLs (total and per site)
  • 💾 Light — 256 MB memory by default; HTTP only, no browser

Measured results

Tested by us on 2026-09-29 with small limits (local runs plus 3 runs on the Apify platform):

TestSitemap setupResult
fec.gov (US government)robots.txt lists 3 sitemaps incl. a PDF-only sitemap69 PDF URLs tagged pdf; HEAD check on all 69: HTTP 200 application/pdf (69/69)
gov.uksitemap index → 5.6 MB child sitemap20,000 URLs in 11 s on the Apify platform at 256 MB (peak memory 107 MB)
nist.govDrupal sitemap index with 56 child sitemapsindex followed, URLs saved up to the limit
developer.mozilla.orgindex of 10 .xml.gz sitemapsgzip sitemaps read
yoast.comYoast SEO index with 21 child sitemaps; /sitemap.xml is a second index to the same fileschildren read once (duplicate sitemaps and URLs skipped)
Non-existent host in the same run–recorded as error (DNS failure, not retried) in SITEMAP_REPORT; the run still succeeded with the other sites' URLs
Local test serverraw gzip file, plain-text sitemap, malformed XML, missing (404) child, index loop, 2-level nestingall read; malformed XML recovered with the tolerant scan; 404 child reported; loop not followed twice

These are small tests on a handful of sites, not a benchmark: sites with unusual setups can still fail — when they do, the reason is in SITEMAP_REPORT.

Use cases

  • Find all PDFs on a website — set Only these file types to pdf (and docx) to list a site's documents: reports, forms, policies, manuals.
  • Documents for RAG / LLMs — feed the PDF and DOCX list straight into a PDF-to-Markdown converter (see Chaining below).
  • SEO audits and content inventories — every indexed URL with lastmod, changefreq, priority and the sitemap it came from.
  • Change monitoring — schedule the Actor with Changed on or after to get only pages updated recently.
  • Crawl seeding — use the URL list as start URLs for a scraper or crawler instead of crawling links.

How to use

  1. Add one or more websites (https://www.example.gov) or sitemap URLs in Websites or sitemap URLs.
  2. Optional: set Max URLs, Only these file types, a lastmod date or URL patterns.
  3. Click Start. URLs appear in the Dataset; the per-sitemap report and the document list are in the Key-value store.

Input example

{
"startUrls": [
{ "url": "https://www.fec.gov" },
{ "url": "https://www.example.com/sitemap_index.xml" }
],
"maxUrls": 5000,
"fileTypes": ["pdf", "docx"],
"checkContentType": "off"
}

Output example (one dataset item per URL)

{
"url": "https://www.fec.gov/resources/cms-content/documents/policy-guidance/fecfrm1.pdf",
"fileType": "pdf",
"extension": "pdf",
"lastmod": "2023-08-17T10:00:00+00:00",
"changefreq": null,
"priority": null,
"sitemapUrl": "https://www.fec.gov/resources/cms-content/documents/sitemap_pdf.xml",
"inputUrl": "https://www.fec.gov",
"host": "www.fec.gov"
}

With Confirm file type with HTTP HEAD enabled, items also get httpStatus, contentType, contentLength, finalUrl, contentTypeFileType and checkError.

Key-value store records

KeyContent
OUTPUTRun summary: URLs saved, counts per file type, sitemaps read / failed / skipped, duplicates, per-input status
SITEMAP_REPORTOne entry per robots.txt and sitemap: status (ok, not_found, error, skipped), HTTP status, type, gzip, URLs found/new, child sitemaps, attempts, error message
DOC_TO_MARKDOWN_INPUT{"urls": [{"url": "...pdf"}, ...]} — every PDF and DOCX URL saved in this run (URLs that failed the HEAD check are left out)

Chaining: sitemap → PDF/DOCX → Markdown

DOC_TO_MARKDOWN_INPUT uses the exact input format of the PDF & DOCX to Markdown Actor (urls list). Example with the Apify Python client:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
# 1) list a site's PDFs and DOCX files
run = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input={
"startUrls": [{"url": "https://www.fec.gov"}],
"fileTypes": ["pdf", "docx"],
"maxUrls": 20,
})
doc_input = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("DOC_TO_MARKDOWN_INPUT")["value"]
# 2) convert them to Markdown / RAG chunks
run2 = client.actor("ingenious_quip_bxq/pdf-docx-to-markdown").call(run_input={**doc_input, "outputFormat": "chunks"})
for item in client.dataset(run2["defaultDatasetId"]).iterate_items():
print(item.get("fileName"), item.get("pageStart"), item.get("text", "")[:80])

No code? In Apify Console, add an Integration → Actor on this Actor's run and pass the key-value store record as input, or copy the DOC_TO_MARKDOWN_INPUT record into the other Actor's JSON input. Mind the second Actor's pricing (per document and per page) before sending hundreds of files.

Pricing

Pay per event:

EventPrice
URL saved to the dataset (primary)$0.0002 per URL (= $0.20 per 1,000 URLs)
Optional HEAD content-type check (off by default)$0.0003 per URL checked
Actor startApify's default synthetic start event

You pay only for unique URLs saved after filters. Duplicates, filtered URLs, sitemaps that fail to load and the per-sitemap report are free. Set Max URLs (or a spending limit on the run) to cap cost — the Actor stops at the limit.

Known limits

  • Only URLs listed in sitemaps (or feeds) are found — this Actor does not crawl links. Sites without a sitemap return no URLs (reported as no_sitemap_found).
  • File type comes from the URL extension unless the HEAD check is on; download handlers such as /download?id=123 are tagged html without it.
  • Image and video URLs inside <image:image> / <video:video> sitemap extensions are not extracted as separate rows.
  • Sites that block data-center IPs may need a proxy (Proxy setting).
  • Sitemaps are read up to 100 MB compressed / 200 MB uncompressed each.

FAQ

Does it respect robots.txt? It reads robots.txt to find sitemaps. Sitemaps exist to be read by bots, so it fetches them; it doesn't fetch pages unless you enable the HEAD check.

Why do I see not_found entries? /sitemap.xml and /sitemap_index.xml are probed on every site; when a site doesn't have them, that is recorded as not_found, not as an error.

Is the output compatible with other tools? Yes — download the dataset as CSV, JSON, Excel or XML, or read it via the API.

License & source code

This Actor is open source under the GNU Affero General Public License v3.0 (AGPL-3.0). The full source code is public: https://github.com/xbox002000/sitemap-url-discovery

Built with the Apify Python SDK, httpx and defusedxml.

See the CHANGELOG.md for version history.