Sitemap Content Extractor — Website to Clean Markdown
Pricing
from $0.50 / 1,000 page extracteds
Go to Apify Store
Sitemap Content Extractor — Website to Clean Markdown
Turn any website's sitemap into clean, LLM-ready Markdown. Auto-discovers sitemaps via robots.txt, extracts article text, tables, headings, links and images while stripping navigation, ads and cookie banners. Filter pages by lastmod date.
Pricing
from $0.50 / 1,000 page extracteds
Rating
0.0
(0)
Developer
XiaoZhi DataTools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
Turn any website's sitemap into clean, LLM-ready Markdown. Give it a domain (or a sitemap URL) and get back the full text content of every page — no browser, no login, no residential proxies, pure HTTP.
What it does
- Discovers sitemaps — reads
robots.txtSitemap:directives and probes common locations (/sitemap.xml,/sitemap_index.xml, …). You can also pass sitemap URLs directly. - Expands sitemap indexes — recursively follows nested sitemaps
(multi-level indexes, gzipped
.xml.gz, plain-text sitemaps). - Fetches every page concurrently — honors
robots.txtDisallow rules, follows redirects, skips non-HTML content, retries transient failures. - Extracts clean content — strips navigation, headers, footers, sidebars,
scripts and cookie banners; keeps the main article content as Markdown
with headings, links (absolute URLs), images and tables rendered as
Markdown pipe tables (
| a | b |). Headings are detected even when nested inside link cards or custom web components (common on news and React/SSR sites).
Input
| Field | Type | Default | Description |
|---|---|---|---|
startUrls | string[] | — | Domains or sitemap URLs, e.g. apify.com, https://example.com/sitemap.xml |
maxPages | integer | 100 | Max pages to fetch and extract (up to 50,000) |
maxSitemapDepth | integer | 3 | Nested sitemap-index levels to expand |
urlPatternInclude | string[] | — | Only keep URLs matching these regexes (one per line), e.g. /docs/ |
urlPatternExclude | string[] | — | Skip URLs matching these regexes, e.g. /tag/ |
minLastmod | string | — | Only keep URLs with sitemap lastmod on/after this date (YYYY-MM-DD); URLs without lastmod are kept |
maxLastmod | string | — | Only keep URLs with sitemap lastmod on/before this date (YYYY-MM-DD); URLs without lastmod are kept |
outputFormat | markdown/text/html | markdown | markdown (best for LLMs), text, or cleaned html |
includeMetadata | boolean | true | Title, meta description, language, headings, outbound links, image counts |
maxContentChars | integer | 0 | Truncate content per page (0 = no limit) |
limitToStartDomain | boolean | true | Only fetch URLs on the start domain |
respectRobotsTxt | boolean | true | Honor robots.txt Disallow rules |
requestConcurrency | integer | 10 | Parallel page fetches (1–30) |
proxyConfiguration | proxy | off | Optional; not needed for typical runs |
Output
One dataset item per page:
{"url": "https://apify.com/about","finalUrl": "https://apify.com/about","statusCode": 200,"sitemapUrl": "https://apify.com/sitemap/pages.xml","lastmod": "2026-09-20","changefreq": null,"priority": null,"contentFormat": "markdown","content": "# About Apify\n\nWe're building the world's largest platform...","contentChars": 4821,"wordCount": 548,"title": "About Apify","metaDescription": "...","language": "en","headings": [{ "level": 1, "text": "About Apify" }],"links": ["https://docs.apify.com/legal"],"linksCount": 31,"imagesCount": 4,"fetchedAt": "2026-09-28T10:00:00+00:00"}
Use cases
- RAG / LLM ingestion — dump an entire docs site or blog into clean Markdown for embeddings or fine-tuning datasets.
- Site content audits — inventory every page with titles, descriptions, word counts and lastmod dates.
- Competitor research — extract a competitor's full public content footprint from their sitemap.
- SEO analysis — pair sitemap metadata (lastmod, priority) with actual on-page headings and word counts.
- Change monitoring — re-run on a schedule and diff
contentper URL.
Cost & performance
- Pure HTTP (no headless browser), so a run costs a fraction of browser-based crawlers — roughly $0.001–0.003 per 1,000 pages in Apify compute.
- ~10 pages/second at default concurrency; 30 pages fetched and extracted in ~2 seconds in testing.
- No residential proxies required: sitemaps and public pages are served to plain HTTP clients.
Limitations
- Only pages listed in sitemaps are fetched (orphan pages not in any sitemap are missed) — this is by design: sitemap-driven, not link-crawling.
- JavaScript-rendered content is not executed; pages that require JS to render body text will yield thin results.
robots.txtDisallow rules are honored by default (can be disabled).- Very large pages are truncated at 5 MB download size.
- Non-HTML responses (PDFs, images, feeds) are skipped.
Changelog
- 0.5 — Fixed gzipped sitemaps (
.xml.gz) served withapplication/gzip/application/octet-streamcontent types being silently skipped; added retries with backoff for transient HTTP statuses (429, 500, 502, 503, 504) instead of failing immediately. - 0.4 — Previous release.