Sitemap Content Extractor — Website to Clean Markdown avatar

Sitemap Content Extractor — Website to Clean Markdown

Pricing

from $0.50 / 1,000 page extracteds

Go to Apify Store
Sitemap Content Extractor — Website to Clean Markdown

Sitemap Content Extractor — Website to Clean Markdown

Turn any website's sitemap into clean, LLM-ready Markdown. Auto-discovers sitemaps via robots.txt, extracts article text, tables, headings, links and images while stripping navigation, ads and cookie banners. Filter pages by lastmod date.

Pricing

from $0.50 / 1,000 page extracteds

Rating

0.0

(0)

Developer

XiaoZhi DataTools

XiaoZhi DataTools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Share

Turn any website's sitemap into clean, LLM-ready Markdown. Give it a domain (or a sitemap URL) and get back the full text content of every page — no browser, no login, no residential proxies, pure HTTP.

What it does

  1. Discovers sitemaps — reads robots.txt Sitemap: directives and probes common locations (/sitemap.xml, /sitemap_index.xml, …). You can also pass sitemap URLs directly.
  2. Expands sitemap indexes — recursively follows nested sitemaps (multi-level indexes, gzipped .xml.gz, plain-text sitemaps).
  3. Fetches every page concurrently — honors robots.txt Disallow rules, follows redirects, skips non-HTML content, retries transient failures.
  4. Extracts clean content — strips navigation, headers, footers, sidebars, scripts and cookie banners; keeps the main article content as Markdown with headings, links (absolute URLs), images and tables rendered as Markdown pipe tables (| a | b |). Headings are detected even when nested inside link cards or custom web components (common on news and React/SSR sites).

Input

FieldTypeDefaultDescription
startUrlsstring[]—Domains or sitemap URLs, e.g. apify.com, https://example.com/sitemap.xml
maxPagesinteger100Max pages to fetch and extract (up to 50,000)
maxSitemapDepthinteger3Nested sitemap-index levels to expand
urlPatternIncludestring[]—Only keep URLs matching these regexes (one per line), e.g. /docs/
urlPatternExcludestring[]—Skip URLs matching these regexes, e.g. /tag/
minLastmodstring—Only keep URLs with sitemap lastmod on/after this date (YYYY-MM-DD); URLs without lastmod are kept
maxLastmodstring—Only keep URLs with sitemap lastmod on/before this date (YYYY-MM-DD); URLs without lastmod are kept
outputFormatmarkdown/text/htmlmarkdownmarkdown (best for LLMs), text, or cleaned html
includeMetadatabooleantrueTitle, meta description, language, headings, outbound links, image counts
maxContentCharsinteger0Truncate content per page (0 = no limit)
limitToStartDomainbooleantrueOnly fetch URLs on the start domain
respectRobotsTxtbooleantrueHonor robots.txt Disallow rules
requestConcurrencyinteger10Parallel page fetches (1–30)
proxyConfigurationproxyoffOptional; not needed for typical runs

Output

One dataset item per page:

{
"url": "https://apify.com/about",
"finalUrl": "https://apify.com/about",
"statusCode": 200,
"sitemapUrl": "https://apify.com/sitemap/pages.xml",
"lastmod": "2026-09-20",
"changefreq": null,
"priority": null,
"contentFormat": "markdown",
"content": "# About Apify\n\nWe're building the world's largest platform...",
"contentChars": 4821,
"wordCount": 548,
"title": "About Apify",
"metaDescription": "...",
"language": "en",
"headings": [{ "level": 1, "text": "About Apify" }],
"links": ["https://docs.apify.com/legal"],
"linksCount": 31,
"imagesCount": 4,
"fetchedAt": "2026-09-28T10:00:00+00:00"
}

Use cases

  • RAG / LLM ingestion — dump an entire docs site or blog into clean Markdown for embeddings or fine-tuning datasets.
  • Site content audits — inventory every page with titles, descriptions, word counts and lastmod dates.
  • Competitor research — extract a competitor's full public content footprint from their sitemap.
  • SEO analysis — pair sitemap metadata (lastmod, priority) with actual on-page headings and word counts.
  • Change monitoring — re-run on a schedule and diff content per URL.

Cost & performance

  • Pure HTTP (no headless browser), so a run costs a fraction of browser-based crawlers — roughly $0.001–0.003 per 1,000 pages in Apify compute.
  • ~10 pages/second at default concurrency; 30 pages fetched and extracted in ~2 seconds in testing.
  • No residential proxies required: sitemaps and public pages are served to plain HTTP clients.

Limitations

  • Only pages listed in sitemaps are fetched (orphan pages not in any sitemap are missed) — this is by design: sitemap-driven, not link-crawling.
  • JavaScript-rendered content is not executed; pages that require JS to render body text will yield thin results.
  • robots.txt Disallow rules are honored by default (can be disabled).
  • Very large pages are truncated at 5 MB download size.
  • Non-HTML responses (PDFs, images, feeds) are skipped.

Changelog

  • 0.5 — Fixed gzipped sitemaps (.xml.gz) served with application/gzip / application/octet-stream content types being silently skipped; added retries with backoff for transient HTTP statuses (429, 500, 502, 503, 504) instead of failing immediately.
  • 0.4 — Previous release.