Website to Markdown for AI: Clean Content, RAG Chunks avatar

Website to Markdown for AI: Clean Content, RAG Chunks

Pricing

$1.00 / 1,000 page converteds

Go to Apify Store
Website to Markdown for AI: Clean Content, RAG Chunks

Website to Markdown for AI: Clean Content, RAG Chunks

Turn web pages or whole websites into clean Markdown for LLMs and RAG: main content only (no menus, ads or share buttons), title, author, date, word and token counts, optional chunks. Crawl a site or give a list. Fixed price per page.

Pricing

$1.00 / 1,000 page converteds

Rating

0.0

(0)

Developer

SwiftKit

SwiftKit

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 hours ago

Last modified

Share

Feed web pages to ChatGPT, Claude, your RAG pipeline or your AI agent without the junk. Give it a list of pages or a site to crawl, and get back:

  • Clean Markdown of the main content: headings, lists, tables, code blocks and links kept; menus, footers, cookie banners, ads and share buttons removed
  • Metadata: title, description, author, publish date, language, site name, image, canonical URL
  • Word count and approximate token count per page, so you can plan context windows and costs
  • Optional RAG chunks: pieces of about N tokens, split on paragraph boundaries with overlap
  • Fixed price per page. No compute-time surprises.

Main content is found with Mozilla Readability, the engine behind Firefox Reader View. Links and images are made absolute, so the Markdown works anywhere.

Who it's for

  • RAG and chatbots: index docs, help centers and blogs as clean chunks.
  • AI agents: read any page as Markdown with one call.
  • Content teams and researchers: archive or analyze articles in a readable format.
  • Training and evaluation data: clean text with word counts and dates.

Input

OptionDefaultWhat it does
URLs–Pages to convert, or start pages to crawl
Crawl the siteoffFollow links on the same site
Max pages50Stop after this many converted pages
Max link depth3How far from the start pages to crawl
Only / skip URLs containing–Keep the crawl to e.g. /docs/
ContentMain contentOr the whole page minus menus and footer
Chunk size, overlap0, 50Tokens per RAG chunk (0 = off)
Include plain text, linksoffExtra fields
Respect robots.txtonSkip pages the site asks bots not to visit

Output

Real result, shortened:

{
"url": "https://github.blog/engineering/the-cost-of-saying-yes-has-changed/",
"status": "ok",
"title": "The cost of saying yes has changed",
"author": "Dalia Abuadas",
"published": "2026-07-17T16:46:47+00:00",
"language": "en-US",
"markdown": "The cost of writing code dropped; the cost of owning it didn’t. …",
"wordCount": 1174,
"approxTokens": 1877,
"chunks": [
{ "index": 0, "text": "…", "approxTokens": "…" }
]
}
statusmeaning
okConverted.
unreachableThe page didn't load. Not charged.
not_htmlNot a web page (PDF, image…). Not charged.
blocked_by_robots_txtThe site asks bots not to visit. Not charged.
blocked_by_bot_protectionThe site showed a bot check. We don't try to get around it. Not charged.

Pricing

You pay per page converted, the same for every page size. Failed and blocked pages are free. See the Pricing tab.

Limits, honestly

  • No browser is used. That keeps it fast and cheap, but pages that build their content with JavaScript after loading (some single-page apps) may come back thin, and crawling them can find few links. Most blogs, docs, news and company sites work well.
  • Token counts are approximate (about 4 characters per token). Exact counts depend on your model.
  • Pages behind a login, a paywall or a bot check aren't converted.
  • You're responsible for how you use the content; respect each site's terms and copyright.

More tools from SwiftKit

Questions?

Open an issue on the Issues tab.