Website to Markdown Crawler for AI, RAG & LLMs avatar

Website to Markdown Crawler for AI, RAG & LLMs

Pricing

from $1.00 / 1,000 pages

Go to Apify Store
Website to Markdown Crawler for AI, RAG & LLMs

Website to Markdown Crawler for AI, RAG & LLMs

Crawl any website and get clean main-content Markdown for every page: no menus or footers. RAG-ready chunks, llms.txt, and an 'only changed pages' mode for cheap scheduled refreshes. $1 per 1,000 pages, all included.

Pricing

from $1.00 / 1,000 pages

Rating

0.0

(0)

Developer

Cronexa Data Tools

Cronexa Data Tools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does Website to Markdown Crawler do?

It crawls any website and turns every page into clean Markdown, keeping only the main content without menus, headers, footers, cookie banners or sidebars. The output is ready for AI, RAG pipelines, vector databases, LLM fine-tuning and custom GPTs.

  • 📝 Clean main-content Markdown, with headings, lists, tables, links and fenced code blocks kept
  • ✂️ RAG-ready chunks (optional): each page split into chunks of your chosen token size, cut at headings and paragraphs (never inside a code block), with overlap
  • 📚 llms.txt and llms-full.txt generated for the site, following the llmstxt.org format
  • 🔄 "Only new or changed pages" mode: schedule it daily or weekly, and only pages whose content changed are saved. Unchanged pages are free, so keeping your RAG index fresh costs almost nothing.
  • 💸 $1 per 1,000 pages. Everything included: no compute, proxy or browser fees on top.

It's fast because it reads the HTML directly instead of starting a full browser. It follows links and the sitemap, respects robots.txt, and slows down automatically if a site asks it to.

What can I use it for?

  • Chatbots and RAG: feed product docs, help centers or knowledge bases into your vector database (Pinecone, Qdrant, Weaviate, pgvector…)
  • Custom GPTs / Claude Projects: download llms-full.txt and upload one file with a whole documentation site
  • Keeping AI knowledge up to date: schedule refresh runs and re-index only the pages that changed
  • Content migration and archiving: move a website's content to Markdown for a new CMS, Notion or Obsidian
  • SEO and content analysis: word counts, titles and descriptions for every page

How do I use it?

  1. Enter a start URL, for example https://docs.example.com/.
  2. (Optional) Limit it with Only crawl URLs matching, for example /docs/.
  3. (Optional) Set Chunk size (for example 800) if you're loading a vector database.
  4. Click Start, then download the pages as JSON, CSV or Excel, or open llms-full.txt from the Output tab.

For scheduled refreshes, turn on Only new or changed pages and add a Schedule. Each run saves only what changed, and marks each row as new or changed.

Example input

{
"startUrls": [{ "url": "https://www.python-httpx.org/" }],
"maxPages": 500,
"includeUrlPatterns": [],
"chunkSize": 800,
"chunkOverlap": 100,
"onlyChangedPages": true
}

Output example

One row per page (Markdown and chunk text shortened here):

{
"url": "https://www.python-httpx.org/quickstart/",
"title": "QuickStart - HTTPX",
"description": "A next-generation HTTP client for Python.",
"language": "en",
"markdown": "# QuickStart\n\nFirst, start by importing HTTPX:\n\n```\n>>> import httpx\n```\n\nNow, let’s try to get a webpage.\n\n```\n>>> r = httpx.get('https://httpbin.org/get')\n...",
"wordCount": 1742,
"tokenEstimate": 3627,
"extraction": "main-content",
"contentHash": "b3e8cef33b8b14e49d1e6d98bba23e7420af901b20ce34365bd53a34a181659d",
"changeStatus": "new",
"httpStatus": 200,
"chunks": [
{ "index": 0, "heading": "QuickStart", "text": "# QuickStart\n\nFirst, start by importing HTTPX: ...", "tokenEstimate": 797 }
]
}

The Output tab also links llms.txt (an index of all pages) and llms-full.txt (all pages in one Markdown file).

How much does it cost?

WhatPrice
Page saved$0.001 ($1 per 1,000 pages)
Unchanged pages in refresh modeFree
Chunks, llms.txt, sitemap, robots.txtFree
Platform usage (compute)Included

A 500-page documentation site costs $0.50. You can set a maximum cost per run, and the Actor stops exactly at your budget.

Tips

  • Documentation sites: use Only crawl URLs matching (for example /docs/) to skip the blog and marketing pages.
  • Chunk size: 500–1,000 tokens works well for most embedding models. The token estimate uses ~4 characters per token.
  • Plain text: turn off Keep links in Markdown, or use the text field.

Limitations

  • Pages are read from their HTML. Content that appears only after JavaScript runs (some single-page apps) may be incomplete. Most docs, blogs and company sites work well.
  • Login-protected pages are not crawled.
  • Some websites block crawlers. Try the Proxy option if pages fail.

It collects publicly available web pages and respects robots.txt by default. Make sure your use of the content respects copyright and each website's terms.

Questions or problems?

Open an issue in the Issues tab, and it will be answered quickly.

Use it from AI assistants (Claude, ChatGPT, Cursor)

AI agents can run this Actor as a tool through the Apify MCP server. Add this to your MCP client (Claude Desktop, Claude Code, Cursor, VS Code…) and sign in with Apify in the browser when asked:

{
"mcpServers": {
"website-to-markdown": { "url": "https://mcp.apify.com?tools=dima_kadirovich/website-to-markdown" }
}
}

Then just ask, for example:

  • "Read the docs at https://www.python-httpx.org/ (up to 30 pages) and explain how to set timeouts."
  • "Turn this help center into Markdown chunks of 800 tokens for my vector database."

The agent fills in the input, runs the Actor and reads the results. You pay the same per-result price.

More tools from Dima Data Tools