Website to Markdown Crawler for AI & RAG avatar

Website to Markdown Crawler for AI & RAG

Pricing

Pay per event

Go to Apify Store
Website to Markdown Crawler for AI & RAG

Website to Markdown Crawler for AI & RAG

Crawl a website or docs site and get clean, LLM-ready Markdown for every page, with optional RAG chunks. Handles PDFs and Word files too. Fast mode from $1 per 1,000 pages.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Yukai Lin

Yukai Lin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

What does Website to Markdown Crawler do?

It crawls a website, documentation portal, blog or knowledge base and returns clean, LLM-ready Markdown for every page, ready for ChatGPT / Claude context, vector databases and RAG pipelines.

  • 🕷️ Crawls for you: start from one URL and follow links within the same folder, the whole site, or only the pages you list
  • 📝 Clean Markdown: headings, lists, tables, links and code blocks preserved; menus can be dropped with a CSS selector
  • 📄 Documents too: linked PDF, Word (DOCX), Excel (XLSX) and CSV files are converted to Markdown as well
  • ✂️ RAG chunks built in: optional chunks array split at paragraph boundaries with overlap, ready for embeddings
  • ⚡ Fast and cheap: pages are fetched over plain HTTP whenever possible ($1 per 1,000 pages) and rendered in a real browser only when a site needs JavaScript or blocks simple requests
  • 🧹 No duplicates, no surprises: identical pages reachable under several URLs are kept once and charged once; blocked and failed pages are free

Who is it for?

  • Building a RAG chatbot or AI assistant over your docs, help center or website
  • Feeding LLM agents with up-to-date documentation
  • Creating a knowledge base export or content audit
  • Migrating a website's content to another CMS

How much does it cost?

Pay per event, only for pages that were converted successfully:

EventPrice
Page (fast mode, plain HTTP)$1.00 / 1,000 pages
Page (browser mode, JavaScript rendering)$2.50 / 1,000 pages

In Auto mode (default) most pages use the fast mode. Duplicates, blocked pages (403, bot checks) and errors are not charged. Your maximum charge limit is always respected.

How to use it

  1. Enter one or more Start URLs, e.g. https://docs.example.com/.
  2. Choose the Crawl scope (same folder is best for documentation).
  3. Set Max pages.
  4. Optional: set a Content selector such as main or article, URL include/exclude patterns, and a RAG chunk size (e.g. 2000).
  5. Click Start and download the results as JSON, CSV or Excel, or fetch them via API.

Input example

{
"startUrls": [{ "url": "https://docs.apify.com/academy" }],
"crawlScope": "path",
"maxPages": 200,
"mode": "auto",
"cssSelector": "main",
"excludeUrlPatterns": ["**/changelog/**"],
"chunkSize": 2000,
"chunkOverlap": 200
}

Output example

{
"url": "https://docs.apify.com/academy/web-scraping-for-beginners",
"finalUrl": "https://docs.apify.com/academy/web-scraping-for-beginners",
"title": "Web scraping basics for JavaScript devs | Academy",
"depth": 0,
"httpStatus": 200,
"contentType": "text/html; charset=utf-8",
"mode": "fast",
"success": true,
"markdown": "# Web scraping basics for JavaScript devs\n\nLearn how to...",
"chunks": [{ "index": 0, "text": "# Web scraping basics..." }],
"crawledAt": "2026-09-29T04:00:00.000Z"
}

Use it from AI agents and code

The Actor works well as a tool for AI agents (via the Apify MCP server, LangChain, LlamaIndex, or the Apify API): give it a URL and a page limit, get Markdown back. Example with the Apify API:

curl -X POST "https://api.apify.com/v2/acts/tidytools~website-markdown-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls":[{"url":"https://docs.example.com"}],"maxPages":20}'

Tips

  • Documentation sites: keep the scope on Same folder and set the content selector to main or article to drop navigation.
  • JavaScript-heavy apps (React, Vue, Angular): Auto mode switches to the browser automatically; choose Browser to force it.
  • Only part of a site: use include patterns like https://example.com/blog/**.
  • Big sites: raise Max pages and Parallel pages; you only pay for pages converted.

Limitations

  • Only public http/https pages; logins, private networks and local addresses are not supported.
  • Sites that block automated access are reported as failed and not charged.
  • Linked documents up to 15 MB are converted; images are not described.

Crawling publicly available pages is generally allowed, but you are responsible for how you use the content. Respect the target site's terms of service, robots rules, copyright and privacy laws.

Support

Open an issue in the Issues tab with the URL and your input. Issues are checked regularly.