Website to Markdown Crawler for AI & RAG
Pricing
Pay per event
Website to Markdown Crawler for AI & RAG
Crawl a website or docs site and get clean, LLM-ready Markdown for every page, with optional RAG chunks. Handles PDFs and Word files too. Fast mode from $1 per 1,000 pages.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Yukai Lin
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
What does Website to Markdown Crawler do?
It crawls a website, documentation portal, blog or knowledge base and returns clean, LLM-ready Markdown for every page, ready for ChatGPT / Claude context, vector databases and RAG pipelines.
- 🕷️ Crawls for you: start from one URL and follow links within the same folder, the whole site, or only the pages you list
- 📝 Clean Markdown: headings, lists, tables, links and code blocks preserved; menus can be dropped with a CSS selector
- 📄 Documents too: linked PDF, Word (DOCX), Excel (XLSX) and CSV files are converted to Markdown as well
- ✂️ RAG chunks built in: optional
chunksarray split at paragraph boundaries with overlap, ready for embeddings - ⚡ Fast and cheap: pages are fetched over plain HTTP whenever possible ($1 per 1,000 pages) and rendered in a real browser only when a site needs JavaScript or blocks simple requests
- 🧹 No duplicates, no surprises: identical pages reachable under several URLs are kept once and charged once; blocked and failed pages are free
Who is it for?
- Building a RAG chatbot or AI assistant over your docs, help center or website
- Feeding LLM agents with up-to-date documentation
- Creating a knowledge base export or content audit
- Migrating a website's content to another CMS
How much does it cost?
Pay per event, only for pages that were converted successfully:
| Event | Price |
|---|---|
| Page (fast mode, plain HTTP) | $1.00 / 1,000 pages |
| Page (browser mode, JavaScript rendering) | $2.50 / 1,000 pages |
In Auto mode (default) most pages use the fast mode. Duplicates, blocked pages (403, bot checks) and errors are not charged. Your maximum charge limit is always respected.
How to use it
- Enter one or more Start URLs, e.g.
https://docs.example.com/. - Choose the Crawl scope (same folder is best for documentation).
- Set Max pages.
- Optional: set a Content selector such as
mainorarticle, URL include/exclude patterns, and a RAG chunk size (e.g. 2000). - Click Start and download the results as JSON, CSV or Excel, or fetch them via API.
Input example
{"startUrls": [{ "url": "https://docs.apify.com/academy" }],"crawlScope": "path","maxPages": 200,"mode": "auto","cssSelector": "main","excludeUrlPatterns": ["**/changelog/**"],"chunkSize": 2000,"chunkOverlap": 200}
Output example
{"url": "https://docs.apify.com/academy/web-scraping-for-beginners","finalUrl": "https://docs.apify.com/academy/web-scraping-for-beginners","title": "Web scraping basics for JavaScript devs | Academy","depth": 0,"httpStatus": 200,"contentType": "text/html; charset=utf-8","mode": "fast","success": true,"markdown": "# Web scraping basics for JavaScript devs\n\nLearn how to...","chunks": [{ "index": 0, "text": "# Web scraping basics..." }],"crawledAt": "2026-09-29T04:00:00.000Z"}
Use it from AI agents and code
The Actor works well as a tool for AI agents (via the Apify MCP server, LangChain, LlamaIndex, or the Apify API): give it a URL and a page limit, get Markdown back. Example with the Apify API:
curl -X POST "https://api.apify.com/v2/acts/tidytools~website-markdown-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"startUrls":[{"url":"https://docs.example.com"}],"maxPages":20}'
Tips
- Documentation sites: keep the scope on Same folder and set the content selector to
mainorarticleto drop navigation. - JavaScript-heavy apps (React, Vue, Angular): Auto mode switches to the browser automatically; choose Browser to force it.
- Only part of a site: use include patterns like
https://example.com/blog/**. - Big sites: raise Max pages and Parallel pages; you only pay for pages converted.
Limitations
- Only public
http/httpspages; logins, private networks and local addresses are not supported. - Sites that block automated access are reported as failed and not charged.
- Linked documents up to 15 MB are converted; images are not described.
Is it legal?
Crawling publicly available pages is generally allowed, but you are responsible for how you use the content. Respect the target site's terms of service, robots rules, copyright and privacy laws.
Support
Open an issue in the Issues tab with the URL and your input. Issues are checked regularly.