Website to Markdown Crawler for AI, RAG & LLMs
Pricing
from $1.00 / 1,000 pages
Website to Markdown Crawler for AI, RAG & LLMs
Crawl any website and get clean main-content Markdown for every page: no menus or footers. RAG-ready chunks, llms.txt, and an 'only changed pages' mode for cheap scheduled refreshes. $1 per 1,000 pages, all included.
Pricing
from $1.00 / 1,000 pages
Rating
0.0
(0)
Developer
Cronexa Data Tools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What does Website to Markdown Crawler do?
It crawls any website and turns every page into clean Markdown, keeping only the main content without menus, headers, footers, cookie banners or sidebars. The output is ready for AI, RAG pipelines, vector databases, LLM fine-tuning and custom GPTs.
- 📝 Clean main-content Markdown, with headings, lists, tables, links and fenced code blocks kept
- ✂️ RAG-ready chunks (optional): each page split into chunks of your chosen token size, cut at headings and paragraphs (never inside a code block), with overlap
- 📚
llms.txtandllms-full.txtgenerated for the site, following the llmstxt.org format - 🔄 "Only new or changed pages" mode: schedule it daily or weekly, and only pages whose content changed are saved. Unchanged pages are free, so keeping your RAG index fresh costs almost nothing.
- 💸 $1 per 1,000 pages. Everything included: no compute, proxy or browser fees on top.
It's fast because it reads the HTML directly instead of starting a full browser. It follows links and the sitemap, respects robots.txt, and slows down automatically if a site asks it to.
What can I use it for?
- Chatbots and RAG: feed product docs, help centers or knowledge bases into your vector database (Pinecone, Qdrant, Weaviate, pgvector…)
- Custom GPTs / Claude Projects: download
llms-full.txtand upload one file with a whole documentation site - Keeping AI knowledge up to date: schedule refresh runs and re-index only the pages that changed
- Content migration and archiving: move a website's content to Markdown for a new CMS, Notion or Obsidian
- SEO and content analysis: word counts, titles and descriptions for every page
How do I use it?
- Enter a start URL, for example
https://docs.example.com/. - (Optional) Limit it with Only crawl URLs matching, for example
/docs/. - (Optional) Set Chunk size (for example
800) if you're loading a vector database. - Click Start, then download the pages as JSON, CSV or Excel, or open
llms-full.txtfrom the Output tab.
For scheduled refreshes, turn on Only new or changed pages and add a Schedule. Each run saves only what changed, and marks each row as new or changed.
Example input
{"startUrls": [{ "url": "https://www.python-httpx.org/" }],"maxPages": 500,"includeUrlPatterns": [],"chunkSize": 800,"chunkOverlap": 100,"onlyChangedPages": true}
Output example
One row per page (Markdown and chunk text shortened here):
{"url": "https://www.python-httpx.org/quickstart/","title": "QuickStart - HTTPX","description": "A next-generation HTTP client for Python.","language": "en","markdown": "# QuickStart\n\nFirst, start by importing HTTPX:\n\n```\n>>> import httpx\n```\n\nNow, let’s try to get a webpage.\n\n```\n>>> r = httpx.get('https://httpbin.org/get')\n...","wordCount": 1742,"tokenEstimate": 3627,"extraction": "main-content","contentHash": "b3e8cef33b8b14e49d1e6d98bba23e7420af901b20ce34365bd53a34a181659d","changeStatus": "new","httpStatus": 200,"chunks": [{ "index": 0, "heading": "QuickStart", "text": "# QuickStart\n\nFirst, start by importing HTTPX: ...", "tokenEstimate": 797 }]}
The Output tab also links llms.txt (an index of all pages) and llms-full.txt (all pages in one Markdown file).
How much does it cost?
| What | Price |
|---|---|
| Page saved | $0.001 ($1 per 1,000 pages) |
| Unchanged pages in refresh mode | Free |
| Chunks, llms.txt, sitemap, robots.txt | Free |
| Platform usage (compute) | Included |
A 500-page documentation site costs $0.50. You can set a maximum cost per run, and the Actor stops exactly at your budget.
Tips
- Documentation sites: use Only crawl URLs matching (for example
/docs/) to skip the blog and marketing pages. - Chunk size: 500–1,000 tokens works well for most embedding models. The token estimate uses ~4 characters per token.
- Plain text: turn off Keep links in Markdown, or use the
textfield.
Limitations
- Pages are read from their HTML. Content that appears only after JavaScript runs (some single-page apps) may be incomplete. Most docs, blogs and company sites work well.
- Login-protected pages are not crawled.
- Some websites block crawlers. Try the Proxy option if pages fail.
Is it legal?
It collects publicly available web pages and respects robots.txt by default. Make sure your use of the content respects copyright and each website's terms.
Questions or problems?
Open an issue in the Issues tab, and it will be answered quickly.
Use it from AI assistants (Claude, ChatGPT, Cursor)
AI agents can run this Actor as a tool through the Apify MCP server. Add this to your MCP client (Claude Desktop, Claude Code, Cursor, VS Code…) and sign in with Apify in the browser when asked:
{"mcpServers": {"website-to-markdown": { "url": "https://mcp.apify.com?tools=dima_kadirovich/website-to-markdown" }}}
Then just ask, for example:
- "Read the docs at https://www.python-httpx.org/ (up to 30 pages) and explain how to set timeouts."
- "Turn this help center into Markdown chunks of 800 tokens for my vector database."
The agent fills in the input, runs the Actor and reads the results. You pay the same per-result price.
More tools from Dima Data Tools
- Medium Articles Scraper & Monitor: Medium articles by tag, author, or publication as clean Markdown, with "only new" monitoring
- Bulk Image Downloader: every image from any web page, as download links or ZIP, with duplicates and icons removed
- Website SEO Audit & Broken Link Checker: crawl a site, score every page 0–100, find broken links, and get a shareable HTML report
- Website Tech Stack & Domain Lookup: technologies, email provider, SPF/DMARC, SaaS tools, SSL expiry and WHOIS for any list of domains