Website to Markdown for RAG & LLMs: Fast Crawler
Pricing
Pay per event
Website to Markdown for RAG & LLMs: Fast Crawler
Crawl any website and turn every page into clean Markdown for LLMs, RAG and AI agents. Fast HTTP crawler, main-content extraction, GFM tables, RAG chunks, llms.txt per domain, robots.txt respected. Flat price per page.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Rod Services
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What does Website to Markdown Crawler do?
Website to Markdown Crawler crawls any website and turns every page into clean Markdown for LLMs, RAG pipelines and AI agents. It follows links on the site, keeps only the main content (no menus, sidebars, footers or cookie banners), converts it to GitHub-flavoured Markdown with tables, code blocks and absolute links, and can split each page into RAG chunks with token estimates. At the end it writes an llms.txt and an llms-full.txt file per domain.
It is a fast HTTP crawler, no browser, so a 200-page documentation site takes about a minute. JavaScript-only pages can optionally be rendered in headless Chrome. You pay a flat price per page, not for compute.
Try it: the prefilled input crawls 10 pages of Apify Academy in under 30 seconds. Run it from the Console, call it from the API, schedule it, or connect it to LangChain, LlamaIndex, Pinecone, Qdrant, Make, n8n or Zapier through Apify integrations.
Why use it for LLM and RAG data?
- Knowledge base for a chatbot or AI agent. Crawl your docs, help center or blog and load the chunks into a vector database.
- LLM training and fine-tuning data. Clean Markdown without boilerplate, one record per page, with language and word count.
- llms.txt for your own site. Get a ready
llms.txtindex and anllms-full.txtfile that AI assistants can read. - Keep an index fresh. Schedule weekly runs;
contentHashandlastModifiedshow which pages changed. - Predictable cost. $0.30 per 1,000 pages over HTTP. No compute units to estimate.
How to crawl a website to Markdown
- Open the Input tab and paste one or more start URLs, for example
https://docs.example.com/. - Set Max pages. This is also your cost cap.
- Optional: turn on Split into RAG chunks and pick a chunk size.
- Click Start. Pages appear in the Output tab while the crawl runs.
- Download the dataset as JSON, CSV, Excel or HTML, or open llms.txt files in the Storage tab.
By default the crawler stays under the start URL path: https://example.com/docs crawls /docs and everything below it.
Input
All options are on the Input tab. The most useful ones:
| Field | What it does | Default |
|---|---|---|
startUrls | Where the crawl starts | required |
maxPages | Stop after this many saved pages | 100 |
maxCrawlDepth | Links away from a start URL | 20 |
includeGlobs / excludeGlobs | URL patterns, ** matches any path | none |
sameDomainOnly / stayWithinStartPath | Crawl scope | on / on |
useSitemaps | Also queue URLs from sitemap.xml | off |
respectRobotsTxt | Skip pages disallowed by robots.txt | on |
extractionMode | readability, main or full page | readability |
removeSelectors | Extra CSS selectors to drop | none |
includeChunks, chunkSize, chunkOverlap | RAG chunks, sizes in estimated tokens | off, 500, 50 |
dedupeContent | Skip same canonical URL or identical content | on |
jsRenderingFallback | Render near-empty pages in headless Chrome | off |
maxConcurrency, maxConcurrencyPerDomain, delayBetweenRequestsMs | Politeness per domain | 10, 5, 0 |
{"startUrls": [{ "url": "https://docs.apify.com/academy" }],"maxPages": 10,"includeChunks": true,"chunkSize": 500,"chunkOverlap": 50}
Output
One dataset item per page. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has four views: Overview, Markdown, Metadata and RAG chunks.
{"url": "https://docs.apify.com/academy/api-scraping","canonicalUrl": "https://docs.apify.com/academy/api-scraping","title": "API scraping | Academy | Apify Documentation","description": "Learn all about how the professionals scrape various types of APIs with various configurations, parameters, and requirements.","language": "en","h1": "API scraping","lastModified": "2026-09-25T10:26:13.000Z","markdown": "# API scraping\n\nAPI scraping is locating a website's API endpoints, and fetching the desired data directly from their API...\n\n## What's an API?\n\n...","wordCount": 868,"chunkCount": 6,"chunks": [{"index": 0,"headings": "API scraping","text": "# API scraping\n\nAPI scraping is locating a website's API endpoints...","charCount": 1790,"tokenEstimate": 448}],"linksCount": 73,"depth": 1,"statusCode": 200,"fetchMode": "http","extractor": "readability","contentHash": "3f1c0a9e5b7d2c4e8a61","crawledAt": "2026-09-27T11:32:57.120Z","warning": null}
The key-value store also holds:
llms-<domain>.txt: an llms.txt index with the title, URL and description of every page.llms-full-<domain>.txt: all Markdown of the domain in one file, each page wrapped in<page url="..." title="...">. Split into parts above 8 MB.RUN-SUMMARY: saved, duplicate, blocked and failed counts, plus the first 500 problem URLs.
Data fields
| Field | Description |
|---|---|
url, canonicalUrl | Final URL after redirects and the rel=canonical URL |
title, description, h1 | Title tag, meta description, first heading |
language | From <html lang>, Content-Language or og:locale |
lastModified | From article:modified_time or the Last-Modified header |
markdown | Main content as GFM Markdown, links and images absolute |
wordCount, linksCount | Words in the Markdown, unique links on the page |
chunks, chunkCount | RAG chunks with heading path and token estimate |
depth, statusCode, fetchMode | Crawl depth, HTTP status, http or browser |
contentHash, crawledAt, warning | Dedupe hash, timestamp, note for thin or empty pages |
How much does it cost to convert a website to Markdown?
Pay per event, no compute charges:
| Event | Price |
|---|---|
| Actor start | $0.001 per run (per GB of memory) |
| Page crawled (HTTP) | $0.0003, that is $0.30 per 1,000 pages |
| Page rendered (Chrome, opt-in) | $0.006, that is $6 per 1,000 pages |
Duplicates, pages blocked by robots.txt, errors and empty pages are not charged. Examples: 10 pages cost $0.004, 1,000 pages $0.30, 10,000 pages $3. With the Apify free plan ($5 credit per month) you can convert about 16,000 pages a month. Set Max cost per run when you start a run and the crawler stops exactly at that budget.
Tips for faster and cheaper crawls
- Use include globs such as
https://example.com/docs/**to skip blogs, tags and login pages. - Turn on Also use sitemap.xml to find pages that are not linked from the menu.
- Keep the default 1024 MB memory for HTTP runs. Use 2048 MB when the JavaScript fallback is on.
- Leave the JavaScript fallback off unless the Overview view shows warnings like "No text content found".
- For small or fragile sites, lower Max concurrency per domain to 1 or 2 and add a delay.
- Use Remove CSS selectors for site-specific clutter, for example
.newsletter, #comments. - Chunks are split at headings, paragraphs and code blocks. 300 to 800 tokens with 10% overlap works well for most embedding models.
FAQ, disclaimers and support
Does it work on JavaScript sites? Most sites, including Docusaurus, GitBook, WordPress, MkDocs and Next.js sites, ship HTML and work over plain HTTP. Pure single-page apps return an empty shell; enable JavaScript rendering fallback for them.
Does it respect robots.txt? Yes, by default. Disallowed pages are skipped and listed in RUN-SUMMARY. You can turn it off only at your own responsibility; the run then logs a warning.
How are duplicates detected? By final URL, by rel=canonical and by a hash of the Markdown.
Which proxies can I use? No proxy (the default), Apify datacenter proxy, or your own proxy URLs. Residential and SERP proxies are not supported. A run that asks for them stops at the start with a clear message and does no work.
Is crawling legal? Crawling public pages is generally allowed, but you are responsible for complying with each site's terms, copyright and privacy law. Do not crawl personal data without a legal basis.
Known limits. No login or forms. PDF and other files are skipped. Hash-routed single-page apps (/#/page) show as one URL. Token counts are estimates (about 4 characters per token).
Found a bug or need a custom crawler, an export to your vector database or a scheduled pipeline? Open an issue in the Issues tab.