Website to Markdown for RAG & LLMs: Fast Crawler avatar

Website to Markdown for RAG & LLMs: Fast Crawler

Pricing

Pay per event

Go to Apify Store
Website to Markdown for RAG & LLMs: Fast Crawler

Website to Markdown for RAG & LLMs: Fast Crawler

Crawl any website and turn every page into clean Markdown for LLMs, RAG and AI agents. Fast HTTP crawler, main-content extraction, GFM tables, RAG chunks, llms.txt per domain, robots.txt respected. Flat price per page.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Rod Services

Rod Services

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does Website to Markdown Crawler do?

Website to Markdown Crawler crawls any website and turns every page into clean Markdown for LLMs, RAG pipelines and AI agents. It follows links on the site, keeps only the main content (no menus, sidebars, footers or cookie banners), converts it to GitHub-flavoured Markdown with tables, code blocks and absolute links, and can split each page into RAG chunks with token estimates. At the end it writes an llms.txt and an llms-full.txt file per domain.

It is a fast HTTP crawler, no browser, so a 200-page documentation site takes about a minute. JavaScript-only pages can optionally be rendered in headless Chrome. You pay a flat price per page, not for compute.

Try it: the prefilled input crawls 10 pages of Apify Academy in under 30 seconds. Run it from the Console, call it from the API, schedule it, or connect it to LangChain, LlamaIndex, Pinecone, Qdrant, Make, n8n or Zapier through Apify integrations.

Why use it for LLM and RAG data?

  • Knowledge base for a chatbot or AI agent. Crawl your docs, help center or blog and load the chunks into a vector database.
  • LLM training and fine-tuning data. Clean Markdown without boilerplate, one record per page, with language and word count.
  • llms.txt for your own site. Get a ready llms.txt index and an llms-full.txt file that AI assistants can read.
  • Keep an index fresh. Schedule weekly runs; contentHash and lastModified show which pages changed.
  • Predictable cost. $0.30 per 1,000 pages over HTTP. No compute units to estimate.

How to crawl a website to Markdown

  1. Open the Input tab and paste one or more start URLs, for example https://docs.example.com/.
  2. Set Max pages. This is also your cost cap.
  3. Optional: turn on Split into RAG chunks and pick a chunk size.
  4. Click Start. Pages appear in the Output tab while the crawl runs.
  5. Download the dataset as JSON, CSV, Excel or HTML, or open llms.txt files in the Storage tab.

By default the crawler stays under the start URL path: https://example.com/docs crawls /docs and everything below it.

Input

All options are on the Input tab. The most useful ones:

FieldWhat it doesDefault
startUrlsWhere the crawl startsrequired
maxPagesStop after this many saved pages100
maxCrawlDepthLinks away from a start URL20
includeGlobs / excludeGlobsURL patterns, ** matches any pathnone
sameDomainOnly / stayWithinStartPathCrawl scopeon / on
useSitemapsAlso queue URLs from sitemap.xmloff
respectRobotsTxtSkip pages disallowed by robots.txton
extractionModereadability, main or full pagereadability
removeSelectorsExtra CSS selectors to dropnone
includeChunks, chunkSize, chunkOverlapRAG chunks, sizes in estimated tokensoff, 500, 50
dedupeContentSkip same canonical URL or identical contenton
jsRenderingFallbackRender near-empty pages in headless Chromeoff
maxConcurrency, maxConcurrencyPerDomain, delayBetweenRequestsMsPoliteness per domain10, 5, 0
{
"startUrls": [{ "url": "https://docs.apify.com/academy" }],
"maxPages": 10,
"includeChunks": true,
"chunkSize": 500,
"chunkOverlap": 50
}

Output

One dataset item per page. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has four views: Overview, Markdown, Metadata and RAG chunks.

{
"url": "https://docs.apify.com/academy/api-scraping",
"canonicalUrl": "https://docs.apify.com/academy/api-scraping",
"title": "API scraping | Academy | Apify Documentation",
"description": "Learn all about how the professionals scrape various types of APIs with various configurations, parameters, and requirements.",
"language": "en",
"h1": "API scraping",
"lastModified": "2026-09-25T10:26:13.000Z",
"markdown": "# API scraping\n\nAPI scraping is locating a website's API endpoints, and fetching the desired data directly from their API...\n\n## What's an API?\n\n...",
"wordCount": 868,
"chunkCount": 6,
"chunks": [
{
"index": 0,
"headings": "API scraping",
"text": "# API scraping\n\nAPI scraping is locating a website's API endpoints...",
"charCount": 1790,
"tokenEstimate": 448
}
],
"linksCount": 73,
"depth": 1,
"statusCode": 200,
"fetchMode": "http",
"extractor": "readability",
"contentHash": "3f1c0a9e5b7d2c4e8a61",
"crawledAt": "2026-09-27T11:32:57.120Z",
"warning": null
}

The key-value store also holds:

  • llms-<domain>.txt: an llms.txt index with the title, URL and description of every page.
  • llms-full-<domain>.txt: all Markdown of the domain in one file, each page wrapped in <page url="..." title="...">. Split into parts above 8 MB.
  • RUN-SUMMARY: saved, duplicate, blocked and failed counts, plus the first 500 problem URLs.

Data fields

FieldDescription
url, canonicalUrlFinal URL after redirects and the rel=canonical URL
title, description, h1Title tag, meta description, first heading
languageFrom <html lang>, Content-Language or og:locale
lastModifiedFrom article:modified_time or the Last-Modified header
markdownMain content as GFM Markdown, links and images absolute
wordCount, linksCountWords in the Markdown, unique links on the page
chunks, chunkCountRAG chunks with heading path and token estimate
depth, statusCode, fetchModeCrawl depth, HTTP status, http or browser
contentHash, crawledAt, warningDedupe hash, timestamp, note for thin or empty pages

How much does it cost to convert a website to Markdown?

Pay per event, no compute charges:

EventPrice
Actor start$0.001 per run (per GB of memory)
Page crawled (HTTP)$0.0003, that is $0.30 per 1,000 pages
Page rendered (Chrome, opt-in)$0.006, that is $6 per 1,000 pages

Duplicates, pages blocked by robots.txt, errors and empty pages are not charged. Examples: 10 pages cost $0.004, 1,000 pages $0.30, 10,000 pages $3. With the Apify free plan ($5 credit per month) you can convert about 16,000 pages a month. Set Max cost per run when you start a run and the crawler stops exactly at that budget.

Tips for faster and cheaper crawls

  • Use include globs such as https://example.com/docs/** to skip blogs, tags and login pages.
  • Turn on Also use sitemap.xml to find pages that are not linked from the menu.
  • Keep the default 1024 MB memory for HTTP runs. Use 2048 MB when the JavaScript fallback is on.
  • Leave the JavaScript fallback off unless the Overview view shows warnings like "No text content found".
  • For small or fragile sites, lower Max concurrency per domain to 1 or 2 and add a delay.
  • Use Remove CSS selectors for site-specific clutter, for example .newsletter, #comments.
  • Chunks are split at headings, paragraphs and code blocks. 300 to 800 tokens with 10% overlap works well for most embedding models.

FAQ, disclaimers and support

Does it work on JavaScript sites? Most sites, including Docusaurus, GitBook, WordPress, MkDocs and Next.js sites, ship HTML and work over plain HTTP. Pure single-page apps return an empty shell; enable JavaScript rendering fallback for them.

Does it respect robots.txt? Yes, by default. Disallowed pages are skipped and listed in RUN-SUMMARY. You can turn it off only at your own responsibility; the run then logs a warning.

How are duplicates detected? By final URL, by rel=canonical and by a hash of the Markdown.

Which proxies can I use? No proxy (the default), Apify datacenter proxy, or your own proxy URLs. Residential and SERP proxies are not supported. A run that asks for them stops at the start with a clear message and does no work.

Is crawling legal? Crawling public pages is generally allowed, but you are responsible for complying with each site's terms, copyright and privacy law. Do not crawl personal data without a legal basis.

Known limits. No login or forms. PDF and other files are skipped. Hash-routed single-page apps (/#/page) show as one URL. Token counts are estimates (about 4 characters per token).

Found a bug or need a custom crawler, an export to your vector database or a scheduled pipeline? Open an issue in the Issues tab.