Web Crawler: Website to Markdown for LLM & RAG avatar

Web Crawler: Website to Markdown for LLM & RAG

Pricing

from $1.60 / 1,000 page crawleds

Go to Apify Store
Web Crawler: Website to Markdown for LLM & RAG

Web Crawler: Website to Markdown for LLM & RAG

Website content crawler for AI. Crawl any site into clean Markdown, text and ready-to-embed RAG chunks for LLMs and vector databases. Removes navigation, footers and cookie banners, renders JavaScript only when needed. $2 per 1,000 pages.

Pricing

from $1.60 / 1,000 page crawleds

Rating

5.0

(1)

Developer

Scrape Lads

Scrape Lads

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does this Web Crawler (Website to Markdown for LLM & RAG) do?

Website to Markdown Crawler turns any website into clean Markdown, plain text and ready-to-embed chunks for LLMs, RAG pipelines and vector databases. Give it a docs site, a blog, a help center or a knowledge base, and it crawls the pages, strips navigation, headers, footers, sidebars and cookie banners, and saves the main content of every page as LLM-ready Markdown with its title, description, language and URL.

It fetches pages over fast HTTP and only starts a headless browser when a page is rendered by JavaScript, so most sites crawl in seconds. You pay a flat, predictable price per page with no compute units to estimate. Run it from the Apify Console, call it through the API, schedule it to keep your index fresh, or plug it into LangChain, LlamaIndex, Pinecone, Qdrant, Weaviate, Zapier or Make with Apify integrations.

Why use this web crawler for AI and RAG?

  • Feed your RAG chatbot or AI agent with your own documentation, product pages or support articles.
  • Build a vector database from a website: turn on chunking and every page arrives already split along headings and paragraphs, with the section heading attached to each chunk.
  • Fine-tuning and LLM datasets: collect clean text without menus, ads and legal footers.
  • Keep an AI knowledge base up to date with a weekly schedule.
  • Predictable cost: $2.00 per 1,000 pages. No surprise proxy or compute bills.

How to crawl a website to Markdown for LLMs

  1. Click Try for free.
  2. Paste one or more Start URLs, for example https://docs.apify.com/academy.
  3. Set Max pages (default 50). This is also your price cap.
  4. Optional: set a Chunk size such as 1000 characters for RAG.
  5. Click Start and download your Markdown dataset as JSON, CSV, Excel or HTML, or read it through the API.

Input

Everything is on the Input tab. The most important fields:

FieldWhat it does
startUrlsWebsites or sections to crawl.
maxPagesStop after this many saved pages (default 50).
maxCrawlDepthHow many links away from a start URL to follow (default 5).
crawlScopepath (default) stays under the start URL's folder, hostname allows the whole site, domain also allows subdomains.
includeUrlGlobs / excludeUrlGlobsOptional glob patterns such as https://example.com/docs/** or **/changelog/**.
crawlerTypeauto (HTTP first, browser only when needed), http, or browser.
outputFormatsAny of markdown, text, html (default Markdown and text).
removeBoilerplateRemove navigation, headers, footers, sidebars and cookie banners (default on).
removeElementsCssSelectorExtra CSS selector of elements to delete.
chunkSize / chunkOverlapSplit pages into chunks of at most N characters with overlap (0 = off).
useSitemapsAlso discover pages from robots.txt and sitemap.xml.
respectRobotsTxtSkip pages disallowed by robots.txt (default on).
{
"startUrls": [{ "url": "https://docs.apify.com/academy" }],
"maxPages": 200,
"chunkSize": 1000,
"chunkOverlap": 100,
"excludeUrlGlobs": [{ "glob": "**/changelog/**" }]
}

Output

One dataset item per page. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has three views: Overview, Markdown and RAG chunks (one row per chunk).

{
"url": "https://crawlee.dev/js/docs/introduction/first-crawler",
"loadedUrl": "https://crawlee.dev/js/docs/introduction/first-crawler",
"canonicalUrl": "https://crawlee.dev/js/docs/introduction/first-crawler",
"title": "First crawler | Crawlee for JavaScript",
"description": "Your first steps into the world of scraping with Crawlee",
"language": "en",
"markdown": "# First crawler\n\nNow, you will build your first crawler...\n\n## How Crawlee works\n\n- [`CheerioCrawler`](https://crawlee.dev/js/api/cheerio-crawler/class/CheerioCrawler)\n...",
"text": "First crawler\n\nNow, you will build your first crawler...",
"wordCount": 1091,
"linksCount": 118,
"crawlDepth": 1,
"httpStatus": 200,
"renderedWith": "http",
"loadedAt": "2026-10-01T09:15:02.114Z",
"chunks": [
{
"index": 0,
"text": "# First crawler\n\nNow, you will build your first crawler...",
"heading": "First crawler",
"charCount": 947
},
{
"index": 1,
"text": "...reasonable defaults for everything else.\n### The Where - `Request` and `RequestQueue`\n\nAll crawlers use...",
"heading": "First crawler > How Crawlee works > The Where - `Request` and `RequestQueue`",
"charCount": 996
}
]
}

A RUN_SUMMARY record in the key-value store lists pages crawled, browser pages, failures, duration and the maximum run price.

Data fields

FieldDescription
url / loadedUrlRequested URL and final URL after redirects
canonicalUrl<link rel="canonical">, if present
title, description, languagePage metadata
markdownMain content as GitHub-flavored Markdown (headings, lists, tables, fenced code with language)
textMain content as plain text, one block per line
htmlCleaned main-content HTML (optional)
wordCount, linksCountWords in the content, unique links on the page
crawlDepth, httpStatusLink distance from the start URL, HTTP status
renderedWithhttp or browser
loadedAtISO timestamp
chunks[]index, text, heading (section path), charCount, when chunking is on

How much does it cost to crawl a website for RAG?

Flat pay-per-event pricing, platform usage included:

  • $0.002 per page ($2.00 per 1,000 pages) fetched over HTTP, which is most pages.
  • $0.004 per page ($4.00 per 1,000) for pages that needed a headless browser.
  • $0.005 per run.

A 500-page documentation site costs about $1.00. Only saved pages are charged: 404s, empty pages, duplicates and failed requests are free. maxPages and the run's maximum cost limit are both hard caps, and the crawler stops as soon as either is reached. Apify's free plan includes monthly credits, enough for thousands of pages.

Tips for better LLM-ready Markdown

  • Scope first. Start from the section you need (/docs, /blog) and keep the default path scope. Use exclude globs for changelogs, tag pages and paginated archives.
  • Chunk size: 800 to 2000 characters suits most embedding models; 10 to 15% overlap keeps context across boundaries.
  • Use http crawler type for static sites (docs, blogs) to guarantee the cheapest price; use browser only for single-page apps.
  • Sitemaps find pages that are not linked from navigation. Combine with maxPages to keep runs bounded.
  • Use removeElementsCssSelector to drop site-specific noise such as .newsletter, #comments.

Use with LangChain and LlamaIndex

Load the dataset with LangChain's ApifyDatasetLoader (map markdown or chunks[].text to page_content and loadedUrl to metadata), or with LlamaIndex's Apify reader. For vector databases, the RAG chunks view gives one row per chunk with its URL, title and section heading.

FAQ, disclaimers and support

Does it handle JavaScript-heavy sites? Yes. In auto mode a page that arrives as an empty app shell is rendered in headless Chrome automatically. Apps that route with #/ fragments expose only their start page.

Does it respect robots.txt? Yes by default. It also only fetches public websites; private and local network addresses are refused.

What about sites that block bots? Blocked pages (403, 429, anti-bot challenges) are retried automatically through a proxy. Sites behind logins or strong bot protection may still fail; failed pages are not charged.

Is it legal? Crawling publicly available pages is generally allowed, but you are responsible for complying with each site's terms of service, copyright and data-protection laws such as GDPR. Do not crawl personal data without a lawful basis.

Found a problem or need a field? Open an issue on the Issues tab. Need a custom crawler or a full RAG ingestion pipeline? Get in touch through the Issues tab.