Web Crawler: Website to Markdown for LLM & RAG
Pricing
from $1.60 / 1,000 page crawleds
Web Crawler: Website to Markdown for LLM & RAG
Website content crawler for AI. Crawl any site into clean Markdown, text and ready-to-embed RAG chunks for LLMs and vector databases. Removes navigation, footers and cookie banners, renders JavaScript only when needed. $2 per 1,000 pages.
Pricing
from $1.60 / 1,000 page crawleds
Rating
5.0
(1)
Developer
Scrape Lads
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What does this Web Crawler (Website to Markdown for LLM & RAG) do?
Website to Markdown Crawler turns any website into clean Markdown, plain text and ready-to-embed chunks for LLMs, RAG pipelines and vector databases. Give it a docs site, a blog, a help center or a knowledge base, and it crawls the pages, strips navigation, headers, footers, sidebars and cookie banners, and saves the main content of every page as LLM-ready Markdown with its title, description, language and URL.
It fetches pages over fast HTTP and only starts a headless browser when a page is rendered by JavaScript, so most sites crawl in seconds. You pay a flat, predictable price per page with no compute units to estimate. Run it from the Apify Console, call it through the API, schedule it to keep your index fresh, or plug it into LangChain, LlamaIndex, Pinecone, Qdrant, Weaviate, Zapier or Make with Apify integrations.
Why use this web crawler for AI and RAG?
- Feed your RAG chatbot or AI agent with your own documentation, product pages or support articles.
- Build a vector database from a website: turn on chunking and every page arrives already split along headings and paragraphs, with the section heading attached to each chunk.
- Fine-tuning and LLM datasets: collect clean text without menus, ads and legal footers.
- Keep an AI knowledge base up to date with a weekly schedule.
- Predictable cost: $2.00 per 1,000 pages. No surprise proxy or compute bills.
How to crawl a website to Markdown for LLMs
- Click Try for free.
- Paste one or more Start URLs, for example
https://docs.apify.com/academy. - Set Max pages (default 50). This is also your price cap.
- Optional: set a Chunk size such as 1000 characters for RAG.
- Click Start and download your Markdown dataset as JSON, CSV, Excel or HTML, or read it through the API.
Input
Everything is on the Input tab. The most important fields:
| Field | What it does |
|---|---|
startUrls | Websites or sections to crawl. |
maxPages | Stop after this many saved pages (default 50). |
maxCrawlDepth | How many links away from a start URL to follow (default 5). |
crawlScope | path (default) stays under the start URL's folder, hostname allows the whole site, domain also allows subdomains. |
includeUrlGlobs / excludeUrlGlobs | Optional glob patterns such as https://example.com/docs/** or **/changelog/**. |
crawlerType | auto (HTTP first, browser only when needed), http, or browser. |
outputFormats | Any of markdown, text, html (default Markdown and text). |
removeBoilerplate | Remove navigation, headers, footers, sidebars and cookie banners (default on). |
removeElementsCssSelector | Extra CSS selector of elements to delete. |
chunkSize / chunkOverlap | Split pages into chunks of at most N characters with overlap (0 = off). |
useSitemaps | Also discover pages from robots.txt and sitemap.xml. |
respectRobotsTxt | Skip pages disallowed by robots.txt (default on). |
{"startUrls": [{ "url": "https://docs.apify.com/academy" }],"maxPages": 200,"chunkSize": 1000,"chunkOverlap": 100,"excludeUrlGlobs": [{ "glob": "**/changelog/**" }]}
Output
One dataset item per page. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has three views: Overview, Markdown and RAG chunks (one row per chunk).
{"url": "https://crawlee.dev/js/docs/introduction/first-crawler","loadedUrl": "https://crawlee.dev/js/docs/introduction/first-crawler","canonicalUrl": "https://crawlee.dev/js/docs/introduction/first-crawler","title": "First crawler | Crawlee for JavaScript","description": "Your first steps into the world of scraping with Crawlee","language": "en","markdown": "# First crawler\n\nNow, you will build your first crawler...\n\n## How Crawlee works\n\n- [`CheerioCrawler`](https://crawlee.dev/js/api/cheerio-crawler/class/CheerioCrawler)\n...","text": "First crawler\n\nNow, you will build your first crawler...","wordCount": 1091,"linksCount": 118,"crawlDepth": 1,"httpStatus": 200,"renderedWith": "http","loadedAt": "2026-10-01T09:15:02.114Z","chunks": [{"index": 0,"text": "# First crawler\n\nNow, you will build your first crawler...","heading": "First crawler","charCount": 947},{"index": 1,"text": "...reasonable defaults for everything else.\n### The Where - `Request` and `RequestQueue`\n\nAll crawlers use...","heading": "First crawler > How Crawlee works > The Where - `Request` and `RequestQueue`","charCount": 996}]}
A RUN_SUMMARY record in the key-value store lists pages crawled, browser pages, failures, duration and the maximum run price.
Data fields
| Field | Description |
|---|---|
url / loadedUrl | Requested URL and final URL after redirects |
canonicalUrl | <link rel="canonical">, if present |
title, description, language | Page metadata |
markdown | Main content as GitHub-flavored Markdown (headings, lists, tables, fenced code with language) |
text | Main content as plain text, one block per line |
html | Cleaned main-content HTML (optional) |
wordCount, linksCount | Words in the content, unique links on the page |
crawlDepth, httpStatus | Link distance from the start URL, HTTP status |
renderedWith | http or browser |
loadedAt | ISO timestamp |
chunks[] | index, text, heading (section path), charCount, when chunking is on |
How much does it cost to crawl a website for RAG?
Flat pay-per-event pricing, platform usage included:
- $0.002 per page ($2.00 per 1,000 pages) fetched over HTTP, which is most pages.
- $0.004 per page ($4.00 per 1,000) for pages that needed a headless browser.
- $0.005 per run.
A 500-page documentation site costs about $1.00. Only saved pages are charged: 404s, empty pages, duplicates and failed requests are free. maxPages and the run's maximum cost limit are both hard caps, and the crawler stops as soon as either is reached. Apify's free plan includes monthly credits, enough for thousands of pages.
Tips for better LLM-ready Markdown
- Scope first. Start from the section you need (
/docs,/blog) and keep the defaultpathscope. Use exclude globs for changelogs, tag pages and paginated archives. - Chunk size: 800 to 2000 characters suits most embedding models; 10 to 15% overlap keeps context across boundaries.
- Use
httpcrawler type for static sites (docs, blogs) to guarantee the cheapest price; usebrowseronly for single-page apps. - Sitemaps find pages that are not linked from navigation. Combine with
maxPagesto keep runs bounded. - Use
removeElementsCssSelectorto drop site-specific noise such as.newsletter, #comments.
Use with LangChain and LlamaIndex
Load the dataset with LangChain's ApifyDatasetLoader (map markdown or chunks[].text to page_content and loadedUrl to metadata), or with LlamaIndex's Apify reader. For vector databases, the RAG chunks view gives one row per chunk with its URL, title and section heading.
FAQ, disclaimers and support
Does it handle JavaScript-heavy sites? Yes. In auto mode a page that arrives as an empty app shell is rendered in headless Chrome automatically. Apps that route with #/ fragments expose only their start page.
Does it respect robots.txt? Yes by default. It also only fetches public websites; private and local network addresses are refused.
What about sites that block bots? Blocked pages (403, 429, anti-bot challenges) are retried automatically through a proxy. Sites behind logins or strong bot protection may still fail; failed pages are not charged.
Is it legal? Crawling publicly available pages is generally allowed, but you are responsible for complying with each site's terms of service, copyright and data-protection laws such as GDPR. Do not crawl personal data without a lawful basis.
Found a problem or need a field? Open an issue on the Issues tab. Need a custom crawler or a full RAG ingestion pipeline? Get in touch through the Issues tab.