# Web Crawler: Website to Markdown for LLM & RAG (`scrape-lads/website-content-crawler-rag`) Actor

Website content crawler for AI. Crawl any site into clean Markdown, text and ready-to-embed RAG chunks for LLMs and vector databases. Removes navigation, footers and cookie banners, renders JavaScript only when needed. $2 per 1,000 pages.

- **URL**: https://apify.com/scrape-lads/website-content-crawler-rag.md
- **Developed by:** [Scrape Lads](https://apify.com/scrape-lads) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $1.60 / 1,000 page crawleds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does this Web Crawler (Website to Markdown for LLM & RAG) do?

**Website to Markdown Crawler** turns any website into **clean Markdown, plain text and ready-to-embed chunks** for **LLMs, RAG pipelines and vector databases**. Give it a docs site, a blog, a help center or a knowledge base, and it crawls the pages, **strips navigation, headers, footers, sidebars and cookie banners**, and saves the main content of every page as LLM-ready Markdown with its title, description, language and URL.

It fetches pages over fast HTTP and **only starts a headless browser when a page is rendered by JavaScript**, so most sites crawl in seconds. You pay a **flat, predictable price per page** with no compute units to estimate. Run it from the Apify Console, call it through the API, schedule it to keep your index fresh, or plug it into LangChain, LlamaIndex, Pinecone, Qdrant, Weaviate, Zapier or Make with Apify integrations.

### Why use this web crawler for AI and RAG?

- **Feed your RAG chatbot or AI agent** with your own documentation, product pages or support articles.
- **Build a vector database** from a website: turn on chunking and every page arrives already split along headings and paragraphs, with the section heading attached to each chunk.
- **Fine-tuning and LLM datasets**: collect clean text without menus, ads and legal footers.
- **Keep an AI knowledge base up to date** with a weekly schedule.
- **Predictable cost**: $2.00 per 1,000 pages. No surprise proxy or compute bills.

### How to crawl a website to Markdown for LLMs

1. Click **Try for free**.
2. Paste one or more **Start URLs**, for example `https://docs.apify.com/academy`.
3. Set **Max pages** (default 50). This is also your price cap.
4. Optional: set a **Chunk size** such as 1000 characters for RAG.
5. Click **Start** and download your Markdown dataset as JSON, CSV, Excel or HTML, or read it through the API.

### Input

Everything is on the Input tab. The most important fields:

| Field                                 | What it does                                                                                                            |
| ------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `startUrls`                           | Websites or sections to crawl.                                                                                          |
| `maxPages`                            | Stop after this many saved pages (default 50).                                                                          |
| `maxCrawlDepth`                       | How many links away from a start URL to follow (default 5).                                                             |
| `crawlScope`                          | `path` (default) stays under the start URL's folder, `hostname` allows the whole site, `domain` also allows subdomains. |
| `includeUrlGlobs` / `excludeUrlGlobs` | Optional glob patterns such as `https://example.com/docs/**` or `**/changelog/**`.                                      |
| `crawlerType`                         | `auto` (HTTP first, browser only when needed), `http`, or `browser`.                                                    |
| `outputFormats`                       | Any of `markdown`, `text`, `html` (default Markdown and text).                                                          |
| `removeBoilerplate`                   | Remove navigation, headers, footers, sidebars and cookie banners (default on).                                          |
| `removeElementsCssSelector`           | Extra CSS selector of elements to delete.                                                                               |
| `chunkSize` / `chunkOverlap`          | Split pages into chunks of at most N characters with overlap (0 = off).                                                 |
| `useSitemaps`                         | Also discover pages from robots.txt and sitemap.xml.                                                                    |
| `respectRobotsTxt`                    | Skip pages disallowed by robots.txt (default on).                                                                       |

```json
{
    "startUrls": [{ "url": "https://docs.apify.com/academy" }],
    "maxPages": 200,
    "chunkSize": 1000,
    "chunkOverlap": 100,
    "excludeUrlGlobs": [{ "glob": "**/changelog/**" }]
}
```

### Output

One dataset item per page. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has three views: **Overview**, **Markdown** and **RAG chunks** (one row per chunk).

```json
{
    "url": "https://crawlee.dev/js/docs/introduction/first-crawler",
    "loadedUrl": "https://crawlee.dev/js/docs/introduction/first-crawler",
    "canonicalUrl": "https://crawlee.dev/js/docs/introduction/first-crawler",
    "title": "First crawler | Crawlee for JavaScript",
    "description": "Your first steps into the world of scraping with Crawlee",
    "language": "en",
    "markdown": "# First crawler\n\nNow, you will build your first crawler...\n\n## How Crawlee works\n\n- [`CheerioCrawler`](https://crawlee.dev/js/api/cheerio-crawler/class/CheerioCrawler)\n...",
    "text": "First crawler\n\nNow, you will build your first crawler...",
    "wordCount": 1091,
    "linksCount": 118,
    "crawlDepth": 1,
    "httpStatus": 200,
    "renderedWith": "http",
    "loadedAt": "2026-10-01T09:15:02.114Z",
    "chunks": [
        {
            "index": 0,
            "text": "# First crawler\n\nNow, you will build your first crawler...",
            "heading": "First crawler",
            "charCount": 947
        },
        {
            "index": 1,
            "text": "...reasonable defaults for everything else.\n### The Where - `Request` and `RequestQueue`\n\nAll crawlers use...",
            "heading": "First crawler > How Crawlee works > The Where - `Request` and `RequestQueue`",
            "charCount": 996
        }
    ]
}
```

A `RUN_SUMMARY` record in the key-value store lists pages crawled, browser pages, failures, duration and the maximum run price.

### Data fields

| Field                              | Description                                                                                   |
| ---------------------------------- | --------------------------------------------------------------------------------------------- |
| `url` / `loadedUrl`                | Requested URL and final URL after redirects                                                   |
| `canonicalUrl`                     | `<link rel="canonical">`, if present                                                          |
| `title`, `description`, `language` | Page metadata                                                                                 |
| `markdown`                         | Main content as GitHub-flavored Markdown (headings, lists, tables, fenced code with language) |
| `text`                             | Main content as plain text, one block per line                                                |
| `html`                             | Cleaned main-content HTML (optional)                                                          |
| `wordCount`, `linksCount`          | Words in the content, unique links on the page                                                |
| `crawlDepth`, `httpStatus`         | Link distance from the start URL, HTTP status                                                 |
| `renderedWith`                     | `http` or `browser`                                                                           |
| `loadedAt`                         | ISO timestamp                                                                                 |
| `chunks[]`                         | `index`, `text`, `heading` (section path), `charCount`, when chunking is on                   |

### How much does it cost to crawl a website for RAG?

Flat pay-per-event pricing, platform usage included:

- **$0.002 per page** ($2.00 per 1,000 pages) fetched over HTTP, which is most pages.
- **$0.004 per page** ($4.00 per 1,000) for pages that needed a headless browser.
- **$0.005 per run**.

A 500-page documentation site costs about **$1.00**. Only saved pages are charged: 404s, empty pages, duplicates and failed requests are free. `maxPages` and the run's maximum cost limit are both hard caps, and the crawler stops as soon as either is reached. Apify's free plan includes monthly credits, enough for thousands of pages.

### Tips for better LLM-ready Markdown

- **Scope first.** Start from the section you need (`/docs`, `/blog`) and keep the default `path` scope. Use exclude globs for changelogs, tag pages and paginated archives.
- **Chunk size:** 800 to 2000 characters suits most embedding models; 10 to 15% overlap keeps context across boundaries.
- **Use `http` crawler type** for static sites (docs, blogs) to guarantee the cheapest price; use `browser` only for single-page apps.
- **Sitemaps** find pages that are not linked from navigation. Combine with `maxPages` to keep runs bounded.
- Use `removeElementsCssSelector` to drop site-specific noise such as `.newsletter, #comments`.

#### Use with LangChain and LlamaIndex

Load the dataset with LangChain's `ApifyDatasetLoader` (map `markdown` or `chunks[].text` to `page_content` and `loadedUrl` to metadata), or with LlamaIndex's Apify reader. For vector databases, the **RAG chunks** view gives one row per chunk with its URL, title and section heading.

### FAQ, disclaimers and support

**Does it handle JavaScript-heavy sites?** Yes. In `auto` mode a page that arrives as an empty app shell is rendered in headless Chrome automatically. Apps that route with `#/` fragments expose only their start page.

**Does it respect robots.txt?** Yes by default. It also only fetches public websites; private and local network addresses are refused.

**What about sites that block bots?** Blocked pages (403, 429, anti-bot challenges) are retried automatically through a proxy. Sites behind logins or strong bot protection may still fail; failed pages are not charged.

**Is it legal?** Crawling publicly available pages is generally allowed, but you are responsible for complying with each site's terms of service, copyright and data-protection laws such as GDPR. Do not crawl personal data without a lawful basis.

Found a problem or need a field? Open an issue on the **Issues** tab. Need a custom crawler or a full RAG ingestion pipeline? Get in touch through the Issues tab.

# Actor input Schema

## `startUrls` (type: `array`):

Websites or sections to crawl. With the default scope the crawler stays under each URL's path, so https://docs.apify.com/academy only crawls pages under /academy.

## `maxPages` (type: `integer`):

Stop after this many pages are saved. Each saved page is one billable event, so this is also your price cap.

## `maxCrawlDepth` (type: `integer`):

How many links away from a start URL to follow. 0 crawls only the start URLs.

## `crawlScope` (type: `string`):

Which discovered links to follow. "Same path" stays under the start URL's folder, "Same hostname" allows the whole host, "Same domain" also allows subdomains.

## `includeUrlGlobs` (type: `array`):

Optional. Only follow links matching one of these glob patterns, e.g. https://example.com/docs/\*\*. Applied on top of the crawl scope.

## `excludeUrlGlobs` (type: `array`):

Optional. Never follow links matching these glob patterns, e.g. **/changelog/** or \*\*/*?page=*.

## `crawlerType` (type: `string`):

"Auto" fetches pages over fast HTTP and only renders a page in a headless browser when it looks JavaScript-rendered (empty app shell). "HTTP only" never uses a browser. "Browser" renders every page in headless Chrome (slower, $0.004 per page).

## `outputFormats` (type: `array`):

Which content fields to save for each page.

## `removeBoilerplate` (type: `boolean`):

Strip navigation, headers, footers, sidebars, cookie banners and share widgets, and keep the main content only.

## `removeElementsCssSelector` (type: `string`):

Optional CSS selector of additional elements to delete before conversion, e.g. .ad-slot, #comments.

## `chunkSize` (type: `integer`):

Set above 0 to split each page into chunks of at most this many characters, along headings and paragraphs, ready for embeddings. 1000 to 2000 works well for most vector databases. 0 turns chunking off.

## `chunkOverlap` (type: `integer`):

Characters from the end of the previous chunk repeated at the start of the next one, so context is not lost at the boundary. Capped at half the chunk size.

## `useSitemaps` (type: `boolean`):

Also read robots.txt Sitemap entries and /sitemap.xml to find pages that are not linked. In-scope sitemap URLs are crawled alongside discovered links.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that the site's robots.txt disallows.

## `maxConcurrency` (type: `integer`):

Maximum parallel HTTP requests. Browser pages are always limited to a few at a time.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy/web-scraping-for-beginners"
    }
  ],
  "maxPages": 50,
  "maxCrawlDepth": 5,
  "crawlScope": "path",
  "includeUrlGlobs": [],
  "excludeUrlGlobs": [],
  "crawlerType": "auto",
  "outputFormats": [
    "markdown",
    "text"
  ],
  "removeBoilerplate": true,
  "chunkSize": 0,
  "chunkOverlap": 200,
  "useSitemaps": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 10
}
```

# Actor output Schema

## `pages` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://docs.apify.com/academy/web-scraping-for-beginners"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrape-lads/website-content-crawler-rag").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://docs.apify.com/academy/web-scraping-for-beginners" }] }

# Run the Actor and wait for it to finish
run = client.actor("scrape-lads/website-content-crawler-rag").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy/web-scraping-for-beginners"
    }
  ]
}' |
apify call scrape-lads/website-content-crawler-rag --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrape-lads/website-content-crawler-rag"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/nPOuWiIrWgbxn1y3p/builds/JgjJqsm1Q14lxlZSD/openapi.json
