# Website to Markdown for RAG & LLMs: Fast Crawler (`rod_analytics/website-to-markdown`) Actor

Crawl any website and turn every page into clean Markdown for LLMs, RAG and AI agents. Fast HTTP crawler, main-content extraction, GFM tables, RAG chunks, llms.txt per domain, robots.txt respected. Flat price per page.

- **URL**: https://apify.com/rod\_analytics/website-to-markdown.md
- **Developed by:** [Rod Services](https://apify.com/rod_analytics) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Website to Markdown Crawler do?

**Website to Markdown Crawler** crawls any website and turns every page into **clean Markdown for LLMs, RAG pipelines and AI agents**. It follows links on the site, keeps only the **main content** (no menus, sidebars, footers or cookie banners), converts it to GitHub-flavoured Markdown with **tables, code blocks and absolute links**, and can split each page into **RAG chunks** with token estimates. At the end it writes an **llms.txt** and an **llms-full.txt** file per domain.

It is a **fast HTTP crawler, no browser**, so a 200-page documentation site takes about a minute. JavaScript-only pages can optionally be rendered in headless Chrome. You pay a **flat price per page**, not for compute.

Try it: the prefilled input crawls 10 pages of [Apify Academy](https://docs.apify.com/academy) in under 30 seconds. Run it from the Console, call it from the API, schedule it, or connect it to LangChain, LlamaIndex, Pinecone, Qdrant, Make, n8n or Zapier through Apify integrations.

### Why use it for LLM and RAG data?

- **Knowledge base for a chatbot or AI agent.** Crawl your docs, help center or blog and load the chunks into a vector database.
- **LLM training and fine-tuning data.** Clean Markdown without boilerplate, one record per page, with language and word count.
- **llms.txt for your own site.** Get a ready `llms.txt` index and an `llms-full.txt` file that AI assistants can read.
- **Keep an index fresh.** Schedule weekly runs; `contentHash` and `lastModified` show which pages changed.
- **Predictable cost.** $0.30 per 1,000 pages over HTTP. No compute units to estimate.

### How to crawl a website to Markdown

1. Open the **Input** tab and paste one or more start URLs, for example `https://docs.example.com/`.
2. Set **Max pages**. This is also your cost cap.
3. Optional: turn on **Split into RAG chunks** and pick a chunk size.
4. Click **Start**. Pages appear in the **Output** tab while the crawl runs.
5. Download the dataset as JSON, CSV, Excel or HTML, or open **llms.txt files** in the Storage tab.

By default the crawler stays under the start URL path: `https://example.com/docs` crawls `/docs` and everything below it.

### Input

All options are on the Input tab. The most useful ones:

| Field                                                                 | What it does                                 | Default      |
| --------------------------------------------------------------------- | -------------------------------------------- | ------------ |
| `startUrls`                                                           | Where the crawl starts                       | required     |
| `maxPages`                                                            | Stop after this many saved pages             | 100          |
| `maxCrawlDepth`                                                       | Links away from a start URL                  | 20           |
| `includeGlobs` / `excludeGlobs`                                       | URL patterns, `**` matches any path          | none         |
| `sameDomainOnly` / `stayWithinStartPath`                              | Crawl scope                                  | on / on      |
| `useSitemaps`                                                         | Also queue URLs from sitemap.xml             | off          |
| `respectRobotsTxt`                                                    | Skip pages disallowed by robots.txt          | on           |
| `extractionMode`                                                      | `readability`, `main` or `full` page         | readability  |
| `removeSelectors`                                                     | Extra CSS selectors to drop                  | none         |
| `includeChunks`, `chunkSize`, `chunkOverlap`                          | RAG chunks, sizes in estimated tokens        | off, 500, 50 |
| `dedupeContent`                                                       | Skip same canonical URL or identical content | on           |
| `jsRenderingFallback`                                                 | Render near-empty pages in headless Chrome   | off          |
| `maxConcurrency`, `maxConcurrencyPerDomain`, `delayBetweenRequestsMs` | Politeness per domain                        | 10, 5, 0     |

```json
{
    "startUrls": [{ "url": "https://docs.apify.com/academy" }],
    "maxPages": 10,
    "includeChunks": true,
    "chunkSize": 500,
    "chunkOverlap": 50
}
```

### Output

One dataset item per page. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has four views: Overview, Markdown, Metadata and RAG chunks.

```json
{
    "url": "https://docs.apify.com/academy/api-scraping",
    "canonicalUrl": "https://docs.apify.com/academy/api-scraping",
    "title": "API scraping | Academy | Apify Documentation",
    "description": "Learn all about how the professionals scrape various types of APIs with various configurations, parameters, and requirements.",
    "language": "en",
    "h1": "API scraping",
    "lastModified": "2026-09-25T10:26:13.000Z",
    "markdown": "# API scraping\n\nAPI scraping is locating a website's API endpoints, and fetching the desired data directly from their API...\n\n## What's an API?\n\n...",
    "wordCount": 868,
    "chunkCount": 6,
    "chunks": [
        {
            "index": 0,
            "headings": "API scraping",
            "text": "# API scraping\n\nAPI scraping is locating a website's API endpoints...",
            "charCount": 1790,
            "tokenEstimate": 448
        }
    ],
    "linksCount": 73,
    "depth": 1,
    "statusCode": 200,
    "fetchMode": "http",
    "extractor": "readability",
    "contentHash": "3f1c0a9e5b7d2c4e8a61",
    "crawledAt": "2026-09-27T11:32:57.120Z",
    "warning": null
}
```

The key-value store also holds:

- `llms-<domain>.txt`: an [llms.txt](https://llmstxt.org) index with the title, URL and description of every page.
- `llms-full-<domain>.txt`: all Markdown of the domain in one file, each page wrapped in `<page url="..." title="...">`. Split into parts above 8 MB.
- `RUN-SUMMARY`: saved, duplicate, blocked and failed counts, plus the first 500 problem URLs.

### Data fields

| Field                                 | Description                                             |
| ------------------------------------- | ------------------------------------------------------- |
| `url`, `canonicalUrl`                 | Final URL after redirects and the rel=canonical URL     |
| `title`, `description`, `h1`          | Title tag, meta description, first heading              |
| `language`                            | From `<html lang>`, Content-Language or og:locale       |
| `lastModified`                        | From article:modified\_time or the Last-Modified header  |
| `markdown`                            | Main content as GFM Markdown, links and images absolute |
| `wordCount`, `linksCount`             | Words in the Markdown, unique links on the page         |
| `chunks`, `chunkCount`                | RAG chunks with heading path and token estimate         |
| `depth`, `statusCode`, `fetchMode`    | Crawl depth, HTTP status, `http` or `browser`           |
| `contentHash`, `crawledAt`, `warning` | Dedupe hash, timestamp, note for thin or empty pages    |

### How much does it cost to convert a website to Markdown?

Pay per event, no compute charges:

| Event                          | Price                                      |
| ------------------------------ | ------------------------------------------ |
| Actor start                    | $0.001 per run (per GB of memory)          |
| Page crawled (HTTP)            | $0.0003, that is **$0.30 per 1,000 pages** |
| Page rendered (Chrome, opt-in) | $0.006, that is $6 per 1,000 pages         |

Duplicates, pages blocked by robots.txt, errors and empty pages are **not charged**. Examples: 10 pages cost $0.004, 1,000 pages $0.30, 10,000 pages $3. With the Apify free plan ($5 credit per month) you can convert about 16,000 pages a month. Set **Max cost per run** when you start a run and the crawler stops exactly at that budget.

### Tips for faster and cheaper crawls

- Use **include globs** such as `https://example.com/docs/**` to skip blogs, tags and login pages.
- Turn on **Also use sitemap.xml** to find pages that are not linked from the menu.
- Keep the default 1024 MB memory for HTTP runs. Use 2048 MB when the JavaScript fallback is on.
- Leave the JavaScript fallback off unless the Overview view shows warnings like "No text content found".
- For small or fragile sites, lower **Max concurrency per domain** to 1 or 2 and add a delay.
- Use **Remove CSS selectors** for site-specific clutter, for example `.newsletter, #comments`.
- Chunks are split at headings, paragraphs and code blocks. 300 to 800 tokens with 10% overlap works well for most embedding models.

### FAQ, disclaimers and support

**Does it work on JavaScript sites?** Most sites, including Docusaurus, GitBook, WordPress, MkDocs and Next.js sites, ship HTML and work over plain HTTP. Pure single-page apps return an empty shell; enable **JavaScript rendering fallback** for them.

**Does it respect robots.txt?** Yes, by default. Disallowed pages are skipped and listed in `RUN-SUMMARY`. You can turn it off only at your own responsibility; the run then logs a warning.

**How are duplicates detected?** By final URL, by `rel=canonical` and by a hash of the Markdown.

**Which proxies can I use?** No proxy (the default), Apify datacenter proxy, or your own proxy URLs. Residential and SERP proxies are not supported. A run that asks for them stops at the start with a clear message and does no work.

**Is crawling legal?** Crawling public pages is generally allowed, but you are responsible for complying with each site's terms, copyright and privacy law. Do not crawl personal data without a legal basis.

**Known limits.** No login or forms. PDF and other files are skipped. Hash-routed single-page apps (`/#/page`) show as one URL. Token counts are estimates (about 4 characters per token).

Found a bug or need a custom crawler, an export to your vector database or a scheduled pipeline? Open an issue in the **Issues** tab.

# Actor input Schema

## `startUrls` (type: `array`):

Pages where the crawl starts. By default only pages under the same path are crawled: https://example.com/docs crawls /docs and everything below it.

## `maxPages` (type: `integer`):

Stop after this many pages are saved. You pay per saved page, so this is also your cost cap.

## `maxCrawlDepth` (type: `integer`):

How many links away from a start URL the crawler may go. 0 crawls only the start URLs (plus sitemap URLs, if enabled).

## `includeGlobs` (type: `array`):

Only crawl URLs matching at least one pattern. Use \*\* for any path, for example https://example.com/blog/\*\*. When set, this replaces the "Stay within start path" rule.

## `excludeGlobs` (type: `array`):

Never crawl URLs matching any of these patterns, for example **/tag/** or https://example.com/login\*\*.

## `sameDomainOnly` (type: `boolean`):

Follow links only on the start URL hosts. www.example.com and example.com count as the same host.

## `stayWithinStartPath` (type: `boolean`):

Only follow links under the start URL path. Turn off to crawl the whole site.

## `useSitemaps` (type: `boolean`):

Find sitemaps via robots.txt and /sitemap.xml and queue the URLs they list (still filtered by the rules above). Finds pages that are not linked from anywhere.

## `respectRobotsTxt` (type: `boolean`):

WARNING: keep this on. Pages disallowed by the site robots.txt are skipped and never charged. Turning it off logs a warning, and you alone are responsible for having permission to crawl the site.

## `extractionMode` (type: `string`):

Readability: finds the main article like Firefox Reader View and drops menus, sidebars, footers and cookie banners (best for RAG). Main element: only removes known boilerplate and keeps the main or article element. Full page: keeps everything except scripts and cookie banners.

## `removeSelectors` (type: `string`):

Extra elements to drop before conversion, as one CSS selector list. Example: .newsletter, #comments, .related-posts

## `dedupeContent` (type: `boolean`):

Skip pages whose canonical URL was already saved or whose Markdown is identical to a saved page. Duplicates are not charged.

## `saveLlmsTxt` (type: `boolean`):

Write llms-<domain>.txt (index of pages) and llms-full-<domain>.txt (all Markdown in one file) to the key-value store at the end of the run.

## `includeChunks` (type: `boolean`):

Add a chunks array to each page, split at headings, paragraphs and code blocks, ready for embeddings and a vector database.

## `chunkSize` (type: `integer`):

Target size of one chunk in estimated tokens (about 4 characters per token).

## `chunkOverlap` (type: `integer`):

How many tokens of the previous chunk are repeated at the start of the next one. Capped at half the chunk size.

## `jsRenderingFallback` (type: `boolean`):

Off by default. When a page has almost no text over HTTP (a JavaScript app shell), open it once in headless Chrome and extract the rendered content. Rendered pages are billed as a separate, more expensive event. Use at least 2048 MB memory.

## `minWordsForHttp` (type: `integer`):

With the fallback on, pages with fewer words than this over HTTP are rendered in the browser.

## `maxConcurrency` (type: `integer`):

Maximum parallel requests in total.

## `maxConcurrencyPerDomain` (type: `integer`):

Maximum parallel requests to one host. Lower it for small sites.

## `delayBetweenRequestsMs` (type: `integer`):

Minimum time between two request starts to the same host. 0 means no delay.

## `requestTimeoutSecs` (type: `integer`):

Give up on a page after this many seconds.

## `maxRequestRetries` (type: `integer`):

Retries for network errors and blocked responses. Failed pages are not charged.

## `proxyConfiguration` (type: `object`):

Optional. Most sites work without a proxy. Use Apify datacenter proxy or your own proxy URLs if a site blocks you. Residential and SERP proxies are not supported.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy"
    }
  ],
  "maxPages": 10,
  "maxCrawlDepth": 20,
  "includeGlobs": [],
  "excludeGlobs": [],
  "sameDomainOnly": true,
  "stayWithinStartPath": true,
  "useSitemaps": false,
  "respectRobotsTxt": true,
  "extractionMode": "readability",
  "dedupeContent": true,
  "saveLlmsTxt": true,
  "includeChunks": false,
  "chunkSize": 500,
  "chunkOverlap": 50,
  "jsRenderingFallback": false,
  "minWordsForHttp": 30,
  "maxConcurrency": 10,
  "maxConcurrencyPerDomain": 5,
  "delayBetweenRequestsMs": 0,
  "requestTimeoutSecs": 30,
  "maxRequestRetries": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `pages` (type: `string`):

No description

## `markdown` (type: `string`):

No description

## `files` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://docs.apify.com/academy"
        }
    ],
    "maxPages": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("rod_analytics/website-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://docs.apify.com/academy" }],
    "maxPages": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("rod_analytics/website-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy"
    }
  ],
  "maxPages": 10
}' |
apify call rod_analytics/website-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rod_analytics/website-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZbhYRQVxEGyfcxxp7/builds/cbls0mPLpmHJ5EPZ9/openapi.json
