# Website to Markdown Crawler for AI & RAG (`tidytools/website-markdown-crawler`) Actor

Crawl a website or docs site and get clean, LLM-ready Markdown for every page, with optional RAG chunks. Handles PDFs and Word files too. Fast mode from $1 per 1,000 pages.

- **URL**: https://apify.com/tidytools/website-markdown-crawler.md
- **Developed by:** [Yukai Lin](https://apify.com/tidytools) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Website to Markdown Crawler do?

It crawls a website, documentation portal, blog or knowledge base and returns **clean, LLM-ready Markdown for every page**, ready for ChatGPT / Claude context, vector databases and RAG pipelines.

- 🕷️ **Crawls for you**: start from one URL and follow links within the same folder, the whole site, or only the pages you list
- 📝 **Clean Markdown**: headings, lists, tables, links and code blocks preserved; menus can be dropped with a CSS selector
- 📄 **Documents too**: linked **PDF, Word (DOCX), Excel (XLSX) and CSV** files are converted to Markdown as well
- ✂️ **RAG chunks built in**: optional `chunks` array split at paragraph boundaries with overlap, ready for embeddings
- ⚡ **Fast and cheap**: pages are fetched over plain HTTP whenever possible (**$1 per 1,000 pages**) and rendered in a real browser only when a site needs JavaScript or blocks simple requests
- 🧹 **No duplicates, no surprises**: identical pages reachable under several URLs are kept once and **charged once**; blocked and failed pages are **free**

### Who is it for?

- Building a **RAG chatbot** or **AI assistant** over your docs, help center or website
- Feeding **LLM agents** with up-to-date documentation
- Creating a **knowledge base export** or content audit
- **Migrating** a website's content to another CMS

### How much does it cost?

Pay per event, only for pages that were converted successfully:

| Event | Price |
|---|---|
| Page (fast mode, plain HTTP) | $1.00 / 1,000 pages |
| Page (browser mode, JavaScript rendering) | $2.50 / 1,000 pages |

In **Auto** mode (default) most pages use the fast mode. Duplicates, blocked pages (403, bot checks) and errors are not charged. Your maximum charge limit is always respected.

### How to use it

1. Enter one or more **Start URLs**, e.g. `https://docs.example.com/`.
2. Choose the **Crawl scope** (same folder is best for documentation).
3. Set **Max pages**.
4. Optional: set a **Content selector** such as `main` or `article`, URL include/exclude patterns, and a **RAG chunk size** (e.g. 2000).
5. Click **Start** and download the results as JSON, CSV or Excel, or fetch them via API.

#### Input example

```json
{
    "startUrls": [{ "url": "https://docs.apify.com/academy" }],
    "crawlScope": "path",
    "maxPages": 200,
    "mode": "auto",
    "cssSelector": "main",
    "excludeUrlPatterns": ["**/changelog/**"],
    "chunkSize": 2000,
    "chunkOverlap": 200
}
```

#### Output example

```json
{
    "url": "https://docs.apify.com/academy/web-scraping-for-beginners",
    "finalUrl": "https://docs.apify.com/academy/web-scraping-for-beginners",
    "title": "Web scraping basics for JavaScript devs | Academy",
    "depth": 0,
    "httpStatus": 200,
    "contentType": "text/html; charset=utf-8",
    "mode": "fast",
    "success": true,
    "markdown": "# Web scraping basics for JavaScript devs\n\nLearn how to...",
    "chunks": [{ "index": 0, "text": "# Web scraping basics..." }],
    "crawledAt": "2026-09-29T04:00:00.000Z"
}
```

### Use it from AI agents and code

The Actor works well as a **tool for AI agents** (via the Apify MCP server, LangChain, LlamaIndex, or the Apify API): give it a URL and a page limit, get Markdown back. Example with the Apify API:

```bash
curl -X POST "https://api.apify.com/v2/acts/tidytools~website-markdown-crawler/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls":[{"url":"https://docs.example.com"}],"maxPages":20}'
```

### Tips

- **Documentation sites**: keep the scope on *Same folder* and set the content selector to `main` or `article` to drop navigation.
- **JavaScript-heavy apps** (React, Vue, Angular): Auto mode switches to the browser automatically; choose *Browser* to force it.
- **Only part of a site**: use include patterns like `https://example.com/blog/**`.
- **Big sites**: raise *Max pages* and *Parallel pages*; you only pay for pages converted.

### Limitations

- Only public `http`/`https` pages; logins, private networks and local addresses are not supported.
- Sites that block automated access are reported as failed and not charged.
- Linked documents up to 15 MB are converted; images are not described.

### Is it legal?

Crawling publicly available pages is generally allowed, but you are responsible for how you use the content. Respect the target site's terms of service, robots rules, copyright and privacy laws.

### Support

Open an issue in the **Issues** tab with the URL and your input. Issues are checked regularly.

# Actor input Schema

## `startUrls` (type: `array`):

Where to start crawling, e.g. the home page of a docs site.

## `crawlScope` (type: `string`):

Which discovered links to follow. "Same folder" keeps the crawl under the start URL's path (best for docs).

## `maxPages` (type: `integer`):

Stop after this many pages. You pay only for pages converted successfully.

## `maxDepth` (type: `integer`):

How many clicks away from a start URL to go. 0 = only the start URLs.

## `mode` (type: `string`):

"Auto" uses the cheap fast mode and switches to a real browser only when a page blocks it or needs JavaScript. "Fast" never uses a browser; "Browser" always does.

## `cssSelector` (type: `string`):

Optional. Convert only this part of each page, e.g. "main" or "article", to drop menus and footers.

## `includeUrlPatterns` (type: `array`):

Optional glob patterns; only matching URLs are crawled. \* matches within a path segment, \*\* matches anything. Example: https://example.com/blog/\*\*

## `excludeUrlPatterns` (type: `array`):

Optional glob patterns to skip, e.g. **/tag/** or \**/login*

## `includeDocuments` (type: `boolean`):

Also convert linked PDF, Word, Excel and CSV files to Markdown.

## `chunkSize` (type: `integer`):

If set (e.g. 2000), each page also gets a "chunks" array split at paragraph boundaries, ready for embeddings. 0 = off.

## `chunkOverlap` (type: `integer`):

Characters repeated between consecutive chunks to keep context.

## `maxConcurrency` (type: `integer`):

How many pages are processed at the same time.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy"
    }
  ],
  "crawlScope": "path",
  "maxPages": 50,
  "maxDepth": 5,
  "mode": "auto",
  "includeDocuments": true,
  "chunkSize": 0,
  "chunkOverlap": 200,
  "maxConcurrency": 10
}
```

# Actor output Schema

## `pages` (type: `string`):

No description

## `full` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://docs.apify.com/academy"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidytools/website-markdown-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://docs.apify.com/academy" }] }

# Run the Actor and wait for it to finish
run = client.actor("tidytools/website-markdown-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy"
    }
  ]
}' |
apify call tidytools/website-markdown-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidytools/website-markdown-crawler"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TpZrf66Jvgs26hMby/builds/jot5XzgD2FtJYX99K/openapi.json
