# Website to Markdown Crawler for AI, RAG & LLMs (`dima_kadirovich/website-to-markdown`) Actor

Crawl any website and get clean main-content Markdown for every page: no menus or footers. RAG-ready chunks, llms.txt, and an 'only changed pages' mode for cheap scheduled refreshes. $1 per 1,000 pages, all included.

- **URL**: https://apify.com/dima\_kadirovich/website-to-markdown.md
- **Developed by:** [Cronexa Data Tools](https://apify.com/dima_kadirovich) (community)
- **Categories:** AI, MCP servers, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Website to Markdown Crawler do?

It **crawls any website and turns every page into clean Markdown**, keeping only the **main content** without menus, headers, footers, cookie banners or sidebars. The output is ready for **AI, RAG pipelines, vector databases, LLM fine-tuning and custom GPTs**.

- 📝 **Clean main-content Markdown**, with headings, lists, tables, links and **fenced code blocks** kept
- ✂️ **RAG-ready chunks** (optional): each page split into chunks of your chosen token size, cut at headings and paragraphs (never inside a code block), with overlap
- 📚 **`llms.txt` and `llms-full.txt`** generated for the site, following the [llmstxt.org](https://llmstxt.org) format
- 🔄 **"Only new or changed pages" mode**: schedule it daily or weekly, and only pages whose content changed are saved. **Unchanged pages are free**, so keeping your RAG index fresh costs almost nothing.
- 💸 **$1 per 1,000 pages. Everything included**: no compute, proxy or browser fees on top.

It's fast because it reads the HTML directly instead of starting a full browser. It follows links **and** the sitemap, respects `robots.txt`, and **slows down automatically** if a site asks it to.

### What can I use it for?

- **Chatbots and RAG**: feed product docs, help centers or knowledge bases into your vector database (Pinecone, Qdrant, Weaviate, pgvector…)
- **Custom GPTs / Claude Projects**: download `llms-full.txt` and upload one file with a whole documentation site
- **Keeping AI knowledge up to date**: schedule refresh runs and re-index only the pages that changed
- **Content migration and archiving**: move a website's content to Markdown for a new CMS, Notion or Obsidian
- **SEO and content analysis**: word counts, titles and descriptions for every page

### How do I use it?

1. Enter a start URL, for example `https://docs.example.com/`.
2. (Optional) Limit it with **Only crawl URLs matching**, for example `/docs/`.
3. (Optional) Set **Chunk size** (for example `800`) if you're loading a vector database.
4. Click **Start**, then download the pages as JSON, CSV or Excel, or open `llms-full.txt` from the Output tab.

For **scheduled refreshes**, turn on **Only new or changed pages** and add a Schedule. Each run saves only what changed, and marks each row as `new` or `changed`.

#### Example input

```json
{
  "startUrls": [{ "url": "https://www.python-httpx.org/" }],
  "maxPages": 500,
  "includeUrlPatterns": [],
  "chunkSize": 800,
  "chunkOverlap": 100,
  "onlyChangedPages": true
}
```

### Output example

One row per page (Markdown and chunk text shortened here):

````json
{
  "url": "https://www.python-httpx.org/quickstart/",
  "title": "QuickStart - HTTPX",
  "description": "A next-generation HTTP client for Python.",
  "language": "en",
  "markdown": "# QuickStart\n\nFirst, start by importing HTTPX:\n\n```\n>>> import httpx\n```\n\nNow, let’s try to get a webpage.\n\n```\n>>> r = httpx.get('https://httpbin.org/get')\n...",
  "wordCount": 1742,
  "tokenEstimate": 3627,
  "extraction": "main-content",
  "contentHash": "b3e8cef33b8b14e49d1e6d98bba23e7420af901b20ce34365bd53a34a181659d",
  "changeStatus": "new",
  "httpStatus": 200,
  "chunks": [
    { "index": 0, "heading": "QuickStart", "text": "# QuickStart\n\nFirst, start by importing HTTPX: ...", "tokenEstimate": 797 }
  ]
}
````

The Output tab also links **`llms.txt`** (an index of all pages) and **`llms-full.txt`** (all pages in one Markdown file).

### How much does it cost?

| What | Price |
|---|---|
| Page saved | **$0.001** ($1 per 1,000 pages) |
| Unchanged pages in refresh mode | **Free** |
| Chunks, llms.txt, sitemap, robots.txt | **Free** |
| Platform usage (compute) | **Included** |

A 500-page documentation site costs **$0.50**. You can set a **maximum cost per run**, and the Actor stops exactly at your budget.

### Tips

- **Documentation sites**: use *Only crawl URLs matching* (for example `/docs/`) to skip the blog and marketing pages.
- **Chunk size**: 500–1,000 tokens works well for most embedding models. The token estimate uses ~4 characters per token.
- **Plain text**: turn off *Keep links in Markdown*, or use the `text` field.

### Limitations

- Pages are read from their HTML. Content that appears **only after JavaScript runs** (some single-page apps) may be incomplete. Most docs, blogs and company sites work well.
- Login-protected pages are not crawled.
- Some websites block crawlers. Try the Proxy option if pages fail.

### Is it legal?

It collects publicly available web pages and respects `robots.txt` by default. Make sure your use of the content respects copyright and each website's terms.

### Questions or problems?

Open an issue in the **Issues** tab, and it will be answered quickly.

### Use it from AI assistants (Claude, ChatGPT, Cursor)

AI agents can run this Actor as a tool through the [Apify MCP server](https://mcp.apify.com). Add this to your MCP client (Claude Desktop, Claude Code, Cursor, VS Code…) and sign in with Apify in the browser when asked:

```json
{
  "mcpServers": {
    "website-to-markdown": { "url": "https://mcp.apify.com?tools=dima_kadirovich/website-to-markdown" }
  }
}
```

Then just ask, for example:

- *"Read the docs at https://www.python-httpx.org/ (up to 30 pages) and explain how to set timeouts."*
- *"Turn this help center into Markdown chunks of 800 tokens for my vector database."*

The agent fills in the input, runs the Actor and reads the results. You pay the same per-result price.

### More tools from Dima Data Tools

- [Medium Articles Scraper & Monitor](https://apify.com/dima_kadirovich/medium-articles-scraper): Medium articles by tag, author, or publication as clean Markdown, with "only new" monitoring
- [Bulk Image Downloader](https://apify.com/dima_kadirovich/bulk-image-downloader): every image from any web page, as download links or ZIP, with duplicates and icons removed
- [Website SEO Audit & Broken Link Checker](https://apify.com/dima_kadirovich/website-seo-audit): crawl a site, score every page 0–100, find broken links, and get a shareable HTML report
- [Website Tech Stack & Domain Lookup](https://apify.com/dima_kadirovich/tech-stack-domain-lookup): technologies, email provider, SPF/DMARC, SaaS tools, SSL expiry and WHOIS for any list of domains

# Actor input Schema

## `startUrls` (type: `array`):

Where to start crawling, for example a documentation homepage. The crawler follows links and the sitemap within the same website.

## `maxPages` (type: `integer`):

Upper limit of pages crawled.

## `includeUrlPatterns` (type: `array`):

Regular expressions. If set, only URLs matching at least one are crawled, for example `/docs/` to get only documentation pages.

## `excludeUrlPatterns` (type: `array`):

Regular expressions for URLs to skip, for example `/blog/tag/` or `\?page=`.

## `onlyChangedPages` (type: `boolean`):

Remember each page's content from earlier runs with the same start URLs, and save only pages that are new or changed. Unchanged pages are skipped and not charged, which makes scheduled refreshes of your RAG index very cheap.

## `chunkSize` (type: `integer`):

Also split each page into chunks of about this many tokens for vector databases, keeping headings and code blocks together. Leave empty or 0 for no chunks. Typical: 500–1000.

## `chunkOverlap` (type: `integer`):

How many tokens each chunk repeats from the previous one, to keep context.

## `generateLlmsTxt` (type: `boolean`):

Create an llms.txt index of the crawled pages and an llms-full.txt with all content, following the llmstxt.org format.

## `includeLinks` (type: `boolean`):

Keep hyperlinks as [text](url). Turn off for plain text-like Markdown.

## `useSitemap` (type: `boolean`):

Also find pages from the website's sitemap.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that robots.txt disallows.

## `maxDepth` (type: `integer`):

How many clicks away from the start page to crawl. 0 = only the start pages.

## `includeSubdomains` (type: `boolean`):

Also crawl subdomains like docs.example.com when starting from example.com.

## `saveHtml` (type: `boolean`):

Add the original HTML of each page to the output.

## `proxyConfiguration` (type: `object`):

Only needed if a website blocks the crawler.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.python-httpx.org/"
    }
  ],
  "maxPages": 100,
  "onlyChangedPages": false,
  "chunkOverlap": 100,
  "generateLlmsTxt": true,
  "includeLinks": true,
  "useSitemap": true,
  "respectRobotsTxt": true,
  "maxDepth": 20,
  "includeSubdomains": false,
  "saveHtml": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `pages` (type: `string`):

One row per page: URL, title, description, Markdown, text, token estimate, optional chunks.

## `llmsTxt` (type: `string`):

Index of the crawled pages in llms.txt format.

## `llmsFullTxt` (type: `string`):

All page content in one Markdown file, ready to paste into an LLM.

## `summary` (type: `string`):

Pages saved, unchanged pages skipped, failures, and total token estimate.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.python-httpx.org/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("dima_kadirovich/website-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.python-httpx.org/" }] }

# Run the Actor and wait for it to finish
run = client.actor("dima_kadirovich/website-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.python-httpx.org/"
    }
  ]
}' |
apify call dima_kadirovich/website-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dima_kadirovich/website-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0KvBke3ntGGJvY9ym/builds/RcRvgWwv4tpDrpShe/openapi.json
