# Website to Markdown – URL & Site Crawler for LLM/RAG (`gazidev/website-to-markdown`) Actor

Convert any URL, sitemap or whole website to clean Markdown for LLMs, RAG and AI agents. Removes nav, footers and cookie banners, keeps headings, tables and code, adds RAG-ready chunks with token counts, llms.txt generation and an only-changed-pages mode. No browser: fast and cheap.

- **URL**: https://apify.com/gazidev/website-to-markdown.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 page converted to markdowns

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website to Markdown – URL & Site Crawler for LLM/RAG

Convert any **URL, sitemap or whole website to clean Markdown** for LLMs, RAG pipelines, vector databases and AI agents. The Actor removes navigation, headers, footers, sidebars and cookie banners. It keeps headings, lists, tables, code blocks and absolute links, and splits every page into **RAG-ready chunks with token counts and heading paths**. It can also write an **`llms.txt` / `llms-full.txt`** for a site and output **only new or changed pages** on scheduled re-indexing runs.

There is no headless browser, so it is fast and costs **$0.40 per 1,000 pages**.

### What it does

- **Three input modes**:
  - A list of URLs.
  - A **same-domain crawl** from a start URL, with max depth and include/exclude globs.
  - Every URL in a **sitemap**: give a `sitemap.xml`, or just the site and the sitemap is found via robots.txt. Sitemap indexes and `.xml.gz` work.
- **Main-content extraction**:
  - Scores `<main>`, `<article>`, `role=main` and common content containers, with a text-density fallback.
  - Strips scripts, styles, forms, nav, footer, aside and cookie or consent banners.
- **High-quality HTML → Markdown**:
  - Headings, nested lists, GFM tables, fenced code blocks with the language, blockquotes and bold/italic.
  - Absolute links. Images are optional.
- **RAG chunks**: `chunks[]` with `{index, headingPath, text, tokens}`. Chunks split on headings and paragraphs, and code blocks are never cut. The size is configurable (default 800 tokens).
- **llms.txt generator** ([llmstxt.org](https://llmstxt.org) format): pages are grouped by path into sections, as `- [title](url): description`. `llms-full.txt` holds all the Markdown.
- **Only-changed mode** for scheduled runs. ETag, Last-Modified and SHA-256 content hash are remembered between runs, so you pay only for pages that are new or changed.
- **JS-only pages are detected** (`needsBrowser: true`) and **not charged**.
- Polite by default:
  - Respects robots.txt and Crawl-delay.
  - 2 requests per host.
  - Retries with backoff on 429/5xx.
  - Honest User-Agent with a contact address.

### Use cases

- Feed documentation sites into a **RAG / vector database** (Pinecone, Qdrant, Weaviate, pgvector) using the ready-made chunks.
- Give an **AI agent a web-fetch tool** that returns clean Markdown instead of raw HTML.
- Generate **`llms.txt`** for your own site or a competitor's docs.
- **Re-index a knowledge base nightly** with `onlyChanged`, so you pay only for pages that changed.
- Build fine-tuning or evaluation datasets from blogs, docs and help centres.

### Input example

```json
{
  "startUrls": [{ "url": "https://docs.apify.com/sitemap_base.xml" }],
  "crawlMode": "sitemap",
  "maxPages": 200,
  "includeGlobs": ["https://docs.apify.com/academy/**"],
  "chunkSize": 800,
  "generateLlmsTxt": true,
  "onlyChanged": false
}
```

The default input converts 2 pages in a few seconds for less than $0.001.

| Field | Default | Notes |
|---|---|---|
| `startUrls` | 2 sample pages | URLs, a site, or a sitemap URL |
| `crawlMode` | `single` | `single` / `sameDomain` / `sitemap` |
| `maxPages` | 20 | Hard cap on pages fetched |
| `maxDepth` | 2 | Same-domain crawl only |
| `includeGlobs` / `excludeGlobs` | – | e.g. `https://example.com/docs/**`, `**/tag/**` |
| `chunkSize` | 800 | Tokens per chunk, 0 = off |
| `contentMode` | `main` | `main` = article/docs body only, `full` = whole page |
| `keepLinksInMarkdown` / `includeImages` | true / false | Token control |
| `generateLlmsTxt` | false | One llms.txt + llms-full.txt per site |
| `onlyChanged` | false | Output only new or changed pages since the last run |
| `respectRobots` | true | |

### Output example

One dataset row per page. The **Overview**, **Markdown** and **Metadata** views are in the Output tab.

```json
{
  "url": "https://docs.apify.com/academy/scraping-basics-javascript",
  "finalUrl": "https://docs.apify.com/academy/scraping-basics-javascript",
  "status": "ok",
  "httpStatus": 200,
  "title": "Web scraping basics for JavaScript devs | Academy | Apify Documentation",
  "description": "Learn how to use JavaScript to extract information from websites ...",
  "lang": "en",
  "canonical": "https://docs.apify.com/academy/scraping-basics-javascript",
  "wordCount": 979,
  "tokenCount": 1580,
  "markdown": "# Web scraping basics for JavaScript devs\n\n**Learn how to use JavaScript ...**\n\n## What we'll do\n\n- Inspect pages using browser DevTools.\n...",
  "chunks": [
    { "index": 0, "headingPath": "Web scraping basics for JavaScript devs", "text": "# Web scraping basics ...", "tokens": 283 },
    { "index": 1, "headingPath": "Web scraping basics for JavaScript devs > Requirements", "text": "## Requirements ...", "tokens": 350 }
  ],
  "links": ["https://docs.apify.com/get-started", "..."],
  "contentHash": "2573bbdf63d0cd97...",
  "depth": 0,
  "needsBrowser": false,
  "fetchedAt": "2026-10-01T09:00:12+00:00",
  "error": null
}
```

- `status` is one of:
  - `ok`: charged.
  - `needsBrowser`: JS-rendered page, not charged.
  - `failed`: HTTP error or timeout, not charged.
  - `skipped`: PDF or other non-HTML, or blocked by robots.txt. Not charged.
- With `onlyChanged`, each row also has `changeStatus` (`new` / `changed`).
- The key-value store holds `llms.txt`, `llms-full.txt` (with several sites: `llms-<host>.txt`) and an `OUTPUT` run summary.

### Pricing

Pay per event. You pay only for pages that are converted successfully.

| Event | Price |
|---|---|
| Page converted to Markdown (chunks, links and metadata included) | **$0.0004** ($0.40 / 1,000) |
| llms.txt + llms-full.txt generated, per site | $0.005 |
| Actor start | $0.00005 |

Compared with other Store Actors (public Store prices, 2026-10-01):

| Actor | Price per 1,000 pages |
|---|---|
| **Website to Markdown (this Actor)** | **$0.40** |
| apify/web-fetch | $1.50 |
| 6sigmag/fast-website-content-crawler | $3.00 |
| parseforge | $25.00 |

Set **Maximum cost per run** on the run options to cap spending. The Actor stops cleanly when the limit is reached, and the money for llms.txt is held back in advance.

### FAQ

**Does it render JavaScript?**
No. It uses plain HTTP, which is why it is fast and cheap. Pages that only render with JavaScript are flagged `needsBrowser: true` and are not charged. Most docs sites, blogs, news sites and help centres are server-rendered and work well.

**How are tokens counted?**
As characters / 4. This is a fast estimate that is close to the OpenAI and Anthropic tokenizers for English text.

**What happens with PDFs, images and other files?**
Links to them appear in `links[]`, but they are not downloaded or converted. A PDF given as a start URL is returned as `skipped`.

**How does only-changed mode work?**
State is kept in a named key-value store (`wtm-state-…`) derived from your start URLs, or set `stateKey`. Re-runs send `If-None-Match` / `If-Modified-Since` and compare content hashes. Unchanged pages are neither output nor charged.

**Can I crawl only part of a site?**
Yes. Use `includeGlobs`, e.g. `https://example.com/docs/**`, together with `maxDepth` and `maxPages`.

**Is it legal?**
The Actor fetches only the public URLs you supply, respects robots.txt by default and identifies itself. You are responsible for having the rights to use the content you convert, for example under the site's terms and copyright.

### Use with AI agents / Apify MCP

The Actor works as a **web-fetch tool for LLM agents**.

- **Apify MCP server**: add `gazidev/website-to-markdown` to your MCP client (Claude Desktop, Cursor, VS Code) through `https://mcp.apify.com?actors=gazidev/website-to-markdown`. The agent can then call it with `{"startUrls":[{"url":"..."}]}` and get Markdown back.
- **API**: `POST https://api.apify.com/v2/acts/gazidev~website-to-markdown/run-sync-get-dataset-items?token=...` with the input JSON. It returns the rows directly, which suits LangChain / LlamaIndex loaders.
- **RAG pipelines**: embed `chunks[].text` and store `url` and `headingPath` as metadata for citations.

# Actor input Schema

## `startUrls` (type: `array`):

Pages to convert. In 'Same-domain crawl' mode these are the crawl entry points; in 'Sitemap' mode give a sitemap URL (e.g. https://example.com/sitemap.xml) or just the site URL and the sitemap is found via robots.txt.

## `crawlMode` (type: `string`):

'Only these URLs' converts exactly the start URLs. 'Same-domain crawl' follows links on the same site up to Max depth. 'Sitemap' converts the URLs listed in the site's sitemap(s).

## `maxPages` (type: `integer`):

Hard cap on pages fetched in this run (start URLs included). Keeps cost predictable: at most maxPages x the page price.

## `maxDepth` (type: `integer`):

Same-domain crawl only: how many clicks away from a start URL to follow (0 = start URLs only).

## `includeGlobs` (type: `array`):

Only crawl / take from the sitemap URLs that match one of these globs, e.g. https://docs.example.com/guides/\*\*. \*\* = anything, \* = anything except '/'. Empty = all URLs on the domain.

## `excludeGlobs` (type: `array`):

Skip URLs matching any of these globs, e.g. **/tag/**, \*\*/page/*, \*\*?replytocom=*.

## `chunkSize` (type: `integer`):

Split each page's Markdown into chunks of at most this many tokens (estimated as characters / 4) on heading and paragraph boundaries. Every chunk carries its heading path. 0 = no chunks. Included in the page price.

## `contentMode` (type: `string`):

'Main content' removes navigation, headers, footers, sidebars, cookie banners and share widgets and keeps the article / docs body. 'Full page' converts the whole <body>.

## `keepLinksInMarkdown` (type: `boolean`):

Render links as \[text]\(absolute URL). Turn off for plain text-only Markdown (fewer tokens).

## `includeImages` (type: `boolean`):

Render images as !\[alt]\(absolute URL). Off by default to save tokens.

## `outputLinks` (type: `boolean`):

Add every absolute link found on the page (including PDFs and other files, which are listed but not fetched) to the dataset row.

## `removeSelectors` (type: `string`):

Optional comma-separated CSS selectors removed before conversion, e.g. .newsletter-box, #comments.

## `generateLlmsTxt` (type: `boolean`):

Build an llms.txt index (llmstxt.org format: sections with - [title](url): description) and an llms-full.txt with all Markdown, one pair per site, saved to the key-value store. Charged once per site.

## `onlyChanged` (type: `boolean`):

For scheduled re-indexing: remembers each page's ETag / Last-Modified / content hash in a named key-value store and outputs (and charges) only pages that are new or changed since the previous run with the same start URLs.

## `stateKey` (type: `string`):

Optional name for the change-tracking state (letters, digits, dashes). By default it is derived from the start URLs and crawl mode.

## `respectRobots` (type: `boolean`):

Skip URLs disallowed by robots.txt and honour Crawl-delay (max 10 s).

## `maxConcurrency` (type: `integer`):

Parallel requests in total.

## `maxConcurrencyPerHost` (type: `integer`):

Parallel requests to one host. Keep it low to be polite.

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout.

## `maxRetries` (type: `integer`):

Retries for timeouts, connection errors, 429 and 5xx responses (with backoff).

## `userAgent` (type: `string`):

Custom User-Agent header. Leave empty for the default, which identifies this Actor with a contact address.

## `proxyConfiguration` (type: `object`):

Optional. Not needed for most sites.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy/scraping-basics-javascript"
    },
    {
      "url": "https://www.paulgraham.com/greatwork.html"
    }
  ],
  "crawlMode": "single",
  "maxPages": 3,
  "maxDepth": 2,
  "includeGlobs": [],
  "excludeGlobs": [],
  "chunkSize": 800,
  "contentMode": "main",
  "keepLinksInMarkdown": true,
  "includeImages": false,
  "outputLinks": true,
  "generateLlmsTxt": false,
  "onlyChanged": false,
  "respectRobots": true,
  "maxConcurrency": 10,
  "maxConcurrencyPerHost": 2,
  "requestTimeoutSecs": 30,
  "maxRetries": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `markdown` (type: `string`):

No description

## `results` (type: `string`):

No description

## `llmsTxt` (type: `string`):

No description

## `llmsFullTxt` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://docs.apify.com/academy/scraping-basics-javascript"
        },
        {
            "url": "https://www.paulgraham.com/greatwork.html"
        }
    ],
    "maxPages": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/website-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        { "url": "https://docs.apify.com/academy/scraping-basics-javascript" },
        { "url": "https://www.paulgraham.com/greatwork.html" },
    ],
    "maxPages": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/website-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://docs.apify.com/academy/scraping-basics-javascript"
    },
    {
      "url": "https://www.paulgraham.com/greatwork.html"
    }
  ],
  "maxPages": 3
}' |
apify call gazidev/website-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/website-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/N4koD9TBqqKj1uB2z/builds/gcIT08Nc2SKhr5v8p/openapi.json
