# Website to Markdown for AI: Clean Content, RAG Chunks (`swiftkit/web-to-markdown`) Actor

Turn web pages or whole websites into clean Markdown for LLMs and RAG: main content only (no menus, ads or share buttons), title, author, date, word and token counts, optional chunks. Crawl a site or give a list. Fixed price per page.

- **URL**: https://apify.com/swiftkit/web-to-markdown.md
- **Developed by:** [SwiftKit](https://apify.com/swiftkit) (community)
- **Categories:** AI, Developer tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 page converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website to Markdown for AI: Clean Content, RAG Chunks

Feed web pages to ChatGPT, Claude, your RAG pipeline or your AI agent **without the junk**. Give it
a list of pages or a site to crawl, and get back:

- **Clean Markdown of the main content:** headings, lists, tables, code blocks and links kept;
  menus, footers, cookie banners, ads and share buttons removed
- **Metadata:** title, description, author, publish date, language, site name, image, canonical URL
- **Word count and approximate token count** per page, so you can plan context windows and costs
- **Optional RAG chunks:** pieces of about N tokens, split on paragraph boundaries with overlap
- **Fixed price per page.** No compute-time surprises.

Main content is found with **Mozilla Readability**, the engine behind Firefox Reader View. Links and
images are made absolute, so the Markdown works anywhere.

### Who it's for

- **RAG and chatbots:** index docs, help centers and blogs as clean chunks.
- **AI agents:** read any page as Markdown with one call.
- **Content teams and researchers:** archive or analyze articles in a readable format.
- **Training and evaluation data:** clean text with word counts and dates.

### Input

| Option | Default | What it does |
|---|---|---|
| URLs | – | Pages to convert, or start pages to crawl |
| Crawl the site | off | Follow links on the same site |
| Max pages | 50 | Stop after this many converted pages |
| Max link depth | 3 | How far from the start pages to crawl |
| Only / skip URLs containing | – | Keep the crawl to e.g. `/docs/` |
| Content | Main content | Or the whole page minus menus and footer |
| Chunk size, overlap | 0, 50 | Tokens per RAG chunk (0 = off) |
| Include plain text, links | off | Extra fields |
| Respect robots.txt | on | Skip pages the site asks bots not to visit |

### Output

Real result, shortened:

```json
{
  "url": "https://github.blog/engineering/the-cost-of-saying-yes-has-changed/",
  "status": "ok",
  "title": "The cost of saying yes has changed",
  "author": "Dalia Abuadas",
  "published": "2026-07-17T16:46:47+00:00",
  "language": "en-US",
  "markdown": "The cost of writing code dropped; the cost of owning it didn’t. …",
  "wordCount": 1174,
  "approxTokens": 1877,
  "chunks": [
    { "index": 0, "text": "…", "approxTokens": "…" }
  ]
}
```

| status | meaning |
|---|---|
| `ok` | Converted. |
| `unreachable` | The page didn't load. Not charged. |
| `not_html` | Not a web page (PDF, image…). Not charged. |
| `blocked_by_robots_txt` | The site asks bots not to visit. Not charged. |
| `blocked_by_bot_protection` | The site showed a bot check. We don't try to get around it. Not charged. |

### Pricing

You pay **per page converted**, the same for every page size. Failed and blocked pages are free.
See the Pricing tab.

### Limits, honestly

- **No browser is used.** That keeps it fast and cheap, but pages that build their content with
  JavaScript after loading (some single-page apps) may come back thin, and crawling them can find
  few links. Most blogs, docs, news and company sites work well.
- Token counts are approximate (about 4 characters per token). Exact counts depend on your model.
- Pages behind a login, a paywall or a bot check aren't converted.
- You're responsible for how you use the content; respect each site's terms and copyright.

### More tools from SwiftKit

- [Sitemap Extractor & Broken Link Checker](https://apify.com/swiftkit/sitemap-urls): every URL from a site’s sitemaps, plus 404 and redirect checks
- [RSS Feed Reader](https://apify.com/swiftkit/rss-feeds): news, blogs and podcasts from any feed or website, only new items
- [Website Screenshot Tool](https://apify.com/swiftkit/website-screenshots): bulk full-page and mobile screenshots, PNG/JPEG/PDF
- [PDF to Text & Markdown](https://apify.com/swiftkit/pdf-text): clean text from PDF links for AI, priced per page

### Questions?

Open an issue on the Issues tab.

# Actor input Schema

## `startUrls` (type: `array`):

Pages to convert, or the start page(s) of a site to crawl.

## `crawl` (type: `boolean`):

Follow links on the same site (and its subdomains) from the start pages.

## `maxPages` (type: `integer`):

Stop after converting this many pages.

## `maxDepth` (type: `integer`):

Crawl: how many clicks away from the start pages to go.

## `includePatterns` (type: `array`):

Crawl: keep only URLs matching any of these texts or regular expressions, e.g. /docs/ or /blog/

## `excludePatterns` (type: `array`):

Crawl: skip URLs matching any of these, e.g. /tag/ or ?page=

## `contentMode` (type: `string`):

Main content uses Mozilla Readability, the engine behind Firefox Reader View.

## `chunkSize` (type: `integer`):

Also split each page into chunks of about this many tokens for RAG. 0 = no chunks.

## `chunkOverlap` (type: `integer`):

Tokens repeated between neighbouring chunks.

## `includeText` (type: `boolean`):

Also add the content as plain text.

## `includeLinks` (type: `boolean`):

Add every link found on the page.

## `respectRobotsTxt` (type: `boolean`):

Skip pages the site asks bots not to visit.

## `maxConcurrency` (type: `integer`):

Parallel requests. Keep it modest to be polite.

## Actor input object example

```json
{
  "startUrls": [
    "https://github.blog/engineering/"
  ],
  "crawl": false,
  "maxPages": 50,
  "maxDepth": 3,
  "contentMode": "article",
  "chunkSize": 0,
  "chunkOverlap": 50,
  "includeText": false,
  "includeLinks": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

One row per page with its Markdown.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://github.blog/engineering/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("swiftkit/web-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://github.blog/engineering/"] }

# Run the Actor and wait for it to finish
run = client.actor("swiftkit/web-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://github.blog/engineering/"
  ]
}' |
apify call swiftkit/web-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,swiftkit/web-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5Q3lgEbUePQJX7EAr/builds/ZXLMjm24XqmyXtXUk/openapi.json
