# Website Content Crawler API: Web Pages to Markdown for RAG (`sauliusautomatesit/website-content-crawler-api`) Actor

Website content crawler API: turn any web page or whole site into clean Markdown and text for RAG, LLMs and AI agents. Crawls links and sitemaps, renders JavaScript pages when needed. RAG Web Browser alternative at $1 per 1,000 pages.

- **URL**: https://apify.com/sauliusautomatesit/website-content-crawler-api.md
- **Developed by:** [Saulius AutomatesIT](https://apify.com/sauliusautomatesit) (community)
- **Categories:** AI, Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.06 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Content Crawler API: Web Pages to Markdown for RAG

Turn any web page, docs section or whole website into clean Markdown and text for RAG pipelines, LLM
context, vector databases and AI agents. Give it URLs; it fetches them, strips navigation, footers, scripts
and cookie banners, keeps the main content, and returns Markdown with headings, lists, links, code blocks and
tables. Crawl links and sitemaps when you need a whole site. **$1.25 per 1,000 pages**, failed and empty pages
free.

A cheaper alternative to RAG Web Browser ($2.28 per 1,000 fetched pages) and to compute-billed website
crawlers: you pay a flat price per page with text, so a big crawl costs what you expect.

### What you get per page

| Field | Example |
|---|---|
| `url`, `requestedUrl`, `statusCode` | final URL after redirects, the URL you gave, 200 |
| `title`, `description` | page title and meta description |
| `markdown` | the page as clean Markdown: `# Title`, `## Sections`, lists, links, fenced code, tables |
| `language`, `author`, `publishedAt` | from the page's meta tags and JSON-LD |
| `canonicalUrl`, `image` | canonical link and share image |
| `wordCount` | words in the text |
| `mainContentOnly` | true when only the main article or docs block was kept |
| `depth`, `referrerUrl`, `startUrl` | where the crawler found the page |
| `text`, `html`, `links` | optional: plain text, the cleaned HTML, every link on the page |

### Input

- **Start URLs**: pages to read. A bare domain (`apify.com`) works.
- **Crawl depth**: 0 (default) reads only your URLs. 1 also reads the pages they link to, and so on. Links
  are followed on the same site and inside the start URL's folder: `https://docs.apify.com/platform/actors`
  crawls `https://docs.apify.com/platform/...`.
- **Use sitemaps**: add every page in the site's `sitemap.xml` (and sitemaps listed in `robots.txt`) inside
  the start folder. The fastest way to get a whole docs site.
- **Max pages per start URL** (100) and **max pages in total**.
- **Only URLs containing / Skip URLs containing**: simple text filters such as `/blog/` or `?page=`.
- **Main content only** (default on): Mozilla Readability keeps the article or docs body on content pages;
  home pages and listings keep everything except navigation, footers, scripts and cookie banners.
- **Remove elements**: extra CSS selectors to drop, such as `.sidebar, #comments`.
- **Include plain text / cleaned HTML / links**: extra fields, same price.

```json
{
  "startUrls": ["https://docs.apify.com/platform/"],
  "maxCrawlDepth": 2,
  "useSitemaps": true,
  "maxPagesPerStartUrl": 500
}
```

### Pricing

**$1.25 per 1,000 pages** with text ($0.00125 a page), plus Apify's $0.00005 Actor start. Apify Store
discounts apply: Bronze $1.19, Silver $1.13, Gold and above $1.06 per 1,000. No charge for pages that fail, are not
found, are blocked, are not HTML (PDF, images, files), have no text or only show text after JavaScript runs.

Examples: a 300 page docs site = $0.38. 100,000 blog posts for a vector database = $125. One page for an AI
agent = $0.0013.

Speed: a run gets one CPU core per 4 GB of memory. The default 1 GB reads about 35 pages a minute; give big
crawls 4 GB (about 4 times faster). The price per page is the same at any memory.

### Use cases

- **RAG and vector databases**: load documentation, help centres, blogs and knowledge bases as Markdown
  chunks with titles and source URLs.
- **AI agents and MCP**: fetch a page as Markdown inside an agent loop; small, cheap, predictable.
- **LLM training and evaluation data**: clean text of whole sites with language and dates.
- **Content monitoring**: schedule a crawl of a competitor's blog or a docs section and diff the Markdown.
- **SEO audits**: titles, descriptions, canonical links, word counts and internal links per page.

### Error rows (free)

| `error` | Meaning |
|---|---|
| `NOT_FOUND` | the page answered 404 or 410 |
| `HTTP_403`, `HTTP_429` and so on | the site refused the page |
| `BLOCKED` | a bot check (Cloudflare challenge, captcha) |
| `NOT_HTML` | a PDF, image, archive or other file |
| `NEEDS_JAVASCRIPT` | the server sends an empty app shell; the text appears only after JavaScript runs |
| `EMPTY` | the page has no text |
| `FAILED` | the site could not be reached |
| `NO_DATA` | the input had no usable URL |

### API, MCP and integrations

Call it from Python, JavaScript, Make, Zapier, n8n, LangChain or LlamaIndex like any Apify Actor, or give
an AI agent this MCP server: `https://mcp.apify.com/?tools=sauliusautomatesit/website-content-crawler-api`.

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("sauliusautomatesit/website-content-crawler-api").call(
    run_input={"startUrls": ["https://docs.apify.com/platform/"], "maxCrawlDepth": 1, "maxPagesPerStartUrl": 50}
)
for page in client.dataset(run["defaultDatasetId"]).iterate_items():
    if page.get("type") != "error":
        print(page["url"], page["wordCount"])
```

### Notes

- Pages are read as the server sends them (plain HTTP with a Chrome fingerprint through Apify datacenter
  proxies). Server rendered sites (docs, blogs, news, Wikipedia, most CMSs and Next.js sites) work; client
  only app shells come back as free `NEEDS_JAVASCRIPT` rows.
- The crawler stays on the start URL's site and folder. Use several start URLs for several sections.
- Respect each site's terms and robots rules for your use; this Actor reads only public pages.

### Related

- [Contact Details Scraper API](https://apify.com/sauliusautomatesit/contact-details-scraper-api): emails,
  phones and social profiles from company websites.
- [YouTube Transcript Scraper](https://apify.com/sauliusautomatesit/youtube-transcript-scraper): video
  transcripts for the same RAG pipelines.

# Actor input Schema

## `startUrls` (type: `array`):

Web pages to read. With crawl depth 0 (default) you get exactly these pages; raise the depth to follow their links. A bare domain (apify.com) works too.

## `maxCrawlDepth` (type: `integer`):

0 reads only the start URLs. 1 also reads the pages they link to, 2 the pages those link to, and so on. Links are followed only on the same site and inside the start URL's folder (https://docs.apify.com/platform/actors crawls https://docs.apify.com/platform/...).

## `maxPagesPerStartUrl` (type: `integer`):

Stop crawling a start URL's site after this many pages.

## `maxPages` (type: `integer`):

Stop the run after this many charged pages. 0 means no limit beyond the per start URL limit.

## `useSitemaps` (type: `boolean`):

Also read the pages listed in the site's sitemap.xml (and sitemaps named in robots.txt), within the start URL's folder. The fastest way to get a whole site or docs section. Used when Max pages per start URL is 10 or more.

## `urlMustContain` (type: `array`):

Crawl only links whose URL contains at least one of these texts, for example /blog/ or /docs/. Start URLs are always read.

## `urlMustNotContain` (type: `array`):

Do not crawl links whose URL contains any of these texts, for example /tag/ or ?page=.

## `mainContentOnly` (type: `boolean`):

On articles and docs, keep the main text and drop sidebars and related links (Mozilla Readability). Pages where the main block is only a small part of the text, like home pages, keep everything except navigation, footers, scripts and cookie banners.

## `removeSelectors` (type: `string`):

Extra elements to drop before conversion, for example .sidebar, .ads, #comments

## `includeText` (type: `boolean`):

Add a text field with the page's text without Markdown.

## `includeHtml` (type: `boolean`):

Add the cleaned HTML that the Markdown was made from.

## `includeLinks` (type: `boolean`):

Add every link found on the page (absolute URLs).

## `maxConcurrency` (type: `integer`):

How many pages to load at once. Capped at one per 128 MB of run memory (8 at the default 1 GB). A run gets one CPU core per 4 GB of memory, so give big crawls 2 to 4 GB: same price, faster.

## Actor input object example

```json
{
  "startUrls": [
    "https://docs.apify.com/platform/actors"
  ],
  "maxCrawlDepth": 0,
  "maxPagesPerStartUrl": 100,
  "maxPages": 0,
  "useSitemaps": false,
  "mainContentOnly": true,
  "includeText": false,
  "includeHtml": false,
  "includeLinks": false,
  "maxConcurrency": 20
}
```

# Actor output Schema

## `results` (type: `string`):

One item per page. Download as JSON, CSV or Excel, or read it from this API endpoint.

## `summary` (type: `string`):

Pages delivered, errors and requests.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://docs.apify.com/platform/actors"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("sauliusautomatesit/website-content-crawler-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://docs.apify.com/platform/actors"] }

# Run the Actor and wait for it to finish
run = client.actor("sauliusautomatesit/website-content-crawler-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://docs.apify.com/platform/actors"
  ]
}' |
apify call sauliusautomatesit/website-content-crawler-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sauliusautomatesit/website-content-crawler-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7Kpy0e1BUj7Q1XflU/builds/aavMGdNTidrNaZr9C/openapi.json
