# Site to Markdown Crawler: Website Content Crawler Alternative (`madrasco/site-markdown-crawler`) Actor

Crawls a website and converts each page to clean Markdown for AI and RAG use. Plain HTTP by default (256 MB), spend cap on, priced per page.

- **URL**: https://apify.com/madrasco/site-markdown-crawler.md
- **Developed by:** [Jack Valmadre](https://apify.com/madrasco) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 page crawleds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Site to Markdown Crawler

Crawls a website and converts each page to clean Markdown, ready for AI ingestion, RAG pipelines or offline archiving. Uses plain HTTP (no browser), so it runs in 512 MB by default and keeps costs predictable.

**Default protections:** a page cap (500 pages) and a spend cap (US$5.00) are on by default. An unattended run stops at whichever limit it hits first. Adjust either value in the input, or set to 0 to disable.

### What you get

One dataset row per crawled page:

| Field | Description |
|---|---|
| `url` | URL as requested |
| `finalUrl` | URL after redirects |
| `httpStatus` | HTTP status code |
| `title` | Page title, if found |
| `markdown` | Page body as Markdown: headings, lists, emphasis, tables, code blocks |
| `wordCount` | Word count of the extracted body |
| `status` | `ok`, `no-content`, `blocked`, `disallowed-by-robots`, or `error` |
| `error` | Error message, when status is `error` |
| `fetchedAt` | ISO timestamp of the fetch |

`no-content` is set when the page is a navigation or index page with too little prose to be useful.

### How it crawls

- Reads `robots.txt` and honours `Crawl-delay` and per-agent disallow rules before fetching any page.
- Seeds the URL queue from the site's `sitemap.xml` (found via `robots.txt` or `/sitemap.xml`), then follows same-domain `<a href>` links in BFS order.
- Extracts body text and converts to Markdown using trafilatura (Apache-2.0). Navigation, header and footer boilerplate are stripped at the library level.
- Stops at the page cap or spend cap, whichever comes first.

### When to use this

- Feeding a documentation site, blog or wiki into a RAG pipeline or AI assistant.
- Archiving a site as structured Markdown.
- Any job where you want clean text at a known cost rather than raw HTML.

### Pricing

**US$0.0005 per page crawled** (US$0.50 per 1,000 pages), plus a one-time **US$0.00005 start charge** per run.

Only pages with successfully extracted Markdown content are charged (`status: ok`). Pages that are blocked, disallowed by robots.txt, have no extractable text, or are skipped due to the spend cap are not charged.

The default spend cap is US$5.00, which covers up to 10,000 pages per run. Adjust `maxTotalChargeUsd` to match your job.

#### Why this is cheaper than Apify's Website Content Crawler

Apify's Website Content Crawler defaults to 8 GB of memory and bills by compute time. On a plain-HTTP crawl, that is roughly US$0.20 per 1,000 pages at its minimum; with a browser it is US$0.50–5 per 1,000 pages (Apify's own published range). An unmonitored run can accumulate large charges.

This actor runs at 512 MB, uses plain HTTP by default, and stops at the page and spend caps. The underlying platform compute cost on our runs was US$0.09–0.14 per 1,000 pages; the US$0.50 per 1,000 pages list price covers overhead and keeps your charges predictable.

| Site | Pages | Run time | Platform cost |
|---|---|---|---|
| flask.palletsprojects.com | 76 ok / 78 total | 252 s | US$0.0070 (US$0.090/1k) |
| www.sphinx-doc.org | 202 ok / 202 total | 754 s | US$0.0222 (US$0.110/1k) |
| docs.djangoproject.com | 204 ok / 204 total | 968 s | US$0.0275 (US$0.135/1k) |

These platform costs are from real Apify runs (actor HrXtDuKlhELYgA9UT, builds 0.1.3–0.1.7) at 512 MB. Your actual charge is the list price (US$0.0005/page), not the underlying compute cost. Your cost will vary with page count, page size and server speed.

### Input

| Field | Default | Description |
|---|---|---|
| `startUrls` | *(required)* | One or more URLs to start crawling from |
| `maxPages` | 500 | Stop after this many pages |
| `maxTotalChargeUsd` | 5.0 | Stop before the run charge exceeds this (pay-per-event billing only) |
| `sameDomainOnly` | true | Only follow links on the same domain |
| `useSitemap` | true | Seed the URL queue from the site's sitemap.xml |
| `timeoutSecs` | 20 | Per-page request timeout in seconds |
| `maxConcurrency` | 2 | Parallel requests (default 2 is safe at 512 MB; increase with memory: up to 4 at 1 GB) |

### Use with AI agents

Pass `startUrls` as a list of `{"url": "..."}` objects, plain URL strings, or a single URL string. All discovered pages are returned in the default dataset as structured rows. The `markdown` field is ready to insert directly into a prompt or embed in a vector store. Filter by `status: "ok"` to drop navigation pages and errors before chunking.

### Robots.txt and terms of service

This actor reads every site's `robots.txt` before crawling and honours `Disallow` rules for `User-agent: *` and `User-agent: MadrascoSiteMarkdownCrawler` (RFC 9309). Sites that block our user agent or return a captcha or 403 are recorded as `blocked`; they are not retried or bypassed.

You are responsible for ensuring your use of this actor complies with the target site's terms of service.

### Limitations

- **JavaScript-rendered content:** the actor uses plain HTTP, not a browser. Pages that require JavaScript to render their main content will return incomplete or empty Markdown. For those sites, expect `no-content` results.
- **Slow servers:** the default 20-second per-page timeout may cause errors on slow hosts. Raise `timeoutSecs` if you see a high error rate.
- **Large sites:** the default page cap is 500. Set `maxPages` and `maxTotalChargeUsd` to match the job; a test run on a sample of URLs is a good first step for unfamiliar sites.
- **Sparse pages:** index and search pages with little prose are returned with `status: no-content` rather than discarded, so you can see what the crawler found.

### Publisher

Built by Madrasco. Support: use the Issues tab on this actor's page.

# Actor input Schema

## `startUrls` (type: `array`):

One or more URLs to start crawling from. The crawler follows links within the same domain (and uses the site's sitemap). Each URL is fetched once; discovered pages are added to the queue.

## `maxPages` (type: `integer`):

Stop crawling after this many pages (regardless of spend cap). Default: 500.

## `maxTotalChargeUsd` (type: `number`):

Stop before the total run charge exceeds this amount. Set to 0 to disable. Applies only when pay-per-event pricing is active.

## `sameDomainOnly` (type: `boolean`):

Only follow links within the same domain as the start URLs. Disabling this allows cross-domain crawls.

## `useSitemap` (type: `boolean`):

Fetch the site's sitemap.xml (from robots.txt or /sitemap.xml) to seed the URL queue.

## `timeoutSecs` (type: `integer`):

Give up on a single page after this long.

## `maxConcurrency` (type: `integer`):

How many pages to fetch and parse at once. Default 2 is safe at the actor's minimum 512 MB memory. Increase together with memory (up to 4 at 1 GB).

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://flask.palletsprojects.com/"
    }
  ],
  "maxPages": 15,
  "maxTotalChargeUsd": 5,
  "sameDomainOnly": true,
  "useSitemap": true,
  "timeoutSecs": 20,
  "maxConcurrency": 2
}
```

# Actor output Schema

## `pages` (type: `string`):

One dataset row per page: URL, title, Markdown content, word count and status.

## `summary` (type: `string`):

Counts of pages by status and whether the page or spend cap was reached.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://flask.palletsprojects.com/"
        }
    ],
    "maxPages": 15
};

// Run the Actor and wait for it to finish
const run = await client.actor("madrasco/site-markdown-crawler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://flask.palletsprojects.com/" }],
    "maxPages": 15,
}

# Run the Actor and wait for it to finish
run = client.actor("madrasco/site-markdown-crawler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://flask.palletsprojects.com/"
    }
  ],
  "maxPages": 15
}' |
apify call madrasco/site-markdown-crawler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,madrasco/site-markdown-crawler"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HrXtDuKlhELYgA9UT/builds/EM4X11Ud1fyhQemTk/openapi.json
