# Website Content Crawler — Clean Markdown for AI (`dallel-data/website-content-to-markdown`) Actor

Crawl public HTML websites into clean Markdown and text for AI, RAG and content research. Export titles, metadata, canonical URLs and links with depth and page limits.

- **URL**: https://apify.com/dallel-data/website-content-to-markdown.md
- **Developed by:** [GUIR Dallel](https://apify.com/dallel-data) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 delivered results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Content Crawler — Clean Markdown for AI

Crawl public HTML websites into clean Markdown and text for AI, RAG and content research. Export titles, metadata, canonical URLs and links with depth and page limits.

### Quick start

Paste your targets into the input form and run. The example below fetches a small sample:

```json
{
  "targets": [
    "https://books.toscrape.com"
  ],
  "maxResults": 10,
  "maxItemsPerTarget": 10,
  "maxPages": 3,
  "maxDepth": 1
}
```

The dataset supports JSON, CSV and Excel export. Use the API, schedules, Make, n8n or other Apify integrations to repeat the same collection. No external API key, browser cookies or login is required by this Actor.

### Input and limits

| Field | Meaning |
|---|---|
| `targets` | 1–50 source URLs. See README for supported URL formats. |
| `maxResults` | Stops the run after this many unique matching results. Always set a sensible cap. |
| `maxItemsPerTarget` | Per source limit after filters. Apple search uses this per query/country, capped at 200. |
| `maxPages` | Caps pagination or sitemap traversal. Ignored for sources returning a single document (Greenhouse, Ashby, RSS, Apple search). |
| `proxyConfiguration` | Optional Apify proxy settings. On Apify, Google Search uses Google SERP proxies; Google News, LinkedIn and Apple reviews use residential proxies. Other collectors connect directly unless configured here. Proxy usage is included in the Actor price when the listing uses pure pay-per-event pricing. |
| `maxDepth` | 0 collects only start URLs; 1 also follows their internal links. Maximum 3. |
| `includePath` | Optional substring required in discovered URL paths, such as /docs/. Start URLs are always processed. |

Duplicate targets are processed once. Filters apply before result billing. Runs are bounded snapshots; limits can stop collection before the source is exhausted.

### Output

Each result includes `source`, `id`, `title`, `url` and `fetchedAt` where applicable, plus source-specific fields. The Results view highlights `title`, `url`, `wordCount`, `language`, `depth`, `markdown`. JSON export retains every field, including nested arrays. Missing source values are null rather than estimated.

#### Example result

An actual test result, with long text shortened for readability:

```json
{
  "source": "website-content",
  "id": "https://books.toscrape.com",
  "url": "https://books.toscrape.com",
  "title": "All products | Books to Scrape - Sandbox",
  "description": "",
  "canonicalUrl": null,
  "language": "en-us",
  "markdown": "* [Home](index.html)\n* All products\n\n# All products\n\n**Warning!** This is a demo website for web scraping purposes. Prices and ratings here were randomly assigned and have no real meaning.\n\n1. ### [A Light in the ...](catalogue/a-light-in-the-attic_1000/index.html \"A Light in the Attic\")\n\n   £51.77\n\n   In stock\n2. ### [Tipping the Velvet](catalogue/tipping-the-velvet_999/index.html \"Tipping the Velvet\")\n\n   £53.74\n\n   In stock\n3. ### [Soumission]…",
  "text": "Home All products All products Warning! This is a demo website for web scraping purposes. Prices and ratings here were randomly assigned and have no real meaning. A Light in the ... £51.77 In stock Tipping the Velvet £53.74 In stock Soumission £50.10 In stock Sharp Objects £47.82 In stock Sapiens: A Brief History ... £54.23 In stock The Requiem Red £22.65 In stock The Dirty Little Secrets ... £33.34 In stock The Coming Woman: A ... £17.93 In stoc…",
  "wordCount": 167,
  "links": [
    "https://books.toscrape.com/index.html",
    "https://books.toscrape.com/catalogue/category/books_1/index.html",
    "https://books.toscrape.com/catalogue/category/books/travel_2/index.html",
    "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html",
    "https://books.toscrape.com/catalogue/category/books/historical-fiction_4/index.html",
    "https://books.toscrape.com/catalogue/category/books/sequential-art_5/index.html",
    "https://books.toscrape.com/catalogue/category/books/classics_6/index.html",
    "https://books.toscrape.com/catalogue/category/books/philosophy_7/index.html",
    "https://books.toscrape.com/catalogue/category/books/romance_8/index.html",
    "https://books.toscrape.com/catalogue/category/books/womens-fiction_9/index.html",
    "https://books.toscrape.com/catalogue/category/books/fiction_10/index.html",
    "https://books.toscrape.com/catalogue/category/books/childrens_11/index.html",
    "https://books.toscrape.com/catalogue/category/books/religion_12/index.html",
    "https://books.toscrape.com/catalogue/category/books/nonfiction_13/index.html",
    "https://books.toscrape.com/catalogue/category/books/music_14/index.html",
    "https://books.toscrape.com/catalogue/category/books/default_15/index.html",
    "https://books.toscrape.com/catalogue/category/books/science-fiction_16/index.html",
    "https://books.toscrape.com/catalogue/category/books/sports-and-games_17/index.html",
    "https://books.toscrape.com/catalogue/category/books/add-a-comment_18/index.html",
    "https://books.toscrape.com/catalogue/category/books/fantasy_19/index.html",
    "https://books.toscrape.com/catalogue/category/books/new-adult_20/index.html",
    "https://books.toscrape.com/catalogue/category/books/young-adult_21/index.html",
    "https://books.toscrape.com/catalogue/category/books/science_22/index.html",
    "https://books.toscrape.com/catalogue/category/books/poetry_23/index.html",
    "https://books.toscrape.com/catalogue/category/books/paranormal_24/index.html",
    "https://books.toscrape.com/catalogue/category/books/art_25/index.html",
    "https://books.toscrape.com/catalogue/category/books/psychology_26/index.html",
    "https://books.toscrape.com/catalogue/category/books/autobiography_27/index.html",
    "https://books.toscrape.com/catalogue/category/books/parenting_28/index.html",
    "https://books.toscrape.com/catalogue/category/books/adult-fiction_29/index.html",
    "https://books.toscrape.com/catalogue/category/books/humor_30/index.html",
    "https://books.toscrape.com/catalogue/category/books/horror_31/index.html",
    "https://books.toscrape.com/catalogue/category/books/history_32/index.html",
    "https://books.toscrape.com/catalogue/category/books/food-and-drink_33/index.html",
    "https://books.toscrape.com/catalogue/category/books/christian-fiction_34/index.html",
    "https://books.toscrape.com/catalogue/category/books/business_35/index.html",
    "https://books.toscrape.com/catalogue/category/books/biography_36/index.html",
    "https://books.toscrape.com/catalogue/category/books/thriller_37/index.html",
    "https://books.toscrape.com/catalogue/category/books/contemporary_38/index.html",
    "https://books.toscrape.com/catalogue/category/books/spirituality_39/index.html",
    "https://books.toscrape.com/catalogue/category/books/academic_40/index.html",
    "https://books.toscrape.com/catalogue/category/books/self-help_41/index.html",
    "https://books.toscrape.com/catalogue/category/books/historical_42/index.html",
    "https://books.toscrape.com/catalogue/category/books/christian_43/index.html",
    "https://books.toscrape.com/catalogue/category/books/suspense_44/index.html",
    "https://books.toscrape.com/catalogue/category/books/short-stories_45/index.html",
    "https://books.toscrape.com/catalogue/category/books/novels_46/index.html",
    "https://books.toscrape.com/catalogue/category/books/health_47/index.html",
    "https://books.toscrape.com/catalogue/category/books/politics_48/index.html",
    "https://books.toscrape.com/catalogue/category/books/cultural_49/index.html",
    "https://books.toscrape.com/catalogue/category/books/erotica_50/index.html",
    "https://books.toscrape.com/catalogue/category/books/crime_51/index.html",
    "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
    "https://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html",
    "https://books.toscrape.com/catalogue/soumission_998/index.html",
    "https://books.toscrape.com/catalogue/sharp-objects_997/index.html",
    "https://books.toscrape.com/catalogue/sapiens-a-brief-history-of-humankind_996/index.html",
    "https://books.toscrape.com/catalogue/the-requiem-red_995/index.html",
    "https://books.toscrape.com/catalogue/the-dirty-little-secrets-of-getting-your-dream-job_994/index.html",
    "https://books.toscrape.com/catalogue/the-coming-woman-a-novel-based-on-the-life-of-the-infamous-feminist-victoria-woodhull_993/index.html",
    "https://books.toscrape.com/catalogue/the-boys-in-the-boat-nine-americans-and-their-epic-quest-for-gold-at-the-1936-berlin-olympics_992/index.html",
    "https://books.toscrape.com/catalogue/the-black-maria_991/index.html",
    "https://books.toscrape.com/catalogue/starving-hearts-triangular-trade-trilogy-1_990/index.html",
    "https://books.toscrape.com/catalogue/shakespeares-sonnets_989/index.html",
    "https://books.toscrape.com/catalogue/set-me-free_988/index.html",
    "https://books.toscrape.com/catalogue/scott-pilgrims-precious-little-life-scott-pilgrim-1_987/index.html",
    "https://books.toscrape.com/catalogue/rip-it-up-and-start-again_986/index.html",
    "https://books.toscrape.com/catalogue/our-band-could-be-your-life-scenes-from-the-american-indie-underground-1981-1991_985/index.html",
    "https://books.toscrape.com/catalogue/olio_984/index.html",
    "https://books.toscrape.com/catalogue/mesaerion-the-best-science-fiction-stories-1800-1849_983/index.html",
    "https://books.toscrape.com/catalogue/libertarianism-for-beginners_982/index.html",
    "https://books.toscrape.com/catalogue/its-only-the-himalayas_981/index.html",
    "https://books.toscrape.com/catalogue/page-2.html"
  ],
  "depth": 0,
  "fetchedAt": "2026-09-19T10:08:43.205661+00:00"
}
```

The free `SUMMARY` record in the default key-value store reports each target's result count and errors, HTTP requests, and whether the overall result or spending limit was reached. Errors are not inserted into the paid dataset. A run in which every source fails is marked failed; mixed success is reported explicitly in the summary.

### Pricing

$2 per 1,000 delivered results ($0.002 each). No Actor start fee. Filtered rows, duplicates within the same context and source errors are not charged. Results are deduplicated within each run. Google Search preserves a URL appearing in distinct queries. Source requests and the default proxy costs are included in paid per-result pricing. Set Apify's maximum charge per run as an additional budget control.

### Source coverage and limitations

HTTP-only crawler: does not render JavaScript, log in, solve CAPTCHAs or read PDFs. Obeys robots.txt for the DallelData user agent. Follows same-host links without query strings to avoid crawl traps; supply a final canonical start URL. Extraction removes navigation, headers, footers and forms, then selects main/article/body; complex layouts may still include boilerplate. maxDepth controls internal-link traversal (0–3); maxPages caps fetched page attempts. Does not store or compare historical page versions.

Public sources can change, throttle requests or return no matches. The Actor retries transient network failures up to twice, follows a bounded number of redirects, and rejects private-network targets. It does not bypass login or CAPTCHA challenges. Use data within the source's applicable terms and your intended permissions.

### Support

Open an issue on this Actor's Issues tab with the failing public target, input, and run link. Never include passwords, tokens or personal account cookies.

# Actor input Schema

## `targets` (type: `array`):

1–50 source URLs. See README for supported URL formats.

## `maxResults` (type: `integer`):

Stops the run after this many unique matching results. Always set a sensible cap.

## `maxItemsPerTarget` (type: `integer`):

Per source limit after filters. Apple search uses this per query/country, capped at 200.

## `maxPages` (type: `integer`):

Caps pagination or sitemap traversal. Ignored for sources returning a single document (Greenhouse, Ashby, RSS, Apple search).

## `proxyConfiguration` (type: `object`):

Optional Apify proxy settings. On Apify, Google Search uses Google SERP proxies; Google News, LinkedIn and Apple reviews use residential proxies. Other collectors connect directly unless configured here. Proxy usage is included in the Actor price when the listing uses pure pay-per-event pricing.

## `maxDepth` (type: `integer`):

0 collects only start URLs; 1 also follows their internal links. Maximum 3.

## `includePath` (type: `string`):

Optional substring required in discovered URL paths, such as /docs/. Start URLs are always processed.

## Actor input object example

```json
{
  "targets": [
    "https://books.toscrape.com"
  ],
  "maxResults": 100,
  "maxItemsPerTarget": 100,
  "maxPages": 10,
  "maxDepth": 1
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "targets": [
        "https://books.toscrape.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("dallel-data/website-content-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "targets": ["https://books.toscrape.com"] }

# Run the Actor and wait for it to finish
run = client.actor("dallel-data/website-content-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "targets": [
    "https://books.toscrape.com"
  ]
}' |
apify call dallel-data/website-content-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dallel-data/website-content-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bkY5z4e0YobiFePSo/builds/3cmn6SvaABeb0rXQf/openapi.json
