# Web Page to Markdown for AI: $1/1K Pages, No Login (`conserving_celerytop/web-page-to-markdown`) Actor

Turn any list of web pages into clean Markdown for LLMs, RAG and AI agents: main content without menus and ads, tables kept, plus title, description, language, dates and links. Plain HTTP first, JavaScript only when a page needs it. $1 per 1,000 pages, failures free.

- **URL**: https://apify.com/conserving\_celerytop/web-page-to-markdown.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 page reads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

For AI agents: pass a list of URLs; get one JSON row per page with clean Markdown (main content, tables kept), plain text, title, description, language, dates and optional links. Pages that cannot be read return a free row with a status and the reason.

Cost: $0.001 per page read with a plain request, $0.0025 per page that needed a browser for JavaScript. Pages that fail, are blocked, or are not allowed by robots.txt are free. No login or API key.

### What does Web Page to Markdown do?

It turns web pages into clean Markdown for LLMs, RAG pipelines and AI agents. For each URL it keeps the main content (the article, documentation page or product text) and drops menus, headers, footers, sidebars, cookie notices and scripts. Headings, lists, links, images and tables stay as Markdown. Each row also has the page title, meta description, language, site name, canonical URL, preview image, and the published and modified dates when the page states them.

- **RAG and search indexes:** feed documentation, help centers and blogs into a vector database as Markdown.
- **AI agents:** give an agent the readable content of any page it finds, through the Apify API or Apify's MCP server.
- **Content and SEO work:** read titles, descriptions, word counts and dates of many pages at once.

**Try it now.** The form opens with two pages. Click **Start**; it takes a few seconds and costs less than a cent.

### How to convert web pages to Markdown

1. Paste your pages into **URLs**, one per line (up to 10,000 per run).
2. Pick the **Output formats**: Markdown, plain text, clean HTML, or several.
3. Keep **Content** at *Main content*, or pick *Whole page* to keep everything in the body.
4. Keep **JavaScript rendering** at *Auto*. Each page is read with a fast plain request first and opened in a browser only when it shows no content without JavaScript.
5. Click **Start**, then open the **Overview** or **Markdown** table, or download the rows as JSON, CSV or Excel.

### How much does it cost?

| Event | Price |
|---|---|
| Page read with a plain HTTP request (`page`) | $0.001 ($1 per 1,000 pages) |
| Page that needed a browser for JavaScript (`page-rendered`) | $0.0025 ($2.50 per 1,000) |
| Page not read (blocked, not allowed by robots.txt, error, not found, PDF, empty) | Free |

Platform usage is included. Apify adds only its small standard fee per run start.

Example: you convert 20,000 documentation pages. 19,400 are read with plain requests, 300 need JavaScript and 300 are not found. The run costs 19,400 x $0.001 + 300 x $0.0025 = **$20.15**. The 300 missing pages cost nothing. You can set a spending limit on any run, and the Actor stops cleanly when it is reached.

### Input

```json
{
  "urls": ["https://en.wikipedia.org/wiki/Markdown", "https://www.gov.uk/government/organisations/hm-revenue-customs"],
  "outputFormats": ["markdown"],
  "contentScope": "main",
  "renderJavaScript": "auto",
  "includeLinks": false
}
```

Advanced options: maximum characters per page (longer content is cut and marked `truncated`), pages at a time (1 to 20; at most 2 at a time go to the same site) and a timeout per page.

### Output

```json
{
  "url": "https://en.wikipedia.org/wiki/Markdown",
  "finalUrl": "https://en.wikipedia.org/wiki/Markdown",
  "status": "ok",
  "httpStatus": 200,
  "fetchMode": "http",
  "title": "Markdown - Wikipedia",
  "language": "en",
  "publishedAt": "2005-08-09T19:56:00.000Z",
  "contentScope": "main",
  "markdown": "From Wikipedia, the free encyclopedia\n\n| Markdown |\n| --- |\n...\n\n**Markdown** is a lightweight markup language ...",
  "wordCount": 2781,
  "approxTokens": 11764,
  "truncated": false,
  "charged": true,
  "error": null
}
```

`status` is `ok` for pages that were read and charged. The free statuses are `robots_disallowed`, `robots_unreachable`, `blocked`, `http_error`, `not_found`, `timeout`, `not_html` (a PDF or other file), `unsupported_content`, `empty` and `invalid_url`, each with an `error` that says why. A STATS record in the key-value store counts pages by status, charges, requests and time.

### Related

- [Website Screenshot API](https://apify.com/conserving_celerytop/website-screenshot-api): a full-page PNG or PDF of the same pages.
- [Article Extractor](https://apify.com/conserving_celerytop/article-extractor): article text, dates and site details from news and blog pages, and from RSS feeds.

### FAQ

**Does it respect robots.txt?** Yes. It reads each site's robots.txt and skips pages the site does not allow for the token `DonMangu-WebToMarkdown` or for all crawlers. Those pages are free. If a site's robots.txt cannot be read because of a server error, its pages are skipped too, as the robots.txt standard asks.

**Does it get past bot protection or CAPTCHAs?** No. A page that answers with a bot check, a CAPTCHA or an access-denied page is returned as `blocked`, free, and the site is not asked again after three blocks in a run. The Actor reads public pages only; it does not log in.

**Why does a page show little text?** Some pages keep their content in a PDF, an image or a frame from another site. PDFs are returned as `not_html`; use a PDF text extractor for them. For very dynamic pages, set JavaScript rendering to *Always*.

**Are author names included?** No. The Actor returns page content and page-level details, not personal details about authors.

**Is there a size limit?** Pages over 5 MB are read up to 5 MB and marked `truncated`. Each output field is cut at the maximum characters you set (200,000 by default).

# Actor input Schema

## `urls` (type: `array`):

Enter the web pages to read, one per line, as full addresses (https://example.com/page) or domains. Up to 10,000 per run. Each page gives one row.

## `outputFormats` (type: `array`):

Pick the formats to return for each page. Markdown keeps headings, lists, links, images and tables. Text is plain text with line breaks. HTML is the cleaned content HTML.

## `contentScope` (type: `string`):

Main content keeps the article or main part of the page and drops menus, headers, footers, sidebars and cookie notices. Whole page keeps everything in the body except scripts and forms.

## `renderJavaScript` (type: `string`):

Auto reads each page with a plain HTTP request and opens it in a browser only when it shows no content without JavaScript. Never uses plain HTTP only. Always opens every page in a browser. Browser pages cost more.

## `includeLinks` (type: `boolean`):

Add a list of the links on each page, with their text, up to 1,000 per page.

## `includeImages` (type: `boolean`):

Keep images in the Markdown as ![alt](address). Turn it off for text-only output.

## `maxCharacters` (type: `integer`):

Cut each output field at this many characters. The row then has truncated set to true.

## `maxConcurrency` (type: `integer`):

How many pages to read in parallel, 1 to 20. At most 2 at a time go to the same site.

## `timeoutSecs` (type: `integer`):

How long to wait for one page, 5 to 120 seconds.

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown"
  ],
  "outputFormats": [
    "markdown"
  ],
  "contentScope": "main",
  "renderJavaScript": "auto",
  "includeLinks": false,
  "includeImages": true,
  "maxCharacters": 200000,
  "maxConcurrency": 5,
  "timeoutSecs": 30
}
```

# Actor output Schema

## `overview` (type: `string`):

Status, title, words and tokens per page.

## `markdown` (type: `string`):

The Markdown of each page.

## `all` (type: `string`):

Every row with every field.

## `stats` (type: `string`):

Pages read, charges, requests and time.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Markdown",
        "https://www.gov.uk/government/organisations/hm-revenue-customs"
    ],
    "outputFormats": [
        "markdown"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/web-page-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Markdown",
        "https://www.gov.uk/government/organisations/hm-revenue-customs",
    ],
    "outputFormats": ["markdown"],
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/web-page-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown",
    "https://www.gov.uk/government/organisations/hm-revenue-customs"
  ],
  "outputFormats": [
    "markdown"
  ]
}' |
apify call conserving_celerytop/web-page-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/web-page-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ngFLLKWp0most13tg/builds/9nX3dqDngFGr2zkQy/openapi.json
