# URL to Markdown for LLMs: Clean Page Content, No Browser (`accountable_eel/url-to-markdown`) Actor

Turn any web page into clean Markdown for LLMs and RAG pipelines. Plain HTTP, no browser needed. Give it a list of URLs and get back readability-extracted, banner/nav/footer-stripped GFM Markdown, with title, description, language, word count, and links.

- **URL**: https://apify.com/accountable\_eel/url-to-markdown.md
- **Developed by:** [Adrian Voss](https://apify.com/accountable_eel) (community)
- **Categories:** Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.95 / 1,000 page converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## URL to Markdown for LLMs: Clean Page Content, No Browser

Converts a web page to Markdown, for a list of URLs at once: give it the list and get back clean Markdown for LLMs and RAG pipelines, without a browser. A plain HTTP fetch plus Mozilla's own Readability algorithm pulls out the real article and drops the cookie banners, navigation, footers, and scripts around it — a lighter, no-browser website content crawler alternative for anyone who just needs one page's text, not a full site crawl.

### How to use

1. **In the Apify Console.** Open the actor page and click **Start** — the `urls` field is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found.
2. **Via the API.** Call it directly with a POST request — no Console needed once you have an API token:
   ```bash
   curl "https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
     -X POST \
     -H "Content-Type: application/json" \
     -d '{"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}'
   ```
3. **On a schedule.** Save this actor as an Apify **Task** with the input you want, then add a **Schedule** (hourly, daily, weekly) so it runs on its own — no server of your own required.

### Input

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown",
    "https://developer.mozilla.org/en-US/docs/Web/HTML"
  ]
}
```

One per line. A full http(s) page URL, e.g. a news article, docs page, or blog post. Accepted formats: https://example.com/blog/some-article, https://docs.example.com/getting-started.

### Output

One row per item, for example:

| query | found | status | url | finalUrl | statusCode | title | description | lang | markdown | wordCount | publishedAt | links | needsBrowser | scrapedAt |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| https://en.wikipedia.org/wiki/Markdown | true | OK | \<url (as submitted)> | \<final url (after redirects)> | <http status code> | <page title> | \<description / excerpt> | <language> | <clean markdown content> | <word count> | \<published date (if given by the page)> | \<links found in the content (if enabled)> | \<needs a real browser (empty/consent-wall page)> | 1970-01-01T00:00:00.000Z |

A miss comes back as a row with `"found": false` and is never charged.

### When a page needs a real browser

This actor is deliberately plain HTTP: it never runs JavaScript. Two kinds of page can't be read that way:

- **An empty client-rendered shell** — the real content is injected by JavaScript after load, so the raw HTML this actor fetches has nothing in it.
- **A consent-wall-only page** — after stripping known cookie/consent-banner containers (OneTrust, Cookiebot, Quantcast Choice, Didomi, TrustArc, Osano, and generic "cookie banner" divs) plus nav/header/footer/scripts, there's nothing left to convert.

Either case comes back as a normal row with `needsBrowser: true`, `markdown: ""`, and `wordCount: 0` — not an error, and **not charged**. Tested against four fixtures during development: a news article (real content, correctly extracted), a docs page (real content, correctly extracted), a page with a cookie banner over real content (banner stripped, article content still extracted and charged normally), and an empty JS shell (correctly flagged `needsBrowser: true`, not charged). For a page that needs a browser, use Apify's own `apify/web-fetch` or `apify/website-content-crawler` instead.

### Pricing

- **Page converted**: $1.9 per 1,000 pages

Plus a $0.00005 start fee per run. Each event above is billed independently, only when it actually returns data — misses (`found:false`) are never charged.

### Use it from Clay, n8n, Make, or an AI agent

This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.

```bash
curl "https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}'
```

**n8n.** Add an HTTP Request node: Method `POST`, URL `https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>`, Body Content Type `JSON`, JSON Body `{"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}` (swap in an expression from an earlier node for a real value).

**Clay.** Add an "HTTP API" column: Method `POST`, URL `https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>`, Body `{"urls":["{{page}}"]}`, mapping the row's page into the `urls` array.

**MCP.** In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "URL to Markdown for LLMs | Apify" — the agent will find and run this actor.

# Actor input Schema

## `urls` (type: `array`):

One per line. A full http(s) page URL, e.g. a news article, docs page, or blog post. Accepted formats: https://example.com/blog/some-article, https://docs.example.com/getting-started. You're only charged for the ones we actually find — a miss costs nothing.

## `testRun` (type: `boolean`):

Turn this on to test your input on a small sample before running the full list. Turn it off to process everything.

## `onlyFound` (type: `boolean`):

Only keep rows where something was actually found. Misses are always free, whether or not you show them here.

## `includeKeywords` (type: `array`):

Optional. Only keep results that mention at least one of these words (e.g. a job title, a city, a product name). Leave empty to keep everything.

## `excludeKeywords` (type: `array`):

Optional. Drop any result that mentions one of these words. Leave empty to skip nothing.

## `maxResults` (type: `integer`):

Optional. Stop the run once this many results have been found — useful for a quick, cheap sample. Leave blank for no limit.

## `columns` (type: `array`):

Choose which pieces of information to include in each result row. All are included by default.

## `includeLinks` (type: `boolean`):

Off by default (smaller rows, cheaper to store/transfer). On: adds a `links` array of `{url, text}` for every link inside the extracted content, with relative links resolved to absolute URLs.

## `maxConcurrency` (type: `integer`):

Parallel requests. Keep conservative — this target has no browser fallback, so getting blocked costs more than slow-and-steady.

## `proxyConfiguration` (type: `object`):

Apify Proxy config. Residential recommended for anti-bot-sensitive targets.

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown",
    "https://developer.mozilla.org/en-US/docs/Web/HTML"
  ],
  "testRun": false,
  "onlyFound": false,
  "includeKeywords": [],
  "excludeKeywords": [],
  "columns": [
    "url",
    "finalUrl",
    "statusCode",
    "title",
    "description",
    "lang",
    "markdown",
    "wordCount",
    "publishedAt",
    "links",
    "needsBrowser"
  ],
  "includeLinks": false,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Markdown",
        "https://developer.mozilla.org/en-US/docs/Web/HTML"
    ],
    "includeKeywords": [],
    "excludeKeywords": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("accountable_eel/url-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Markdown",
        "https://developer.mozilla.org/en-US/docs/Web/HTML",
    ],
    "includeKeywords": [],
    "excludeKeywords": [],
}

# Run the Actor and wait for it to finish
run = client.actor("accountable_eel/url-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown",
    "https://developer.mozilla.org/en-US/docs/Web/HTML"
  ],
  "includeKeywords": [],
  "excludeKeywords": []
}' |
apify call accountable_eel/url-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,accountable_eel/url-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CEqr0CUGMPjdBMu4X/builds/lFdHagk6Vyd5ppdWn/openapi.json
