# URL to Markdown - Web Page to LLM-Ready Text for AI Agents (`tenfoldfleet/url-to-markdown`) Actor

Convert a web page to Markdown for an LLM: any URL to clean, LLM-ready text for AI agents, RAG pipelines and MCP tools. Strips navigation, ads and cookie banners; keeps headings, lists, tables, code and links, plus title and word count. HTTP-only. $1 per 1,000 pages; failed pages and PDFs are free.

- **URL**: https://apify.com/tenfoldfleet/url-to-markdown.md
- **Developed by:** [Tenfold Fleet](https://apify.com/tenfoldfleet) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 page converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does URL to Markdown do?

**URL to Markdown** turns any **web page into clean, LLM-ready Markdown**. Give it one URL or thousands and it returns the page's **main content as GitHub-flavoured Markdown** (headings, lists, tables, code blocks and links), plus the title, description, author, publish date, language, word count and every link on the page.

It's a single-purpose **HTML to Markdown API** built for **AI agents, RAG ingestion and LLM pipelines**: navigation menus, headers, footers, sidebars, cookie banners, ads, share buttons and scripts are stripped out with a Readability-style main-content extractor, so you don't pay for junk tokens. It is a fast, low-cost **Firecrawl scrape and Jina Reader alternative** that runs on plain HTTP requests (no browser) for **$1 per 1,000 pages**.

Use it as an **MCP tool for AI agents** through the [Apify MCP server](https://mcp.apify.com), from the **API**, or from the Apify Console.

### Why use a web page to Markdown converter?

- 🤖 **AI agents**: give Claude, ChatGPT, Cursor or your own agent a reliable "read this URL" tool that returns compact Markdown instead of raw HTML.
- 📚 **RAG ingestion**: convert documentation, help centers, blogs and Wikipedia articles into clean chunks for your vector database.
- 🧠 **LLM context**: paste whole articles into a prompt with `maxCharacters` to fit the context window.
- 🔎 **Research and monitoring**: archive articles as text, with author and publish date.
- 🧰 **Developers**: one HTTP call, one JSON item per URL, predictable fields.

Runs on the Apify platform, so you get an **API**, scheduling, webhooks and integrations with Make, Zapier, n8n, LangChain and LlamaIndex.

### What data can it extract?

| Field | Example |
|---|---|
| `markdown` | `# Overview of HTTP\n\n**HTTP** is a [protocol](https://...) for fetching resources...` |
| `title`, `description` | page title and meta description |
| `author`, `publishedAt` | from meta tags and JSON-LD (when the page has them) |
| `language` | `en-US` |
| `wordCount` | `2400` |
| `links` | all absolute `http(s)` URLs on the page (for agents that crawl on) |
| `url`, `finalUrl`, `statusCode` | requested URL, URL after redirects, HTTP status |
| `error` | why a page failed (`HTTP 404`, `PDF not supported`, ...), otherwise `null` |

Plain-text and Markdown files are returned as-is, JSON responses are pretty-printed in a fenced `json` code block.

### How to convert a URL to Markdown

1. Click **Try for free**.
2. Paste one or more URLs into **URLs** (one per line).
3. Keep **Main content only** on for articles and docs; turn it off to convert the whole page.
4. Click **Start**, then download results as **JSON, CSV, Excel or HTML**, or read them through the API.

#### Call it from an AI agent (Apify MCP server)

Add the Actor as a tool in any MCP client (Claude Desktop, Claude Code, Cursor, VS Code, ...):

```json
{
    "mcpServers": {
        "apify": { "url": "https://mcp.apify.com/?tools=tenfoldfleet/url-to-markdown" }
    }
}
```

The agent then calls the tool with `{"urls": ["https://example.com/article"]}` and gets the Markdown back in the same response.

#### Call it from the API (synchronous)

`run-sync-get-dataset-items` runs the Actor and returns the dataset items in one request:

```bash
curl -X POST "https://api.apify.com/v2/acts/tenfoldfleet~url-to-markdown/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://en.wikipedia.org/wiki/Markdown"], "maxCharacters": 20000}'
```

Python (`pip install "apify-client>=3"`):

```python
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("tenfoldfleet/url-to-markdown").call(run_input={"urls": ["https://react.dev/learn"]})
for page in client.dataset(run.default_dataset_id).iterate_items():
    print(page["title"], page["wordCount"])
    print(page["markdown"][:500])
```

On apify-client 1.x or 2.x, `call()` returns a dict, so use `run["defaultDatasetId"]` instead.

### How much does it cost?

You pay **$0.001 per page successfully converted** (**$1 per 1,000 pages**). Pages that fail (timeouts, 404s, blocked pages), PDFs and pages with no text are **not charged**. Apify's free plan includes monthly credit, so you can convert thousands of pages at no cost. Set a **maximum cost per run** and the Actor stops as soon as it is reached.

### Input

See the **Input** tab for all options. Example:

```json
{
    "urls": [
        "https://en.wikipedia.org/wiki/Markdown",
        "https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview",
        "https://github.com/mozilla/readability"
    ],
    "mainContentOnly": true,
    "includeLinks": true,
    "includeImages": false,
    "maxCharacters": 0,
    "maxConcurrency": 25
}
```

### Output

One item per URL. Example (Markdown shortened):

```json
{
    "url": "https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview",
    "finalUrl": "https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Overview",
    "statusCode": 200,
    "title": "Overview of HTTP - HTTP | MDN",
    "description": "HTTP is a protocol for fetching resources such as HTML documents.",
    "language": "en-US",
    "author": null,
    "publishedAt": null,
    "markdown": "# Overview of HTTP\n\n**HTTP** is a [protocol](https://developer.mozilla.org/en-US/docs/Glossary/Protocol) for fetching resources such as HTML documents...\n\n## Components of HTTP-based systems\n\n...",
    "wordCount": 2400,
    "links": ["https://developer.mozilla.org/en-US/docs/Glossary/Protocol", "..."],
    "error": null
}
```

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

### Tips

- **Save tokens**: turn off **Include links** for plain prose, and set **Max characters per page** to cap long pages.
- **Whole page**: turn off **Main content only** for landing pages, pricing pages or link directories where every block matters.
- **JavaScript-only sites**: this Actor reads the HTML the server sends. Single-page apps that render text only in the browser may return little content; use a browser-based crawler such as Website Content Crawler for those.
- **Blocked sites**: enable **Apify Proxy** if many URLs return 403 or 429.

### FAQ

#### Is this a Jina Reader or Firecrawl alternative?

Yes for the common case "give me this URL as clean Markdown". It does one thing: fetch the page over HTTP, extract the main content and convert it to Markdown, with simple per-page pricing and no subscription.

#### How is the main content detected?

It uses Mozilla's Readability algorithm (the engine behind Firefox Reader View) plus extra cleanup of navigation, cookie banners, ads and share widgets. Short pages such as product or docs index pages fall back to the cleaned `<main>` element so content isn't lost.

#### Are tables and code blocks preserved?

Yes. Tables become GitHub-flavoured Markdown tables (including tables without a header row) and code blocks become fenced blocks with the language when the page declares it.

#### Does it support PDFs?

Not yet. PDF URLs return an item with `error: "PDF not supported"` and are not charged.

#### Does it crawl whole websites?

No, it converts exactly the URLs you give it. Use the `links` field to pick the next pages, or use a crawler Actor for full-site crawls.

#### Support

Found a page that converts badly? Open an issue in the **Issues** tab with the URL.

This Actor only reads publicly available web pages. It does not log in, solve CAPTCHAs or collect personal data. You are responsible for respecting the terms of the websites you convert.

### More tools from Tenfold Fleet

| Actor | Price |
|---|---|
| [ATS Jobs Scraper - Greenhouse, Lever, Ashby & 5 More](https://apify.com/tenfoldfleet/ats-jobs-scraper) | $2 per 1,000 (job posting) |
| [Website Contact Scraper - Emails, Phones & Socials](https://apify.com/tenfoldfleet/company-contact-finder) | $8 per 1,000 (website with contacts) |
| [Website Technology Detector - Wappalyzer Alternative](https://apify.com/tenfoldfleet/tech-stack-detector) | $10 per 1,000 (website analyzed) |
| [YouTube Transcript Scraper - Captions & Subtitles API](https://apify.com/tenfoldfleet/youtube-transcript-scraper) | $3 per 1,000 (transcript extracted) |

# Actor input Schema

## `urls` (type: `array`):

Web pages to convert to Markdown. One URL = one result. HTML pages are converted; plain text and JSON are passed through; PDFs are skipped (free).

## `mainContentOnly` (type: `boolean`):

Keep only the article or main content (Readability-style extraction). Removes navigation, header, footer, sidebars, cookie banners, ads, scripts and share buttons. Turn off to convert the whole page.

## `includeLinks` (type: `boolean`):

Keep hyperlinks in the Markdown as [text](url) and return all absolute page URLs in the links field. Turn off for plain text with fewer tokens.

## `includeImages` (type: `boolean`):

Keep images as ![alt](url) in the Markdown. Off by default to save LLM tokens.

## `maxCharacters` (type: `integer`):

Truncate each page's Markdown to this many characters (useful to fit an LLM context window). 0 = no limit.

## `maxConcurrency` (type: `integer`):

How many pages to fetch in parallel.

## `proxyConfiguration` (type: `object`):

Optional. Most sites work without a proxy. Use Apify Proxy if many sites block you.

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown",
    "https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview",
    "https://github.com/mozilla/readability"
  ],
  "mainContentOnly": true,
  "includeLinks": true,
  "includeImages": false,
  "maxCharacters": 0,
  "maxConcurrency": 25,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Markdown",
        "https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview",
        "https://github.com/mozilla/readability"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tenfoldfleet/url-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://en.wikipedia.org/wiki/Markdown",
        "https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview",
        "https://github.com/mozilla/readability",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tenfoldfleet/url-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/Markdown",
    "https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview",
    "https://github.com/mozilla/readability"
  ]
}' |
apify call tenfoldfleet/url-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tenfoldfleet/url-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3hgdWJUAonHKebvWw/builds/cxhG3UJEpEMa9QHa2/openapi.json
