Web Page to Markdown for LLMs - Fast & Cheap avatar

Web Page to Markdown for LLMs - Fast & Cheap

Pricing

from $0.50 / 1,000 page converteds

Go to Apify Store
Web Page to Markdown for LLMs - Fast & Cheap

Web Page to Markdown for LLMs - Fast & Cheap

Turns any URL into clean Markdown and plain text: main content only (no menus, footers or cookie banners), absolute links, tables and code blocks, plus title, author, date and language. No browser, sub-second, $0.0005 per page. MCP-ready for AI agents.

Pricing

from $0.50 / 1,000 page converteds

Rating

0.0

(0)

Developer

Antonio del Olmo Tena

Antonio del Olmo Tena

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Give it URLs and get back clean Markdown you can feed straight into an LLM, a RAG index or an agent: only the main content of each page (no menus, headers, footers, sidebars, cookie banners or share buttons), with headings, lists, tables, code blocks, absolute links and images preserved. Each page also comes with plain text, word count, title, description, author, publish and modified dates, language and canonical URL.

No browser and no proxies: a plain HTTP request and a few milliseconds of processing per page, which is why it costs $0.0005 per page and answers in well under a second. Built for AI agents and automation (it is available through Apify's MCP integration) as much as for people who just need readable text.

What you get

One Dataset item per URL with:

  • markdown: the main content as Markdown (ATX headings, - lists, fenced code blocks, GFM tables, absolute links and image URLs).
  • text: the same content as plain text, and wordCount.
  • title, description, siteName, author, publishedAt, modifiedAt, language, canonical.
  • links and images found in the content (absolute, deduplicated).
  • mainContentUsed: false when no main block was detected and the whole page was converted.
  • likelyJsRendered: true when the page seems to render its content only in the browser, so the Markdown may be empty.
  • truncated: true when the Markdown was cut at Max Markdown characters.

Pages that cannot be fetched are returned with an error field and are not charged. Pages with no readable text are not charged either.

How to use it

  1. Open the Input tab.
  2. Paste the page URLs in Page URLs, one per line.
  3. Leave Main content only on for articles, docs and posts; turn it off to convert the whole page.
  4. Click Start. Read the Markdown in Output or export as JSON.
  5. From code or an agent, call the API below; schedule the Actor to monitor pages and get CONTENT_CHANGED and TITLE_CHANGED events when they change.

Input

{
"urls": ["https://docs.apify.com/platform/actors", "https://en.wikipedia.org/wiki/Markdown"],
"mainContent": true,
"maxChars": 100000
}
FieldDescription
urlsPublic http(s) URLs. Up to 50 per run. Private or internal addresses are rejected.
mainContentDefault true. Keeps the article/main block and removes site chrome.
maxCharsDefault 100,000, maximum 500,000. Longer Markdown is cut at a line break.

Output example

{
"url": "https://blog.example/posts/hello",
"title": "Hello World | Example Blog",
"description": "A post about things",
"author": "Jane Doe",
"publishedAt": "2026-10-01T10:00:00Z",
"language": "en",
"markdown": "# Hello World\n\nFirst paragraph with a [link](https://blog.example/related) and **bold** text.\n\n- Item one\n- Item two\n\n| Col A | Col B |\n| --- | --- |\n| 1 | 2 |",
"text": "Hello World First paragraph with a link and bold text. Item one Item two Col A Col B 1 2",
"wordCount": 18,
"truncated": false,
"mainContentUsed": true,
"likelyJsRendered": false,
"links": ["https://blog.example/related"],
"images": [],
"timestamp": "2026-10-04T10:00:00.000Z"
}

Pricing

You pay per page, nothing else:

  • Page converted ($0.0005): one page with readable content turned into Markdown. 1,000 pages cost $0.50.
  • Content changed ($0.002): only in scheduled runs, when a monitored page's content changed since the previous run.
  • Actor start ($0.00005): Apify's standard start fee.

Failed pages and pages with no readable text are free. Comparable Actors charge $0.002 to $0.003 per page.

Use it as an API

curl -X POST "https://api.apify.com/v2/acts/USERNAME~page-to-markdown/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://docs.apify.com/platform/actors"]}'

Replace USERNAME~page-to-markdown with the Actor id shown on this page. The call returns the items directly. Add &format=json&clean=true to get a compact response for prompts.

FAQ

The Markdown is empty or very short. Check likelyJsRendered: some sites render their content only in the browser. This Actor reads the HTML the server sends, which keeps it fast and cheap; for those sites use a browser-based crawler.

It kept something that is not content, or dropped something that is. Set Main content only to false to get the whole page, then filter on your side. Detection relies on article/main blocks and text density, which works for most articles, documentation and blog posts.

Does it follow links or crawl a whole site? No, it converts exactly the URLs you give it. Combine it with a sitemap or a list of links when you need a whole section.

Is anything bypassed? No. Plain requests to public pages, never to private or internal addresses, and no evasion of anti-bot systems.