Web Page to Markdown for LLMs - Fast & Cheap
Pricing
from $0.50 / 1,000 page converteds
Web Page to Markdown for LLMs - Fast & Cheap
Turns any URL into clean Markdown and plain text: main content only (no menus, footers or cookie banners), absolute links, tables and code blocks, plus title, author, date and language. No browser, sub-second, $0.0005 per page. MCP-ready for AI agents.
Pricing
from $0.50 / 1,000 page converteds
Rating
0.0
(0)
Developer
Antonio del Olmo Tena
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Give it URLs and get back clean Markdown you can feed straight into an LLM, a RAG index or an agent: only the main content of each page (no menus, headers, footers, sidebars, cookie banners or share buttons), with headings, lists, tables, code blocks, absolute links and images preserved. Each page also comes with plain text, word count, title, description, author, publish and modified dates, language and canonical URL.
No browser and no proxies: a plain HTTP request and a few milliseconds of processing per page, which is why it costs $0.0005 per page and answers in well under a second. Built for AI agents and automation (it is available through Apify's MCP integration) as much as for people who just need readable text.
What you get
One Dataset item per URL with:
markdown: the main content as Markdown (ATX headings,-lists, fenced code blocks, GFM tables, absolute links and image URLs).text: the same content as plain text, andwordCount.title,description,siteName,author,publishedAt,modifiedAt,language,canonical.linksandimagesfound in the content (absolute, deduplicated).mainContentUsed:falsewhen no main block was detected and the whole page was converted.likelyJsRendered:truewhen the page seems to render its content only in the browser, so the Markdown may be empty.truncated:truewhen the Markdown was cut at Max Markdown characters.
Pages that cannot be fetched are returned with an error field and are not charged. Pages with no readable text are not charged either.
How to use it
- Open the Input tab.
- Paste the page URLs in Page URLs, one per line.
- Leave Main content only on for articles, docs and posts; turn it off to convert the whole page.
- Click Start. Read the Markdown in Output or export as JSON.
- From code or an agent, call the API below; schedule the Actor to monitor pages and get
CONTENT_CHANGEDandTITLE_CHANGEDevents when they change.
Input
{"urls": ["https://docs.apify.com/platform/actors", "https://en.wikipedia.org/wiki/Markdown"],"mainContent": true,"maxChars": 100000}
| Field | Description |
|---|---|
urls | Public http(s) URLs. Up to 50 per run. Private or internal addresses are rejected. |
mainContent | Default true. Keeps the article/main block and removes site chrome. |
maxChars | Default 100,000, maximum 500,000. Longer Markdown is cut at a line break. |
Output example
{"url": "https://blog.example/posts/hello","title": "Hello World | Example Blog","description": "A post about things","author": "Jane Doe","publishedAt": "2026-10-01T10:00:00Z","language": "en","markdown": "# Hello World\n\nFirst paragraph with a [link](https://blog.example/related) and **bold** text.\n\n- Item one\n- Item two\n\n| Col A | Col B |\n| --- | --- |\n| 1 | 2 |","text": "Hello World First paragraph with a link and bold text. Item one Item two Col A Col B 1 2","wordCount": 18,"truncated": false,"mainContentUsed": true,"likelyJsRendered": false,"links": ["https://blog.example/related"],"images": [],"timestamp": "2026-10-04T10:00:00.000Z"}
Pricing
You pay per page, nothing else:
- Page converted ($0.0005): one page with readable content turned into Markdown. 1,000 pages cost $0.50.
- Content changed ($0.002): only in scheduled runs, when a monitored page's content changed since the previous run.
- Actor start ($0.00005): Apify's standard start fee.
Failed pages and pages with no readable text are free. Comparable Actors charge $0.002 to $0.003 per page.
Use it as an API
curl -X POST "https://api.apify.com/v2/acts/USERNAME~page-to-markdown/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"urls":["https://docs.apify.com/platform/actors"]}'
Replace USERNAME~page-to-markdown with the Actor id shown on this page. The call returns the items directly. Add &format=json&clean=true to get a compact response for prompts.
FAQ
The Markdown is empty or very short. Check likelyJsRendered: some sites render their content only in the browser. This Actor reads the HTML the server sends, which keeps it fast and cheap; for those sites use a browser-based crawler.
It kept something that is not content, or dropped something that is. Set Main content only to false to get the whole page, then filter on your side. Detection relies on article/main blocks and text density, which works for most articles, documentation and blog posts.
Does it follow links or crawl a whole site? No, it converts exactly the URLs you give it. Combine it with a sitemap or a list of links when you need a whole section.
Is anything bypassed? No. Plain requests to public pages, never to private or internal addresses, and no evasion of anti-bot systems.