URL to Markdown for LLMs: Clean Page Content, No Browser
Pricing
from $0.95 / 1,000 page converteds
URL to Markdown for LLMs: Clean Page Content, No Browser
Turn any web page into clean Markdown for LLMs and RAG pipelines. Plain HTTP, no browser needed. Give it a list of URLs and get back readability-extracted, banner/nav/footer-stripped GFM Markdown, with title, description, language, word count, and links.
Pricing
from $0.95 / 1,000 page converteds
Rating
0.0
(0)
Developer
Adrian Voss
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Converts a web page to Markdown, for a list of URLs at once: give it the list and get back clean Markdown for LLMs and RAG pipelines, without a browser. A plain HTTP fetch plus Mozilla's own Readability algorithm pulls out the real article and drops the cookie banners, navigation, footers, and scripts around it — a lighter, no-browser website content crawler alternative for anyone who just needs one page's text, not a full site crawl.
How to use
- In the Apify Console. Open the actor page and click Start — the
urlsfield is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found. - Via the API. Call it directly with a POST request — no Console needed once you have an API token:
curl "https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \-X POST \-H "Content-Type: application/json" \-d '{"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}'
- On a schedule. Save this actor as an Apify Task with the input you want, then add a Schedule (hourly, daily, weekly) so it runs on its own — no server of your own required.
Input
{"urls": ["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}
One per line. A full http(s) page URL, e.g. a news article, docs page, or blog post. Accepted formats: https://example.com/blog/some-article, https://docs.example.com/getting-started.
Output
One row per item, for example:
| query | found | status | url | finalUrl | statusCode | title | description | lang | markdown | wordCount | publishedAt | links | needsBrowser | scrapedAt |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| https://en.wikipedia.org/wiki/Markdown | true | OK | <url (as submitted)> | <final url (after redirects)> | <description / excerpt> | <published date (if given by the page)> | <links found in the content (if enabled)> | <needs a real browser (empty/consent-wall page)> | 1970-01-01T00:00:00.000Z |
A miss comes back as a row with "found": false and is never charged.
When a page needs a real browser
This actor is deliberately plain HTTP: it never runs JavaScript. Two kinds of page can't be read that way:
- An empty client-rendered shell — the real content is injected by JavaScript after load, so the raw HTML this actor fetches has nothing in it.
- A consent-wall-only page — after stripping known cookie/consent-banner containers (OneTrust, Cookiebot, Quantcast Choice, Didomi, TrustArc, Osano, and generic "cookie banner" divs) plus nav/header/footer/scripts, there's nothing left to convert.
Either case comes back as a normal row with needsBrowser: true, markdown: "", and wordCount: 0 — not an error, and not charged. Tested against four fixtures during development: a news article (real content, correctly extracted), a docs page (real content, correctly extracted), a page with a cookie banner over real content (banner stripped, article content still extracted and charged normally), and an empty JS shell (correctly flagged needsBrowser: true, not charged). For a page that needs a browser, use Apify's own apify/web-fetch or apify/website-content-crawler instead.
Pricing
- Page converted: $1.9 per 1,000 pages
Plus a $0.00005 start fee per run. Each event above is billed independently, only when it actually returns data — misses (found:false) are never charged.
Use it from Clay, n8n, Make, or an AI agent
This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.
curl "https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \-X POST \-H "Content-Type: application/json" \-d '{"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}'
n8n. Add an HTTP Request node: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body Content Type JSON, JSON Body {"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]} (swap in an expression from an earlier node for a real value).
Clay. Add an "HTTP API" column: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body {"urls":["{{page}}"]}, mapping the row's page into the urls array.
MCP. In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "URL to Markdown for LLMs | Apify" — the agent will find and run this actor.