URL to Markdown for LLMs: Clean Page Content, No Browser avatar

URL to Markdown for LLMs: Clean Page Content, No Browser

Pricing

from $0.95 / 1,000 page converteds

Go to Apify Store
URL to Markdown for LLMs: Clean Page Content, No Browser

URL to Markdown for LLMs: Clean Page Content, No Browser

Turn any web page into clean Markdown for LLMs and RAG pipelines. Plain HTTP, no browser needed. Give it a list of URLs and get back readability-extracted, banner/nav/footer-stripped GFM Markdown, with title, description, language, word count, and links.

Pricing

from $0.95 / 1,000 page converteds

Rating

0.0

(0)

Developer

Adrian Voss

Adrian Voss

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Converts a web page to Markdown, for a list of URLs at once: give it the list and get back clean Markdown for LLMs and RAG pipelines, without a browser. A plain HTTP fetch plus Mozilla's own Readability algorithm pulls out the real article and drops the cookie banners, navigation, footers, and scripts around it — a lighter, no-browser website content crawler alternative for anyone who just needs one page's text, not a full site crawl.

How to use

  1. In the Apify Console. Open the actor page and click Start — the urls field is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found.
  2. Via the API. Call it directly with a POST request — no Console needed once you have an API token:
    curl "https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
    -X POST \
    -H "Content-Type: application/json" \
    -d '{"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}'
  3. On a schedule. Save this actor as an Apify Task with the input you want, then add a Schedule (hourly, daily, weekly) so it runs on its own — no server of your own required.

Input

{
"urls": [
"https://en.wikipedia.org/wiki/Markdown",
"https://developer.mozilla.org/en-US/docs/Web/HTML"
]
}

One per line. A full http(s) page URL, e.g. a news article, docs page, or blog post. Accepted formats: https://example.com/blog/some-article, https://docs.example.com/getting-started.

Output

One row per item, for example:

queryfoundstatusurlfinalUrlstatusCodetitledescriptionlangmarkdownwordCountpublishedAtlinksneedsBrowserscrapedAt
https://en.wikipedia.org/wiki/MarkdowntrueOK<url (as submitted)><final url (after redirects)><description / excerpt><published date (if given by the page)><links found in the content (if enabled)><needs a real browser (empty/consent-wall page)>1970-01-01T00:00:00.000Z

A miss comes back as a row with "found": false and is never charged.

When a page needs a real browser

This actor is deliberately plain HTTP: it never runs JavaScript. Two kinds of page can't be read that way:

  • An empty client-rendered shell — the real content is injected by JavaScript after load, so the raw HTML this actor fetches has nothing in it.
  • A consent-wall-only page — after stripping known cookie/consent-banner containers (OneTrust, Cookiebot, Quantcast Choice, Didomi, TrustArc, Osano, and generic "cookie banner" divs) plus nav/header/footer/scripts, there's nothing left to convert.

Either case comes back as a normal row with needsBrowser: true, markdown: "", and wordCount: 0 — not an error, and not charged. Tested against four fixtures during development: a news article (real content, correctly extracted), a docs page (real content, correctly extracted), a page with a cookie banner over real content (banner stripped, article content still extracted and charged normally), and an empty JS shell (correctly flagged needsBrowser: true, not charged). For a page that needs a browser, use Apify's own apify/web-fetch or apify/website-content-crawler instead.

Pricing

  • Page converted: $1.9 per 1,000 pages

Plus a $0.00005 start fee per run. Each event above is billed independently, only when it actually returns data — misses (found:false) are never charged.

Use it from Clay, n8n, Make, or an AI agent

This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.

curl "https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
-X POST \
-H "Content-Type: application/json" \
-d '{"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]}'

n8n. Add an HTTP Request node: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body Content Type JSON, JSON Body {"urls":["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTML"]} (swap in an expression from an earlier node for a real value).

Clay. Add an "HTTP API" column: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~url-to-markdown/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body {"urls":["{{page}}"]}, mapping the row's page into the urls array.

MCP. In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "URL to Markdown for LLMs | Apify" — the agent will find and run this actor.