Web Page to Markdown for AI: $1/1K Pages, No Login avatar

Web Page to Markdown for AI: $1/1K Pages, No Login

Pricing

from $1.00 / 1,000 page reads

Go to Apify Store
Web Page to Markdown for AI: $1/1K Pages, No Login

Web Page to Markdown for AI: $1/1K Pages, No Login

Turn any list of web pages into clean Markdown for LLMs, RAG and AI agents: main content without menus and ads, tables kept, plus title, description, language, dates and links. Plain HTTP first, JavaScript only when a page needs it. $1 per 1,000 pages, failures free.

Pricing

from $1.00 / 1,000 page reads

Rating

0.0

(0)

Developer

Don Mangu

Don Mangu

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 hours ago

Last modified

Share

For AI agents: pass a list of URLs; get one JSON row per page with clean Markdown (main content, tables kept), plain text, title, description, language, dates and optional links. Pages that cannot be read return a free row with a status and the reason.

Cost: $0.001 per page read with a plain request, $0.0025 per page that needed a browser for JavaScript. Pages that fail, are blocked, or are not allowed by robots.txt are free. No login or API key.

What does Web Page to Markdown do?

It turns web pages into clean Markdown for LLMs, RAG pipelines and AI agents. For each URL it keeps the main content (the article, documentation page or product text) and drops menus, headers, footers, sidebars, cookie notices and scripts. Headings, lists, links, images and tables stay as Markdown. Each row also has the page title, meta description, language, site name, canonical URL, preview image, and the published and modified dates when the page states them.

  • RAG and search indexes: feed documentation, help centers and blogs into a vector database as Markdown.
  • AI agents: give an agent the readable content of any page it finds, through the Apify API or Apify's MCP server.
  • Content and SEO work: read titles, descriptions, word counts and dates of many pages at once.

Try it now. The form opens with two pages. Click Start; it takes a few seconds and costs less than a cent.

How to convert web pages to Markdown

  1. Paste your pages into URLs, one per line (up to 10,000 per run).
  2. Pick the Output formats: Markdown, plain text, clean HTML, or several.
  3. Keep Content at Main content, or pick Whole page to keep everything in the body.
  4. Keep JavaScript rendering at Auto. Each page is read with a fast plain request first and opened in a browser only when it shows no content without JavaScript.
  5. Click Start, then open the Overview or Markdown table, or download the rows as JSON, CSV or Excel.

How much does it cost?

EventPrice
Page read with a plain HTTP request (page)$0.001 ($1 per 1,000 pages)
Page that needed a browser for JavaScript (page-rendered)$0.0025 ($2.50 per 1,000)
Page not read (blocked, not allowed by robots.txt, error, not found, PDF, empty)Free

Platform usage is included. Apify adds only its small standard fee per run start.

Example: you convert 20,000 documentation pages. 19,400 are read with plain requests, 300 need JavaScript and 300 are not found. The run costs 19,400 x $0.001 + 300 x $0.0025 = $20.15. The 300 missing pages cost nothing. You can set a spending limit on any run, and the Actor stops cleanly when it is reached.

Input

{
"urls": ["https://en.wikipedia.org/wiki/Markdown", "https://www.gov.uk/government/organisations/hm-revenue-customs"],
"outputFormats": ["markdown"],
"contentScope": "main",
"renderJavaScript": "auto",
"includeLinks": false
}

Advanced options: maximum characters per page (longer content is cut and marked truncated), pages at a time (1 to 20; at most 2 at a time go to the same site) and a timeout per page.

Output

{
"url": "https://en.wikipedia.org/wiki/Markdown",
"finalUrl": "https://en.wikipedia.org/wiki/Markdown",
"status": "ok",
"httpStatus": 200,
"fetchMode": "http",
"title": "Markdown - Wikipedia",
"language": "en",
"publishedAt": "2005-08-09T19:56:00.000Z",
"contentScope": "main",
"markdown": "From Wikipedia, the free encyclopedia\n\n| Markdown |\n| --- |\n...\n\n**Markdown** is a lightweight markup language ...",
"wordCount": 2781,
"approxTokens": 11764,
"truncated": false,
"charged": true,
"error": null
}

status is ok for pages that were read and charged. The free statuses are robots_disallowed, robots_unreachable, blocked, http_error, not_found, timeout, not_html (a PDF or other file), unsupported_content, empty and invalid_url, each with an error that says why. A STATS record in the key-value store counts pages by status, charges, requests and time.

FAQ

Does it respect robots.txt? Yes. It reads each site's robots.txt and skips pages the site does not allow for the token DonMangu-WebToMarkdown or for all crawlers. Those pages are free. If a site's robots.txt cannot be read because of a server error, its pages are skipped too, as the robots.txt standard asks.

Does it get past bot protection or CAPTCHAs? No. A page that answers with a bot check, a CAPTCHA or an access-denied page is returned as blocked, free, and the site is not asked again after three blocks in a run. The Actor reads public pages only; it does not log in.

Why does a page show little text? Some pages keep their content in a PDF, an image or a frame from another site. PDFs are returned as not_html; use a PDF text extractor for them. For very dynamic pages, set JavaScript rendering to Always.

Are author names included? No. The Actor returns page content and page-level details, not personal details about authors.

Is there a size limit? Pages over 5 MB are read up to 5 MB and marked truncated. Each output field is cut at the maximum characters you set (200,000 by default).