URL to Markdown - Web Page to LLM-Ready Text for AI Agents
Pricing
from $1.00 / 1,000 page converteds
URL to Markdown - Web Page to LLM-Ready Text for AI Agents
Convert a web page to Markdown for an LLM: any URL to clean, LLM-ready text for AI agents, RAG pipelines and MCP tools. Strips navigation, ads and cookie banners; keeps headings, lists, tables, code and links, plus title and word count. HTTP-only. $1 per 1,000 pages; failed pages and PDFs are free.
Pricing
from $1.00 / 1,000 page converteds
Rating
0.0
(0)
Developer
Tenfold Fleet
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
What does URL to Markdown do?
URL to Markdown turns any web page into clean, LLM-ready Markdown. Give it one URL or thousands and it returns the page's main content as GitHub-flavoured Markdown (headings, lists, tables, code blocks and links), plus the title, description, author, publish date, language, word count and every link on the page.
It's a single-purpose HTML to Markdown API built for AI agents, RAG ingestion and LLM pipelines: navigation menus, headers, footers, sidebars, cookie banners, ads, share buttons and scripts are stripped out with a Readability-style main-content extractor, so you don't pay for junk tokens. It is a fast, low-cost Firecrawl scrape and Jina Reader alternative that runs on plain HTTP requests (no browser) for $1 per 1,000 pages.
Use it as an MCP tool for AI agents through the Apify MCP server, from the API, or from the Apify Console.
Why use a web page to Markdown converter?
- 🤖 AI agents: give Claude, ChatGPT, Cursor or your own agent a reliable "read this URL" tool that returns compact Markdown instead of raw HTML.
- 📚 RAG ingestion: convert documentation, help centers, blogs and Wikipedia articles into clean chunks for your vector database.
- 🧠 LLM context: paste whole articles into a prompt with
maxCharactersto fit the context window. - 🔎 Research and monitoring: archive articles as text, with author and publish date.
- 🧰 Developers: one HTTP call, one JSON item per URL, predictable fields.
Runs on the Apify platform, so you get an API, scheduling, webhooks and integrations with Make, Zapier, n8n, LangChain and LlamaIndex.
What data can it extract?
| Field | Example |
|---|---|
markdown | # Overview of HTTP\n\n**HTTP** is a [protocol](https://...) for fetching resources... |
title, description | page title and meta description |
author, publishedAt | from meta tags and JSON-LD (when the page has them) |
language | en-US |
wordCount | 2400 |
links | all absolute http(s) URLs on the page (for agents that crawl on) |
url, finalUrl, statusCode | requested URL, URL after redirects, HTTP status |
error | why a page failed (HTTP 404, PDF not supported, ...), otherwise null |
Plain-text and Markdown files are returned as-is, JSON responses are pretty-printed in a fenced json code block.
How to convert a URL to Markdown
- Click Try for free.
- Paste one or more URLs into URLs (one per line).
- Keep Main content only on for articles and docs; turn it off to convert the whole page.
- Click Start, then download results as JSON, CSV, Excel or HTML, or read them through the API.
Call it from an AI agent (Apify MCP server)
Add the Actor as a tool in any MCP client (Claude Desktop, Claude Code, Cursor, VS Code, ...):
{"mcpServers": {"apify": { "url": "https://mcp.apify.com/?tools=tenfoldfleet/url-to-markdown" }}}
The agent then calls the tool with {"urls": ["https://example.com/article"]} and gets the Markdown back in the same response.
Call it from the API (synchronous)
run-sync-get-dataset-items runs the Actor and returns the dataset items in one request:
curl -X POST "https://api.apify.com/v2/acts/tenfoldfleet~url-to-markdown/run-sync-get-dataset-items" \-H "Authorization: Bearer $APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"urls": ["https://en.wikipedia.org/wiki/Markdown"], "maxCharacters": 20000}'
Python (pip install "apify-client>=3"):
from apify_client import ApifyClientclient = ApifyClient("<APIFY_TOKEN>")run = client.actor("tenfoldfleet/url-to-markdown").call(run_input={"urls": ["https://react.dev/learn"]})for page in client.dataset(run.default_dataset_id).iterate_items():print(page["title"], page["wordCount"])print(page["markdown"][:500])
On apify-client 1.x or 2.x, call() returns a dict, so use run["defaultDatasetId"] instead.
How much does it cost?
You pay $0.001 per page successfully converted ($1 per 1,000 pages). Pages that fail (timeouts, 404s, blocked pages), PDFs and pages with no text are not charged. Apify's free plan includes monthly credit, so you can convert thousands of pages at no cost. Set a maximum cost per run and the Actor stops as soon as it is reached.
Input
See the Input tab for all options. Example:
{"urls": ["https://en.wikipedia.org/wiki/Markdown","https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview","https://github.com/mozilla/readability"],"mainContentOnly": true,"includeLinks": true,"includeImages": false,"maxCharacters": 0,"maxConcurrency": 25}
Output
One item per URL. Example (Markdown shortened):
{"url": "https://developer.mozilla.org/en-US/docs/Web/HTTP/Overview","finalUrl": "https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Overview","statusCode": 200,"title": "Overview of HTTP - HTTP | MDN","description": "HTTP is a protocol for fetching resources such as HTML documents.","language": "en-US","author": null,"publishedAt": null,"markdown": "# Overview of HTTP\n\n**HTTP** is a [protocol](https://developer.mozilla.org/en-US/docs/Glossary/Protocol) for fetching resources such as HTML documents...\n\n## Components of HTTP-based systems\n\n...","wordCount": 2400,"links": ["https://developer.mozilla.org/en-US/docs/Glossary/Protocol", "..."],"error": null}
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
Tips
- Save tokens: turn off Include links for plain prose, and set Max characters per page to cap long pages.
- Whole page: turn off Main content only for landing pages, pricing pages or link directories where every block matters.
- JavaScript-only sites: this Actor reads the HTML the server sends. Single-page apps that render text only in the browser may return little content; use a browser-based crawler such as Website Content Crawler for those.
- Blocked sites: enable Apify Proxy if many URLs return 403 or 429.
FAQ
Is this a Jina Reader or Firecrawl alternative?
Yes for the common case "give me this URL as clean Markdown". It does one thing: fetch the page over HTTP, extract the main content and convert it to Markdown, with simple per-page pricing and no subscription.
How is the main content detected?
It uses Mozilla's Readability algorithm (the engine behind Firefox Reader View) plus extra cleanup of navigation, cookie banners, ads and share widgets. Short pages such as product or docs index pages fall back to the cleaned <main> element so content isn't lost.
Are tables and code blocks preserved?
Yes. Tables become GitHub-flavoured Markdown tables (including tables without a header row) and code blocks become fenced blocks with the language when the page declares it.
Does it support PDFs?
Not yet. PDF URLs return an item with error: "PDF not supported" and are not charged.
Does it crawl whole websites?
No, it converts exactly the URLs you give it. Use the links field to pick the next pages, or use a crawler Actor for full-site crawls.
Support
Found a page that converts badly? Open an issue in the Issues tab with the URL.
This Actor only reads publicly available web pages. It does not log in, solve CAPTCHAs or collect personal data. You are responsible for respecting the terms of the websites you convert.
More tools from Tenfold Fleet
| Actor | Price |
|---|---|
| ATS Jobs Scraper - Greenhouse, Lever, Ashby & 5 More | $2 per 1,000 (job posting) |
| Website Contact Scraper - Emails, Phones & Socials | $8 per 1,000 (website with contacts) |
| Website Technology Detector - Wappalyzer Alternative | $10 per 1,000 (website analyzed) |
| YouTube Transcript Scraper - Captions & Subtitles API | $3 per 1,000 (transcript extracted) |