HTML Readability to Markdown Converter avatar

HTML Readability to Markdown Converter

Pricing

from $1.44 / 1,000 item extracteds

Go to Apify Store
HTML Readability to Markdown Converter

HTML Readability to Markdown Converter

Convert raw HTML or one anonymous public webpage into clean Markdown with headings, links, source metadata, and a content hash for RAG ingestion or archives.

Pricing

from $1.44 / 1,000 item extracteds

Rating

0.0

(0)

Developer

Stas Persiianenko

Stas Persiianenko

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Convert supplied raw HTML or one anonymous public page URL into clean Markdown for RAG ingestion, documentation archives, content migration, and knowledge-base pipelines.

The Actor removes common navigation and boilerplate with Mozilla Readability, preserves useful headings and links, and returns source metadata plus a deterministic SHA-256 content hash.

What does HTML Readability to Markdown Converter do?

The Actor accepts exactly one of two sources:

  • a public HTTP(S) page URL; or
  • a raw HTML string supplied directly in the run input.

It produces one typed dataset record containing:

  • clean Markdown;
  • document title and description;
  • ordered headings;
  • unique absolute links;
  • canonical URL and language when declared;
  • word and character counts;
  • a content hash for deduplication or change detection; and
  • source and conversion timestamps.

For URL input, the Actor follows a bounded number of public redirects, rejects private-network destinations, retries temporary failures, and accepts HTML responses only.

Who is this HTML to Markdown converter for?

RAG and AI engineers

Normalize public documentation or supplied HTML before chunking, embedding, retrieval, or model evaluation.

Documentation teams

Archive important web documents in a portable text format while retaining headings, links, title, and canonical source information.

Data engineers

Feed stable dataset records into Apify integrations, webhooks, Make, Zapier, cloud storage, or custom ETL jobs.

Content migration teams

Turn legacy HTML fragments into readable Markdown without copying navigation, forms, scripts, or styling markup.

Researchers and compliance teams

Capture public policies, standards, or reference pages with a content hash that helps identify later changes.

Why use this Actor?

A generic HTML downloader returns markup that is noisy for language models and archives.

A basic tag replacer often retains menus, cookie notices, sidebars, and repeated site chrome.

This Actor combines:

  1. safe direct HTML retrieval for one public page;
  2. Readability-based main-content extraction;
  3. optional CSS-selector targeting for structured documentation;
  4. normalized absolute links;
  5. Markdown conversion with ATX headings and fenced code blocks; and
  6. typed source metadata for downstream automation.

It is intentionally focused.

It does not crawl a site, execute JavaScript, bypass login walls, or take screenshots.

Input parameters

FieldTypeRequiredDescription
urlstringOne source requiredAnonymous public HTTP(S) page to fetch and convert.
htmlstringOne source requiredRaw HTML to convert without fetching a page.
baseUrlstringNoPublic URL used to resolve relative links in raw HTML.
contentSelectorstringNoCSS selector such as main or article; overrides automatic Readability selection.
requestTimeoutSecsintegerNoURL request timeout from 5 to 120 seconds. Default: 30.
maxRequestRetriesintegerNoRetries for temporary failures from 0 to 5. Default: 2.
maxContentBytesintegerNoMaximum raw or fetched HTML size from 10,000 to 5,000,000 bytes.

Provide exactly one of url or html.

baseUrl is valid only with html.

Getting started

  1. Open the Actor in Apify Console.
  2. Enter a public page in Public page URL, or clear it and paste Raw HTML.
  3. Optionally set contentSelector when you know the page's main-content selector.
  4. Click Start.
  5. Open the Dataset tab after the run succeeds.
  6. Copy markdown, or download the record as JSON, JSONL, CSV, XML, Excel, or RSS through Apify's dataset tools.
  7. Use contentHash to detect duplicate or changed converted content in recurring workflows.

The prefilled Wikipedia URL is a working small example.

URL input example

{
"url": "https://developer.mozilla.org/en-US/docs/Web/HTML",
"contentSelector": "main"
}

This fetches the public MDN page, selects its main element, converts relative links to absolute links, and emits one Markdown document.

Raw HTML input example

{
"html": "<!doctype html><html><head><title>HTTP archive note</title></head><body><nav>Menu</nav><main><h1>HTTP Semantics</h1><p>Read the <a href='/rfc/rfc9110.html'>complete specification</a>.</p></main></body></html>",
"baseUrl": "https://www.rfc-editor.org/"
}

For your own input, replace the sample host with the real public base URL represented by the supplied HTML.

Relative links such as /guide become absolute when baseUrl is present.

Output fields

FieldMeaning
sourceTypeurl or raw_html.
sourceUrlNormalized requested URL or raw-HTML base URL; nullable.
finalUrlFinal fetched URL after redirects; null for raw HTML.
statusCodeHTTP status for URL input; null for raw HTML.
titleReadability title or HTML document title.
descriptionHTML meta description when present.
canonicalUrlAbsolute canonical URL when declared.
languageLanguage declared on the root HTML element.
bylineReadability byline when detected.
excerptReadability excerpt when detected.
markdownClean converted Markdown.
headingsOrdered objects with level and text.
linksUnique objects with visible text and absolute url.
wordCountApproximate words in selected readable content.
characterCountCharacters in the Markdown output.
contentHashSHA-256 hash of the Markdown.
convertedAtUTC ISO timestamp of conversion.

Nullable metadata remains null rather than being invented.

Output example

The current implementation produces records shaped like this:

{
"sourceType": "url",
"sourceUrl": "https://developer.mozilla.org/en-US/docs/Web/HTML",
"finalUrl": "https://developer.mozilla.org/en-US/docs/Web/HTML",
"statusCode": 200,
"title": "HTML: HyperText Markup Language | MDN",
"description": "HTML is the most basic building block of the Web.",
"canonicalUrl": "https://developer.mozilla.org/en-US/docs/Web/HTML",
"language": "en-US",
"byline": null,
"excerpt": null,
"markdown": "# HTML: HyperText Markup Language\n\n**HTML** defines the meaning and structure of web content...",
"headings": [
{ "level": 1, "text": "HTML: HyperText Markup Language" },
{ "level": 2, "text": "Key resources" }
],
"links": [
{ "text": "HTML element reference", "url": "https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Elements" }
],
"wordCount": 1089,
"characterCount": 13434,
"contentHash": "9d20a7f6d837c05663c9e65cf6714d698d5a9da9ab3fe09284a0835736cefd4e",
"convertedAt": "2026-08-29T14:05:00.000Z"
}

Page content and counts can change when the source changes.

How much does it cost to convert HTML to Markdown?

The Actor uses pay-per-event pricing:

  • $0.005 once when a run starts; and
  • the applicable tiered price for each converted Markdown document.

At the BRONZE tier, the current document price is $0.0024 per document.

Because this Actor returns one document per run, a successful BRONZE URL or raw-HTML conversion costs $0.0074 before any Apify subscription credits.

Higher subscription tiers receive lower per-document event prices.

Failed validation, failed retrieval, or empty conversion can still incur the one-time start event, but does not emit or charge a document event.

Always check the live pricing panel for the tier that applies to your account.

Readability and CSS selector behavior

When contentSelector is absent, Mozilla Readability identifies the main article-like content.

If the page is short or does not look like an article, the Actor falls back to a main, article, or document-body element.

When contentSelector is present:

  • it must be valid CSS;
  • it must match at least one element; and
  • the first matching element becomes the conversion source.

Selector mode is useful for API documentation, standards, and known templates.

It can be less portable if the source site changes its markup.

RAG ingestion workflow

A recurring RAG pipeline can:

  1. run the Actor for a public documentation page;
  2. read the single default-dataset record;
  3. compare contentHash with the previously stored hash;
  4. skip unchanged documents;
  5. chunk markdown by the objects in headings;
  6. retain sourceUrl and canonicalUrl as citation metadata; and
  7. embed only new or changed chunks.

The Actor does not create embeddings or choose a vector database.

This keeps conversion reusable across AI stacks.

Document archiving workflow

For repeatable archives:

  1. schedule one task per source page;
  2. export the dataset record to cloud storage;
  3. name the archive object with the date and contentHash;
  4. retain convertedAt, finalUrl, and statusCode; and
  5. alert in your own workflow when the hash differs.

Apify schedules can start tasks hourly, daily, weekly, or with a custom cron expression.

This Actor does not itself send alerts or retain history outside normal Apify run storage.

API usage with cURL

Start a run and wait for completion:

curl -X POST \
"https://api.apify.com/v2/acts/automation-lab~html-readability-markdown-converter/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url":"https://en.wikipedia.org/wiki/Markdown"}'

Keep your Apify token in a secret or environment variable.

Do not commit it to source control.

API usage with JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const input = {
url: 'https://developer.mozilla.org/en-US/docs/Web/HTML',
contentSelector: 'main',
};
const run = await client
.actor('automation-lab/html-readability-markdown-converter')
.call(input);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].markdown);

API usage with Python

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor(
"automation-lab/html-readability-markdown-converter"
).call(run_input={
"url": "https://nodejs.org/api/http.html",
"maxContentBytes": 5_000_000,
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(items[0]["title"])
print(items[0]["markdown"][:500])

Use with Apify MCP

Add this Actor to Claude Code:

claude mcp add --transport http apify \
"https://mcp.apify.com?tools=automation-lab/html-readability-markdown-converter"

Claude Desktop

Use this remote MCP server configuration in Claude Desktop:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=automation-lab/html-readability-markdown-converter"
}
}
}

Cursor

Add the same mcpServers.apify.url value in Cursor's MCP settings.

VS Code

Add the same remote Apify MCP URL in your VS Code MCP configuration, then enable the automation-lab/html-readability-markdown-converter tool.

Example prompts:

  • "Convert this public documentation URL to Markdown and summarize its headings."
  • "Run the HTML Readability to Markdown Converter for this page and return its canonical URL and content hash."
  • "Convert this raw HTML to Markdown using the supplied public base URL."

Integrations

The default dataset works with:

  • Apify webhooks;
  • Make;
  • Zapier;
  • Google Drive;
  • Google Sheets;
  • Slack;
  • GitHub Actions;
  • AWS S3; and
  • custom applications through the Apify API.

Large Markdown strings are generally best consumed as JSON or JSONL rather than spreadsheet cells.

Reliability and retry behavior

URL mode retries only temporary failures such as timeouts, HTTP 429, and selected HTTP 5xx responses.

Retries use bounded exponential backoff.

The Actor does not blindly retry:

  • malformed URLs;
  • private-network URLs;
  • permanent HTTP errors;
  • non-HTML content; or
  • anti-bot challenge pages.

Redirect destinations are validated again before fetching.

Raw HTML mode performs no network request except DNS validation when a baseUrl is supplied.

Legality and responsible use

Only process pages and HTML you are authorized to access.

Respect website terms, robots guidance, copyright, privacy, and applicable data-protection law.

Do not use the Actor to access internal services or private infrastructure.

The Actor rejects localhost, credentials in URLs, private IP ranges, link-local destinations, and unsafe redirect destinations.

Output may still contain personal or copyrighted information present in the supplied public content.

You are responsible for retention and downstream use.

Limitations

  • One raw HTML document or one public page URL is accepted per run.
  • URL mode performs direct HTTP retrieval and does not execute JavaScript.
  • Login-required, CAPTCHA-protected, and strongly bot-defended pages can fail.
  • The maximum HTML input or response size is 5 MB.
  • Readability is heuristic and can omit content on unusual layouts.
  • A custom selector uses only its first match.
  • Images are represented as Markdown references; image files are not downloaded.
  • Tables are converted using Turndown's standard behavior and may need downstream cleanup for complex layouts.
  • Content hashes change whenever the resulting Markdown changes.
  • The Actor is a converter, not a multi-page crawler or monitoring service.

Troubleshooting

The run says to provide exactly one input source

Clear either url or html.

The Actor deliberately rejects runs containing both or neither.

The selector did not match

Inspect the current page markup and update contentSelector, or remove it to use automatic Readability extraction.

The page returned HTTP 403 or an anti-bot challenge

The source does not permit this direct anonymous retrieval path.

Use supplied raw HTML when you can lawfully obtain it, or choose an Actor designed for browser rendering.

The page is JavaScript-rendered and Markdown is empty

This Actor does not render JavaScript.

Use the related Public Webpage HTML Downloader in rendered mode, then pass lawfully obtained HTML into this Actor.

Set baseUrl to the real public source URL represented by the HTML.

Without a base URL, only already-absolute HTTP(S) links can be retained.

The HTML is too large

Reduce the supplied HTML to the useful content, or use contentSelector with URL input.

The 5 MB ceiling is intentional for bounded memory and predictable runs.

Choose this Actor when the primary output you need is one clean Markdown document with headings, links, and source metadata.

FAQ

Can I convert multiple URLs in one run?

No.

The accepted product scope is one source document per run.

Create multiple Apify tasks or call the Actor once per document from your orchestrator.

Yes.

HTTP(S) links retained in the selected content are normalized to absolute URLs when a source or base URL is available.

Does it preserve headings?

Yes.

The Markdown contains ATX headings, and the headings array exposes their level and text separately.

Can I use raw HTML without a URL?

Yes.

Leave url empty and provide html.

Add baseUrl only when relative links need resolution.

Does it render client-side JavaScript?

No.

It converts the HTML returned by direct HTTP or the raw HTML supplied in input.

Can I detect document changes?

Yes.

Compare contentHash values from successful runs for the same source and settings.

Does the Actor store a Markdown file?

The Markdown is stored in the default dataset record.

Use the API or an integration to write it to a .md file or external archive.

What happens when conversion fails?

The run exits with a failure status and a concise reason.

It does not emit or charge a document record for failed conversion.