# HTML to Markdown Converter — Clean RAG Output (`zenomastro/html-to-markdown-cleaner`) Actor

HTML to Markdown converter and cleaner for LLM/RAG pipelines. Convert static or JavaScript pages, preserve headings, tables, links and images, remove noisy selectors, output clean text, and create overlap-aware chunks with stable IDs, SHA-256 fingerprints, word counts and token estimates.

- **URL**: https://apify.com/zenomastro/html-to-markdown-cleaner.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 converted documents

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## HTML to Markdown Converter API — JS, Tables & RAG Chunks

### Why use this Actor?

HTML to Markdown converter and cleaner for LLM/RAG pipelines. Convert static or JavaScript pages, preserve headings, tables, links and images, remove noisy selectors, output clean text, and create overlap-aware chunks with stable IDs, SHA-256 fingerprints, word counts and token estimates.

### Features

- **Page URLs** — Public HTTP/HTTPS page URLs to fetch and convert.
- **Optional CSS content selector** — When provided, only the first matching element is converted.
- **Maximum Markdown characters per page** — Maximum Markdown characters retained for each converted page.
- **Include cleaned HTML** — Include cleaned HTML alongside the Markdown result.
- **Request timeout** — Maximum seconds allowed for each network request.
- **User agent** — HTTP User-Agent header sent to target websites.
- **Raw HTML documents** — Optional objects with html, baseUrl and id. Convert raw HTML without fetching a URL.
- **Extra selectors to remove** — Optional CSS selectors for ads, menus, widgets or other page noise.
- **Extract links** — Return deduplicated absolute links with anchor text and rel metadata.
- **Extract images** — Return image URLs, alt text, title and dimensions.
- **Include plain text** — Add a cleaned plain-text representation alongside Markdown.
- **Create RAG chunks** — Split Markdown into overlapping chunks ready for embeddings or LLM retrieval.

### Use cases

- Llm and rag ingestion.
- Knowledge-base cleaning.
- Web content migration.
- Markdown dataset creation.

### Example input

```json
{
  "urls": [
    "https://example.com"
  ],
  "maxMarkdownChars": 500000,
  "includeCleanHtml": false,
  "requestTimeoutSecs": 20,
  "userAgent": "Mozilla/5.0 (compatible; ApifyHtmlMarkdown/1.0)",
  "includeLinks": true
}
```

### Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

### FAQ

**What is this Actor for?**\
It is designed for LLM and RAG ingestion, knowledge-base cleaning, web content migration.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

### Search keywords

html to markdown, html to markdown converter, html to markdown python, html to markdown online, html to markdown table, html to markdown npm, html to markdown github, html to markdown converter online, html to markdown chrome extension, html to markdown c#, webpage to markdown, webpage to markdown chrome extension, webpage to markdown converter, webpage to markdown online

Turn public web pages into clean Markdown for LLMs, vector databases, RAG pipelines, knowledge bases, research tools, and content-processing automations.

The Actor fetches each URL, removes scripts, styles, templates and common navigation noise, prefers the page's main or article content when available, and converts cleaned HTML to Markdown. It also returns title, meta description, language, final URL and output size.

### Input example

```json
{"urls":["https://example.com"],"contentSelector":"","maxMarkdownChars":200000,"includeCleanHtml":false}
```

A CSS selector can target an exact content block. Optional cleaned HTML can be included for downstream parsers.

### Reliability and spend controls

Inputs are deduplicated, only HTTP/HTTPS URLs are accepted, requests use explicit timeouts, one failed page does not abort a batch, and `maxMarkdownChars` prevents unexpectedly huge outputs.

### Output

Successful `document` rows contain clean Markdown and metadata. Invalid or unavailable pages produce free diagnostic `error` rows.

### Pricing

Target launch price: **$0.001 per successfully converted document**, competitive with the current Store niche. Failed pages are not billed as successful documents.

### Responsible use

Process public pages in accordance with applicable website terms, copyright, privacy, robots policies and rate limits.

### Support

For reproducible issues provide the public page URL, input options and Apify run ID. Never include private tokens or credentials.

### Extended capabilities

- Convert public URLs or raw HTML into Markdown, plain text, metadata, links, images, and bounded overlapping RAG chunks.
- Use fast HTTP, full JavaScript browser rendering, or automatic browser fallback for thin scripted pages.
- Preserve tables and resolve relative resources against the final/base URL.

# Actor input Schema

## `urls` (type: `array`):

Public HTTP/HTTPS page URLs to fetch and convert.

## `contentSelector` (type: `string`):

When provided, only the first matching element is converted.

## `maxMarkdownChars` (type: `integer`):

Maximum Markdown characters retained for each converted page.

## `includeCleanHtml` (type: `boolean`):

Include cleaned HTML alongside the Markdown result.

## `requestTimeoutSecs` (type: `integer`):

Maximum seconds allowed for each network request.

## `userAgent` (type: `string`):

HTTP User-Agent header sent to target websites.

## `htmlDocuments` (type: `array`):

Optional objects with html, baseUrl and id. Convert raw HTML without fetching a URL.

## `removeSelectors` (type: `array`):

Optional CSS selectors for ads, menus, widgets or other page noise.

## `includeLinks` (type: `boolean`):

Return deduplicated absolute links with anchor text and rel metadata.

## `includeImages` (type: `boolean`):

Return image URLs, alt text, title and dimensions.

## `includePlainText` (type: `boolean`):

Add a cleaned plain-text representation alongside Markdown.

## `includeChunks` (type: `boolean`):

Split Markdown into overlapping chunks ready for embeddings or LLM retrieval.

## `chunkSizeChars` (type: `integer`):

Approximate maximum characters per Markdown chunk.

## `chunkOverlapChars` (type: `integer`):

Characters overlapped between consecutive chunks.

## `renderMode` (type: `string`):

HTTP is fastest. Browser renders client-side JavaScript. Auto starts with HTTP and falls back to browser for failures or very thin scripted pages.

## `minStaticTextChars` (type: `integer`):

In auto mode, browser rendering is used when static body text is below this size and scripts are present.

## `browserWaitUntil` (type: `string`):

Readiness strategy for browser-rendered pages.

## `browserDelayMs` (type: `integer`):

Extra milliseconds after browser navigation before converting HTML.

## `browserWaitForSelector` (type: `string`):

Optional CSS selector that must become visible before browser HTML is captured.

## `proxyConfiguration` (type: `object`):

Optional Apify/custom proxy for browser rendering. Fast HTTP mode remains direct.

## Actor input object example

```json
{
  "urls": [
    "https://example.com"
  ],
  "contentSelector": "",
  "maxMarkdownChars": 500000,
  "includeCleanHtml": false,
  "requestTimeoutSecs": 20,
  "userAgent": "Mozilla/5.0 (compatible; ApifyHtmlMarkdown/1.0)",
  "htmlDocuments": [],
  "removeSelectors": [],
  "includeLinks": true,
  "includeImages": true,
  "includePlainText": true,
  "includeChunks": false,
  "chunkSizeChars": 6000,
  "chunkOverlapChars": 500,
  "renderMode": "http",
  "minStaticTextChars": 300,
  "browserWaitUntil": "domcontentloaded",
  "browserDelayMs": 500,
  "browserWaitForSelector": "",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/html-to-markdown-cleaner").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/html-to-markdown-cleaner").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call zenomastro/html-to-markdown-cleaner --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/html-to-markdown-cleaner"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gnwB7Qver4WYybuvi/builds/Cr8PY2T3U7jLgUO3Z/openapi.json
