HTML to Markdown Converter — Clean RAG Output avatar

HTML to Markdown Converter — Clean RAG Output

Pricing

from $3.00 / 1,000 converted documents

Go to Apify Store
HTML to Markdown Converter — Clean RAG Output

HTML to Markdown Converter — Clean RAG Output

HTML to Markdown converter and cleaner for LLM/RAG pipelines. Convert static or JavaScript pages, preserve headings, tables, links and images, remove noisy selectors, output clean text, and create overlap-aware chunks with stable IDs, SHA-256 fingerprints, word counts and token estimates.

Pricing

from $3.00 / 1,000 converted documents

Rating

0.0

(0)

Developer

Rosario Vitale

Rosario Vitale

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

HTML to Markdown Converter API — JS, Tables & RAG Chunks

Why use this Actor?

HTML to Markdown converter and cleaner for LLM/RAG pipelines. Convert static or JavaScript pages, preserve headings, tables, links and images, remove noisy selectors, output clean text, and create overlap-aware chunks with stable IDs, SHA-256 fingerprints, word counts and token estimates.

Features

  • Page URLs — Public HTTP/HTTPS page URLs to fetch and convert.
  • Optional CSS content selector — When provided, only the first matching element is converted.
  • Maximum Markdown characters per page — Maximum Markdown characters retained for each converted page.
  • Include cleaned HTML — Include cleaned HTML alongside the Markdown result.
  • Request timeout — Maximum seconds allowed for each network request.
  • User agent — HTTP User-Agent header sent to target websites.
  • Raw HTML documents — Optional objects with html, baseUrl and id. Convert raw HTML without fetching a URL.
  • Extra selectors to remove — Optional CSS selectors for ads, menus, widgets or other page noise.
  • Extract links — Return deduplicated absolute links with anchor text and rel metadata.
  • Extract images — Return image URLs, alt text, title and dimensions.
  • Include plain text — Add a cleaned plain-text representation alongside Markdown.
  • Create RAG chunks — Split Markdown into overlapping chunks ready for embeddings or LLM retrieval.

Use cases

  • Llm and rag ingestion.
  • Knowledge-base cleaning.
  • Web content migration.
  • Markdown dataset creation.

Example input

{
"urls": [
"https://example.com"
],
"maxMarkdownChars": 500000,
"includeCleanHtml": false,
"requestTimeoutSecs": 20,
"userAgent": "Mozilla/5.0 (compatible; ApifyHtmlMarkdown/1.0)",
"includeLinks": true
}

Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

FAQ

What is this Actor for?
It is designed for LLM and RAG ingestion, knowledge-base cleaning, web content migration.

Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

Search keywords

html to markdown, html to markdown converter, html to markdown python, html to markdown online, html to markdown table, html to markdown npm, html to markdown github, html to markdown converter online, html to markdown chrome extension, html to markdown c#, webpage to markdown, webpage to markdown chrome extension, webpage to markdown converter, webpage to markdown online

Turn public web pages into clean Markdown for LLMs, vector databases, RAG pipelines, knowledge bases, research tools, and content-processing automations.

The Actor fetches each URL, removes scripts, styles, templates and common navigation noise, prefers the page's main or article content when available, and converts cleaned HTML to Markdown. It also returns title, meta description, language, final URL and output size.

Input example

{"urls":["https://example.com"],"contentSelector":"","maxMarkdownChars":200000,"includeCleanHtml":false}

A CSS selector can target an exact content block. Optional cleaned HTML can be included for downstream parsers.

Reliability and spend controls

Inputs are deduplicated, only HTTP/HTTPS URLs are accepted, requests use explicit timeouts, one failed page does not abort a batch, and maxMarkdownChars prevents unexpectedly huge outputs.

Output

Successful document rows contain clean Markdown and metadata. Invalid or unavailable pages produce free diagnostic error rows.

Pricing

Target launch price: $0.001 per successfully converted document, competitive with the current Store niche. Failed pages are not billed as successful documents.

Responsible use

Process public pages in accordance with applicable website terms, copyright, privacy, robots policies and rate limits.

Support

For reproducible issues provide the public page URL, input options and Apify run ID. Never include private tokens or credentials.

Extended capabilities

  • Convert public URLs or raw HTML into Markdown, plain text, metadata, links, images, and bounded overlapping RAG chunks.
  • Use fast HTTP, full JavaScript browser rendering, or automatic browser fallback for thin scripted pages.
  • Preserve tables and resolve relative resources against the final/base URL.