HTML to Markdown Converter — Clean RAG Output
Pricing
from $3.00 / 1,000 converted documents
HTML to Markdown Converter — Clean RAG Output
HTML to Markdown converter and cleaner for LLM/RAG pipelines. Convert static or JavaScript pages, preserve headings, tables, links and images, remove noisy selectors, output clean text, and create overlap-aware chunks with stable IDs, SHA-256 fingerprints, word counts and token estimates.
Pricing
from $3.00 / 1,000 converted documents
Rating
0.0
(0)
Developer
Rosario Vitale
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
HTML to Markdown Converter API — JS, Tables & RAG Chunks
Why use this Actor?
HTML to Markdown converter and cleaner for LLM/RAG pipelines. Convert static or JavaScript pages, preserve headings, tables, links and images, remove noisy selectors, output clean text, and create overlap-aware chunks with stable IDs, SHA-256 fingerprints, word counts and token estimates.
Features
- Page URLs — Public HTTP/HTTPS page URLs to fetch and convert.
- Optional CSS content selector — When provided, only the first matching element is converted.
- Maximum Markdown characters per page — Maximum Markdown characters retained for each converted page.
- Include cleaned HTML — Include cleaned HTML alongside the Markdown result.
- Request timeout — Maximum seconds allowed for each network request.
- User agent — HTTP User-Agent header sent to target websites.
- Raw HTML documents — Optional objects with html, baseUrl and id. Convert raw HTML without fetching a URL.
- Extra selectors to remove — Optional CSS selectors for ads, menus, widgets or other page noise.
- Extract links — Return deduplicated absolute links with anchor text and rel metadata.
- Extract images — Return image URLs, alt text, title and dimensions.
- Include plain text — Add a cleaned plain-text representation alongside Markdown.
- Create RAG chunks — Split Markdown into overlapping chunks ready for embeddings or LLM retrieval.
Use cases
- Llm and rag ingestion.
- Knowledge-base cleaning.
- Web content migration.
- Markdown dataset creation.
Example input
{"urls": ["https://example.com"],"maxMarkdownChars": 500000,"includeCleanHtml": false,"requestTimeoutSecs": 20,"userAgent": "Mozilla/5.0 (compatible; ApifyHtmlMarkdown/1.0)","includeLinks": true}
Pricing & cost control
Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.
FAQ
What is this Actor for?
It is designed for LLM and RAG ingestion, knowledge-base cleaning, web content migration.
Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.
How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.
Search keywords
html to markdown, html to markdown converter, html to markdown python, html to markdown online, html to markdown table, html to markdown npm, html to markdown github, html to markdown converter online, html to markdown chrome extension, html to markdown c#, webpage to markdown, webpage to markdown chrome extension, webpage to markdown converter, webpage to markdown online
Turn public web pages into clean Markdown for LLMs, vector databases, RAG pipelines, knowledge bases, research tools, and content-processing automations.
The Actor fetches each URL, removes scripts, styles, templates and common navigation noise, prefers the page's main or article content when available, and converts cleaned HTML to Markdown. It also returns title, meta description, language, final URL and output size.
Input example
{"urls":["https://example.com"],"contentSelector":"","maxMarkdownChars":200000,"includeCleanHtml":false}
A CSS selector can target an exact content block. Optional cleaned HTML can be included for downstream parsers.
Reliability and spend controls
Inputs are deduplicated, only HTTP/HTTPS URLs are accepted, requests use explicit timeouts, one failed page does not abort a batch, and maxMarkdownChars prevents unexpectedly huge outputs.
Output
Successful document rows contain clean Markdown and metadata. Invalid or unavailable pages produce free diagnostic error rows.
Pricing
Target launch price: $0.001 per successfully converted document, competitive with the current Store niche. Failed pages are not billed as successful documents.
Responsible use
Process public pages in accordance with applicable website terms, copyright, privacy, robots policies and rate limits.
Support
For reproducible issues provide the public page URL, input options and Apify run ID. Never include private tokens or credentials.
Extended capabilities
- Convert public URLs or raw HTML into Markdown, plain text, metadata, links, images, and bounded overlapping RAG chunks.
- Use fast HTTP, full JavaScript browser rendering, or automatic browser fallback for thin scripted pages.
- Preserve tables and resolve relative resources against the final/base URL.