Microsoft MarkItDown Document Converter avatar

Microsoft MarkItDown Document Converter

Pricing

from $4.32 / 1,000 document converteds

Go to Apify Store
Microsoft MarkItDown Document Converter

Microsoft MarkItDown Document Converter

Convert bounded batches of supplied PDF, Office, HTML, CSV, JSON and other supported documents into Markdown with per-file status for RAG ingestion.

Pricing

from $4.32 / 1,000 document converteds

Rating

0.0

(0)

Developer

Automation Lab

Automation Lab

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Convert a bounded batch of documents to Markdown with Microsoft's open-source MarkItDown library. Supply file bytes as base64 or direct public HTTPS URLs; receive one dataset row per attempted file, with a conversion status and usable Markdown or an error. This document to markdown converter is designed for repeatable LLM and RAG ingestion, not website crawling.

Who is it for?

  • Knowledge-base maintainers turn public HTML documentation and office documents into text for indexing.
  • Data engineers normalize mixed-format uploads before chunking and embedding.
  • Analysts convert spreadsheet and CSV tables to Markdown for review alongside reports.

Why use this converter?

A single batch can mix PDF, DOCX, PPTX, XLSX, HTML, CSV, JSON and plain text instead of maintaining a separate extraction job per format. Failed items remain visible in the dataset and do not generate an item event. Position, original filename and source URL make it straightforward to reconcile the result with an upstream manifest. MarkItDown runs inside the Actor; no Microsoft account or Microsoft-hosted conversion endpoint is required. This independent Actor is not affiliated with or endorsed by Microsoft.

Getting started

  1. Add 1–20 entries to files, each with a name ending in a supported extension.
  2. Give each entry either a direct public HTTPS url or the file's base64-encoded base64 contents, not both.
  3. Run the Actor and download its default dataset in JSON, CSV or another supported Apify format.
  4. Filter for status === "success" before passing Markdown to your splitter, embeddings model or retrieval index.

For example, this public page is a small HTML conversion test:

{"files":[{"name":"python-home.html","url":"https://www.python.org/"}],"maxItems":1}

Supported input

FieldRequiredMeaning
filesYesArray of 1–20 objects. Each needs name and exactly one of url or base64.
files[].nameYesFilename ending in .pdf, .docx, .pptx, .xlsx, .html, .htm, .csv, .json, .png, .jpg, .jpeg, .wav, .mp3 or .txt.
files[].urlAlternativeDirect public HTTPS URL. No redirects, login, IP literals or private-network hosts.
files[].base64AlternativeBase64 file bytes; 15 MiB decoded maximum. Do not include a data-URL prefix.
maxItemsNoProcess only the first N entries, between 1 and 20; default 20.

Each downloaded file is limited to 15 MiB. The download timeout is 20 seconds and the individual converter timeout is 90 seconds. Input validation errors fail the run; per-file download or conversion errors instead produce an error row so other documents can still complete.

Output fields

The default dataset has one row per attempted file (including failures). It does not contain separate document or page datasets.

FieldMeaning
indexZero-based position in the processed batch.
nameFilename supplied in the input.
sourceUrlDownload URL or null for inline base64.
statussuccess when nonempty Markdown was extracted, otherwise error.
markdownExtracted Markdown; null on error.
titleConverter title when available; may be null.
bytesOriginal file byte count on success; null on failure.
errorError detail on failure; null on success.

A CSV with headers name,score and one row Ada,10 produces a Markdown table similar to:

{"index":0,"name":"scores.csv","sourceUrl":null,"status":"success","markdown":"| name | score |\n| --- | --- |\n| Ada | 10 |","title":null,"bytes":18,"error":null}

Exact whitespace, byte counts and extracted formatting vary by source file and library version. Use status to select successful rows, not title.

How much does it cost to convert documents to Markdown?

The Actor charges one $0.005 start event per run and one item event per successfully converted file. Error rows have no item charge. At the BRONZE spend tier, an item is $0.0072; FREE is $0.00828, SILVER $0.005616, and GOLD, PLATINUM and DIAMOND $0.00432 each. These are Apify Store spend tiers, based on the customer's qualifying aggregate monthly Store spend, not a discount for more files in one run. For example, at BRONZE, 1, 5 and 20 successful files cost approximately $0.0122, $0.041 and $0.149 per run respectively (one start event included). If an item fails, only successful files incur item events. The pricing panel is authoritative; estimated totals can vary with account tier, optional platform charges and refunds. The 20-document maximum bounds one run; split larger batches into separate runs.

Integrations and workflow ideas

  • Feed successful Markdown into a document chunker and vector index; store name and sourceUrl as citation metadata.
  • Run a scheduled Apify Task with public document URLs, then compare file-indexed results in downstream automation.
  • Supply base64 from an authorized ingestion pipeline when the original file is private; never expose private files through a temporary public URL merely for conversion.
  • Use error rows to retry or route unsupported scans to a dedicated OCR/transcription process.

Use the Apify API

Start a run with the Actor API (replace YOUR_APIFY_TOKEN with your own token):

curl -X POST 'https://api.apify.com/v2/acts/automation-lab~microsoft-markitdown-document-converter/runs?token=YOUR_APIFY_TOKEN' \
-H 'Content-Type: application/json' \
-d '{"files":[{"name":"python-home.html","url":"https://www.python.org/"}],"maxItems":1}'

JavaScript:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/microsoft-markitdown-document-converter').call({
files: [{ name: 'python-home.html', url: 'https://www.python.org/' }],
maxItems: 1,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.filter((item) => item.status === 'success'));

Python:

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/microsoft-markitdown-document-converter').call(run_input={
'files': [{'name': 'python-home.html', 'url': 'https://www.python.org/'}],
'maxItems': 1,
})
items = client.dataset(run['defaultDatasetId']).list_items().items
print([item for item in items if item['status'] == 'success'])

Use it through MCP

Connect Apify to Claude Code:

claude mcp add --transport http apify \
'https://mcp.apify.com?tools=automation-lab/microsoft-markitdown-document-converter'

Claude Desktop, Cursor and VS Code can use this MCP server URL in their remote-server settings. For a compatible desktop/editor JSON configuration:

{"mcpServers":{"apify":{"url":"https://mcp.apify.com?tools=automation-lab/microsoft-markitdown-document-converter"}}}

Example prompt for MCP: “Convert the public Python homepage HTML to Markdown and show its headings and links.” For a dataset workflow, ask “Convert the public Vega Seattle weather CSV to Markdown, then show only successful rows.” Avoid putting private documents or API tokens into an assistant prompt unless your organization's data policy permits it.

Limits and quality

MarkItDown preserves supported headings, links and tables when the source contains extractable structure, but this is not a pixel-perfect document renderer. Scanned PDFs and arbitrary images may require OCR; audio may require a separate transcription service. A synthetic image-with-text and silent WAV both produced error records in local tests, so do not rely on this Actor for OCR or transcription. Encrypted, malformed or empty documents may likewise fail. HTML is transformed as a document, not crawled for linked pages. Only direct HTTPS downloads are accepted; redirects are rejected. A failed item does not stop the rest of the batch.

Document format choices

Use DOCX for editable reports, PPTX for presentations, XLSX or CSV for tabular data, PDF for machine-readable exports, and HTML for direct pages. JSON can be converted to a textual representation, but this tool does not apply your own data model or schema. If table layout or exact page fidelity matters, inspect the Markdown before ingesting it. Test a representative document in your pipeline before scheduling repeated batches.

Legality, privacy and retention

Convert only documents you are entitled to process. Inputs (including base64 bytes), converted Markdown, filenames and source URLs can contain sensitive or personal information. Apify stores the input and default dataset under your account's storage retention and access settings; review those settings and delete runs/datasets when no longer needed. Source URLs, including query strings, are repeated in output rows. The Actor writes temporary files inside its runtime and removes them after each item; it does not deliberately retain a separate document archive or send files to Microsoft or an AI model. For URL inputs it downloads directly from the supplied public HTTPS host, so that host receives a normal network request. Apify's platform provides hosting and storage; the installed open-source MarkItDown package does not need a remote conversion provider. Do not put tokens in URLs, supply confidential documents to a public dataset, or treat this tool as a redaction service. Respect licensing and terms for public URLs.

Troubleshooting

  • URL rejected? Use a public direct HTTPS file URL without a redirect, IP address or authentication requirement, or pass authorized bytes as base64 instead.
  • Empty Markdown? Check whether the file is scanned, image-only, encrypted or unsupported by the installed converter; route it to OCR/transcription if appropriate.
  • Unexpected missing files? maxItems processes a prefix of files; increase it up to 20 or split your workload.
  • File too large? Reduce the file under 15 MiB or split it before upload. Look at each row's error field for item-specific failure details.

For single-format work, consider PDF to Structured Markdown Converter for PDF-specific output, or HTML Readability to Markdown Converter for web-page readability extraction. This Actor instead accepts heterogeneous supplied file batches.

FAQ

Does a conversion error charge for an item? No. The start event still applies to the run, but item events are emitted for successful conversions only.

Can I provide a login-protected URL? No. Provide the file bytes as base64 from your authorized workflow instead; the Actor will not log in or follow redirects.

Does it transcribe audio or recognize text in images? Not reliably. Some media formats may be recognized by the underlying library, but OCR and transcription are not guaranteed and the Actor reports failures as error rows.

Can I convert more than 20 documents? Run multiple batches of up to 20 files each.