# Microsoft MarkItDown Document Converter (`automation-lab/microsoft-markitdown-document-converter`) Actor

Convert bounded batches of supplied PDF, Office, HTML, CSV, JSON and other supported documents into Markdown with per-file status for RAG ingestion.

- **URL**: https://apify.com/automation-lab/microsoft-markitdown-document-converter.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.32 / 1,000 document converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Microsoft MarkItDown Document Converter

Convert a bounded batch of documents to Markdown with Microsoft's open-source MarkItDown library. Supply file bytes as base64 or direct public HTTPS URLs; receive one dataset row per attempted file, with a conversion status and usable Markdown or an error. This document to markdown converter is designed for repeatable LLM and RAG ingestion, not website crawling.

### Who is it for?

- Knowledge-base maintainers turn public HTML documentation and office documents into text for indexing.
- Data engineers normalize mixed-format uploads before chunking and embedding.
- Analysts convert spreadsheet and CSV tables to Markdown for review alongside reports.

### Why use this converter?

A single batch can mix PDF, DOCX, PPTX, XLSX, HTML, CSV, JSON and plain text instead of maintaining a separate extraction job per format. Failed items remain visible in the dataset and do not generate an item event. Position, original filename and source URL make it straightforward to reconcile the result with an upstream manifest. MarkItDown runs inside the Actor; no Microsoft account or Microsoft-hosted conversion endpoint is required. This independent Actor is not affiliated with or endorsed by Microsoft.

### Getting started

1. Add 1–20 entries to `files`, each with a `name` ending in a supported extension.
2. Give each entry either a direct public HTTPS `url` or the file's base64-encoded `base64` contents, not both.
3. Run the Actor and download its default dataset in JSON, CSV or another supported Apify format.
4. Filter for `status === "success"` before passing Markdown to your splitter, embeddings model or retrieval index.

For example, this public page is a small HTML conversion test:

```json
{"files":[{"name":"python-home.html","url":"https://www.python.org/"}],"maxItems":1}
```

### Supported input

| Field | Required | Meaning |
| --- | --- | --- |
| `files` | Yes | Array of 1–20 objects. Each needs `name` and exactly one of `url` or `base64`. |
| `files[].name` | Yes | Filename ending in `.pdf`, `.docx`, `.pptx`, `.xlsx`, `.html`, `.htm`, `.csv`, `.json`, `.png`, `.jpg`, `.jpeg`, `.wav`, `.mp3` or `.txt`. |
| `files[].url` | Alternative | Direct public HTTPS URL. No redirects, login, IP literals or private-network hosts. |
| `files[].base64` | Alternative | Base64 file bytes; 15 MiB decoded maximum. Do not include a data-URL prefix. |
| `maxItems` | No | Process only the first N entries, between 1 and 20; default 20. |

Each downloaded file is limited to 15 MiB. The download timeout is 20 seconds and the individual converter timeout is 90 seconds. Input validation errors fail the run; per-file download or conversion errors instead produce an error row so other documents can still complete.

### Output fields

The default dataset has one row per attempted file (including failures). It does not contain separate document or page datasets.

| Field | Meaning |
| --- | --- |
| `index` | Zero-based position in the processed batch. |
| `name` | Filename supplied in the input. |
| `sourceUrl` | Download URL or null for inline base64. |
| `status` | `success` when nonempty Markdown was extracted, otherwise `error`. |
| `markdown` | Extracted Markdown; null on error. |
| `title` | Converter title when available; may be null. |
| `bytes` | Original file byte count on success; null on failure. |
| `error` | Error detail on failure; null on success. |

A CSV with headers `name,score` and one row `Ada,10` produces a Markdown table similar to:

```json
{"index":0,"name":"scores.csv","sourceUrl":null,"status":"success","markdown":"| name | score |\n| --- | --- |\n| Ada | 10 |","title":null,"bytes":18,"error":null}
```

Exact whitespace, byte counts and extracted formatting vary by source file and library version. Use `status` to select successful rows, not `title`.

### How much does it cost to convert documents to Markdown?

The Actor charges one $0.005 `start` event per run and one `item` event per successfully converted file. Error rows have no item charge. At the BRONZE spend tier, an item is $0.0072; FREE is $0.00828, SILVER $0.005616, and GOLD, PLATINUM and DIAMOND $0.00432 each. These are **Apify Store spend tiers**, based on the customer's qualifying aggregate monthly Store spend, not a discount for more files in one run. For example, at BRONZE, 1, 5 and 20 successful files cost approximately $0.0122, $0.041 and $0.149 per run respectively (one start event included). If an item fails, only successful files incur item events. The pricing panel is authoritative; estimated totals can vary with account tier, optional platform charges and refunds. The 20-document maximum bounds one run; split larger batches into separate runs.

### Integrations and workflow ideas

- Feed successful Markdown into a document chunker and vector index; store `name` and `sourceUrl` as citation metadata.
- Run a scheduled Apify Task with public document URLs, then compare file-indexed results in downstream automation.
- Supply base64 from an authorized ingestion pipeline when the original file is private; never expose private files through a temporary public URL merely for conversion.
- Use error rows to retry or route unsupported scans to a dedicated OCR/transcription process.

### Use the Apify API

Start a run with the Actor API (replace `YOUR_APIFY_TOKEN` with your own token):

```bash
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~microsoft-markitdown-document-converter/runs?token=YOUR_APIFY_TOKEN' \
  -H 'Content-Type: application/json' \
  -d '{"files":[{"name":"python-home.html","url":"https://www.python.org/"}],"maxItems":1}'
```

JavaScript:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/microsoft-markitdown-document-converter').call({
  files: [{ name: 'python-home.html', url: 'https://www.python.org/' }],
  maxItems: 1,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.filter((item) => item.status === 'success'));
```

Python:

```python
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/microsoft-markitdown-document-converter').call(run_input={
    'files': [{'name': 'python-home.html', 'url': 'https://www.python.org/'}],
    'maxItems': 1,
})
items = client.dataset(run['defaultDatasetId']).list_items().items
print([item for item in items if item['status'] == 'success'])
```

### Use it through MCP

Connect Apify to Claude Code:

```bash
claude mcp add --transport http apify \
  'https://mcp.apify.com?tools=automation-lab/microsoft-markitdown-document-converter'
```

Claude Desktop, Cursor and VS Code can use this MCP server URL in their remote-server settings. For a compatible desktop/editor JSON configuration:

```json
{"mcpServers":{"apify":{"url":"https://mcp.apify.com?tools=automation-lab/microsoft-markitdown-document-converter"}}}
```

Example prompt for MCP: “Convert the public Python homepage HTML to Markdown and show its headings and links.” For a dataset workflow, ask “Convert the public Vega Seattle weather CSV to Markdown, then show only successful rows.” Avoid putting private documents or API tokens into an assistant prompt unless your organization's data policy permits it.

### Limits and quality

MarkItDown preserves supported headings, links and tables when the source contains extractable structure, but this is not a pixel-perfect document renderer. Scanned PDFs and arbitrary images may require OCR; audio may require a separate transcription service. A synthetic image-with-text and silent WAV both produced error records in local tests, so do not rely on this Actor for OCR or transcription. Encrypted, malformed or empty documents may likewise fail. HTML is transformed as a document, not crawled for linked pages. Only direct HTTPS downloads are accepted; redirects are rejected. A failed item does not stop the rest of the batch.

### Document format choices

Use DOCX for editable reports, PPTX for presentations, XLSX or CSV for tabular data, PDF for machine-readable exports, and HTML for direct pages. JSON can be converted to a textual representation, but this tool does not apply your own data model or schema. If table layout or exact page fidelity matters, inspect the Markdown before ingesting it. Test a representative document in your pipeline before scheduling repeated batches.

### Legality, privacy and retention

Convert only documents you are entitled to process. Inputs (including base64 bytes), converted Markdown, filenames and source URLs can contain sensitive or personal information. Apify stores the input and default dataset under your account's storage retention and access settings; review those settings and delete runs/datasets when no longer needed. Source URLs, including query strings, are repeated in output rows. The Actor writes temporary files inside its runtime and removes them after each item; it does not deliberately retain a separate document archive or send files to Microsoft or an AI model. For URL inputs it downloads directly from the supplied public HTTPS host, so that host receives a normal network request. Apify's platform provides hosting and storage; the installed open-source MarkItDown package does not need a remote conversion provider. Do not put tokens in URLs, supply confidential documents to a public dataset, or treat this tool as a redaction service. Respect licensing and terms for public URLs.

### Troubleshooting

- **URL rejected?** Use a public direct HTTPS file URL without a redirect, IP address or authentication requirement, or pass authorized bytes as base64 instead.
- **Empty Markdown?** Check whether the file is scanned, image-only, encrypted or unsupported by the installed converter; route it to OCR/transcription if appropriate.
- **Unexpected missing files?** `maxItems` processes a prefix of `files`; increase it up to 20 or split your workload.
- **File too large?** Reduce the file under 15 MiB or split it before upload. Look at each row's `error` field for item-specific failure details.

### Related automation-lab Actors

For single-format work, consider [PDF to Structured Markdown Converter](https://apify.com/automation-lab/pdf-to-structured-markdown-converter) for PDF-specific output, or [HTML Readability to Markdown Converter](https://apify.com/automation-lab/html-readability-markdown-converter) for web-page readability extraction. This Actor instead accepts heterogeneous supplied file batches.

### FAQ

**Does a conversion error charge for an item?** No. The start event still applies to the run, but item events are emitted for successful conversions only.

**Can I provide a login-protected URL?** No. Provide the file bytes as base64 from your authorized workflow instead; the Actor will not log in or follow redirects.

**Does it transcribe audio or recognize text in images?** Not reliably. Some media formats may be recognized by the underlying library, but OCR and transcription are not guaranteed and the Actor reports failures as error rows.

**Can I convert more than 20 documents?** Run multiple batches of up to 20 files each.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/microsoft-markitdown-document-converter/changelog.md

# Actor input Schema

## `files` (type: `array`):

Provide 1–20 files. Each item needs a filename with a supported extension and either url or base64 (not both). URLs must be public HTTPS direct downloads without redirects; each file is limited to 15 MiB.

## `maxItems` (type: `integer`):

Process at most this many files from the beginning of the list (1–20).

## Actor input object example

```json
{
  "files": [
    {
      "name": "python-home.html",
      "url": "https://www.python.org/"
    }
  ],
  "maxItems": 20
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset of conversions and individual errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "files": [
        {
            "name": "python-home.html",
            "url": "https://www.python.org/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/microsoft-markitdown-document-converter").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "files": [{
            "name": "python-home.html",
            "url": "https://www.python.org/",
        }] }

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/microsoft-markitdown-document-converter").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "files": [
    {
      "name": "python-home.html",
      "url": "https://www.python.org/"
    }
  ]
}' |
apify call automation-lab/microsoft-markitdown-document-converter --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/microsoft-markitdown-document-converter"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UICPMHUo95ndJcata/builds/v6Aqrx209u0Dq4k0r/openapi.json
