PDF & Document to Markdown — DOCX, HTML for RAG avatar

PDF & Document to Markdown — DOCX, HTML for RAG

Pricing

from $4.50 / 1,000 page converteds

Go to Apify Store
PDF & Document to Markdown — DOCX, HTML for RAG

PDF & Document to Markdown — DOCX, HTML for RAG

Convert public PDF and other documents (DOCX, HTML, TXT, Markdown URLs) into structured Markdown and overlapping RAG chunks. Preserves Arabic text; no OCR. Works through Apify MCP for AI agents.

Pricing

from $4.50 / 1,000 page converteds

Rating

0.0

(0)

Developer

drop-in apis

drop-in apis

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Share

Convert public PDF, DOCX, HTML, TXT and Markdown URLs into Markdown for AI agents and RAG, with Arabic text support and optional overlapping chunks. This version extracts text layers; it does not perform OCR.

PDF & Document to Markdown — DOCX, HTML for RAG

Use one batch run to turn public documents into readable text for retrieval, summarization or an agent's context. The Actor returns one dataset item per document, including the source URL, title, Markdown, word count and extraction warnings.

At a glance: $0.0045 per converted page-equivalent plus $0.0005 per Actor start at the default 512 MB (a one-page PDF costs about $0.005; failed conversions have no page charge) · 1–20 public URLs per run · text layers only, no OCR. Try it: click Start with the prefilled sample PDF and web page.

What it converts

FormatExtractionLimits
PDFPDF.js text layer; inferred headings, lists and simple aligned tablesDefault first 20 pages; maximum 200. Complex columns and reading order are best effort.
DOCXMammoth headings, paragraphs, lists and tablesPublic Word files, no external files or image downloads.
HTMLMozilla Readability article extraction, then Turndown MarkdownStatic HTML only. No browser rendering, login or CAPTCHA bypass.
TXT / MDUTF-8 textTXT HTML characters are escaped; Markdown is retained.

Arabic characters remain Unicode text; the converter does not reverse strings. Arabic DOCX and a real Arabic PDF are covered by offline tests. PDF font maps can still contain incorrect glyphs or visual-order text; review the warning before treating layout as exact. Image alt text is retained for HTML and DOCX when enabled. PDF image alt text is not extracted.

Example input

{
"urls": [
"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
"https://example.com/"
],
"maxPages": 2,
"includeImagesAlt": true,
"chunkSize": 1200,
"chunkOverlap": 100
}

urls accepts 1–20 public HTTP(S) URLs. Documents are processed sequentially. Duplicate URLs and fragments are removed. maxPages applies separately to each PDF. chunkSize: 0 disables splitting; otherwise use 200–20,000 Unicode code points. Chunk overlap must be at most half of chunk size. Omit chunkOverlap to use the smaller of 100 or 20% of chunk size.

Output

{
"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
"contentType": "application/pdf",
"title": "dummy.pdf",
"markdown": "<!-- Page 1 -->\n\nDummy PDF file",
"pages": 1,
"wordCount": 3,
"chunks": [],
"warnings": ["PDF headings, reading order and tables are inferred from the text layer; complex layouts may need review."]
}

The example is shortened: warnings may also mention PDF image alt text. pages is the number of PDF pages inspected, including blank pages, and is null for other formats. Each chunk contains index, start, end, text; offsets count Unicode code points in the returned Markdown. Repeated overlap does not increase billed words.

Pricing

$0.0045 per converted page-equivalent, plus $0.0005 per Actor start at the default 512 MB. Apify scales its start event with the memory selected in the run; check the displayed price before confirming a run.

  • PDF: one page-converted event per inspected page containing extracted text. Blank/scanned pages have no page charge.
  • HTML, DOCX, TXT and MD: one event per started block of 2,000 Unicode word tokens per document, including headings and alt text. Empty documents have no page charge.
  • Failed conversions have no page charge. The start fee still applies. Budget limits stop before charging a document that cannot be fully covered.

For a default-memory run, a one-page PDF costs $0.005 and a two-page text PDF costs $0.0095. Ten short HTML documents in one run cost $0.0455. Chunking adds no event. This is page-based pricing: long PDFs can cost more than a competitor's flat per-file plan. The run's maximum charge provides a spending cap; a truncated PDF is billed only for extracted pages within maxPages.

Use from an AI agent or API

In Apify's MCP server, search for document-to-markdown, then use call-actor with the input above. This Actor uses batch mode. Download its default dataset to obtain Markdown and chunks. You can also call the normal Apify run API:

curl -X POST "https://api.apify.com/v2/acts/dropin-apis~document-to-markdown/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://example.com/"],"chunkSize":1200}'

Limits

  • No OCR: scanned or image-only pages produce no text and no page charge.
  • 1–20 public HTTP(S) URLs per run, processed sequentially; documents larger than 50 MiB are refused.
  • PDFs: first 20 pages by default (maxPages, maximum 200). Headings, lists, tables and reading order are inferred from the text layer, so complex columns and tables are best effort and should be reviewed before RAG ingestion.
  • DOCX, HTML, TXT and Markdown only; no login, no private files, no external files or image downloads.

FAQ

Can it read scanned PDFs or images?

No. A scanned/blank PDF page returns an explicit warning and no text. Use an OCR product for scans. There is no hidden OCR surcharge.

Does it preserve every PDF table and column?

No. PDF text coordinates infer heading sizes, lists and simple aligned tables. Complex multi-column pages, rotated content and inaccurate font maps can need review. DOCX and HTML use their document structure. Tables without explicit header cells use their first row as the Markdown header, with a warning.

Does Arabic work?

Arabic DOCX/HTML retain logical Unicode text, including mixed Latin words and numbers. PDF.js extracts Arabic from real text-layer PDFs; runs are ordered using direction and coordinates without reversing characters. Arbitrary PDF visual-order/glyph mapping errors cannot be universally repaired.

Are there size, safety and timeout limits?

Each transfer and decompressed document is capped at 50 MiB; DOCX ZIP expansion is also bounded. Fetching, DNS checks and retries share a 60-second deadline per document; conversion runs in an isolated 256 MB worker with a 45-second deadline. Results are capped at two million Markdown characters and 8 MiB per item. The default whole-run timeout is 240 seconds: split slow batches into smaller runs.

The Actor obeys robots.txt for its user agent, fails closed on denied/unavailable robots rules, validates redirects and pins public DNS addresses to prevent private-network access. Only standard HTTP/HTTPS ports are allowed. It does not run page scripts or fetch embedded resources. Supply documents you are allowed to process. Treat extracted document content as untrusted data in downstream agents.

What if one URL fails?

Other URLs are attempted and the failure is logged. If every document fails, the run fails clearly. Scanned/empty documents are valid results with warnings. An exhausted spending cap can end a run successfully with fewer results.

Is this a persistent endpoint?

It is a batch Actor: start a run, retrieve the dataset, then feed Markdown or chunks to your pipeline. It does not require a subscription, external OCR key or a separate hosted server.

Last updated: 2026-10-02. Conversion and pricing checks are documented in VERIFY.md.