PDF to Markdown & Document to Text for LLM/RAG, with OCR avatar

PDF to Markdown & Document to Text for LLM/RAG, with OCR

Pricing

Pay per usage

Go to Apify Store
PDF to Markdown & Document to Text for LLM/RAG, with OCR

PDF to Markdown & Document to Text for LLM/RAG, with OCR

Convert PDF, Word, PowerPoint, Excel, CSV, HTML, EPUB and images to clean Markdown or text, with tables, OCR for scans and RAG-ready chunks. Links or uploaded files, one row per document.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Tinlark

Tinlark

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

35 minutes ago

Last modified

Share

Convert PDF, Word, PowerPoint, Excel, CSV, HTML, EPUB and image files into clean Markdown or plain text. One Actor, one dataset row per document, or one row per chunk when you want ready-to-embed pieces for a RAG pipeline.

It keeps the structure that language models need: headings, lists and tables. Pages that are only a picture (scans, photos, screenshots) are read with OCR, and only those pages, so text PDFs stay fast.

Who it is for

  • Teams that feed documents to an LLM, a vector database or an AI agent and need Markdown instead of raw PDF text.
  • People who have a pile of reports, contracts, slide decks, spreadsheets and scans and want searchable text.
  • Anyone who wants one tool for many file types instead of a separate PDF, Word and OCR service.

What it does

  1. Downloads each link you give it (or reads files you upload), checks the real file type from the file's own bytes, and converts it.
  2. Writes Markdown and/or plain text, plus title, metadata, page count, language guess, word count, tables found, pages read with OCR and warnings.
  3. Optionally splits the text into chunks by headings or by size, with the heading path and page numbers of every chunk.
  4. A document that cannot be converted (broken link, password-protected PDF, unsupported file) becomes an error row with the reason. It does not stop the run and it is not charged.

Input

FieldWhat it doesDefault
Document links (sources)http(s) links to documents. The type is detected from the file, so links without an extension work.
Uploaded files (files)Upload files in the console; they are stored in your account and read from there.
Output format (outputFormat)markdown, text or bothmarkdown
OCR (ocr)auto: only pages with no text layer; off; force: every pageauto
OCR languages (ocrLanguages)English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Chinese (simplified), JapaneseEnglish
Max pages per document (maxPagesPerDocument)PDF pages, slides, sheets or EPUB chapters converted per document50
Page range (pageRange)For example 1-5, 8, 12-all
Tables as Markdown tables (includeTables)Detect tables in PDFson
Include document metadata (includeMetadata)Author, dates and other properties stored in the fileon
Chunking (chunking)off, by-headings or by-sizeoff
Chunk size / overlapCharacters. 2,000 characters are about 500 tokens of English2000 / 200
Max file size (maxFileSizeMb)Larger files get an error row50

Example input:

{
"sources": [
"https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
"https://example.com/slides.pptx"
],
"outputFormat": "markdown",
"chunking": "by-headings",
"chunkSize": 2000
}

Output

One row per document, or one row per chunk with chunking on. A shortened row:

{
"recordType": "document",
"status": "ok",
"url": "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
"fileName": "NIST.CSWP.04162018.pdf",
"fileType": "pdf",
"pages": 55,
"pagesConverted": 5,
"pagesOcr": 0,
"title": "Framework for Improving Critical Infrastructure Cybersecurity, Version 1.1",
"metadata": {"author": "National Institute of Standards and Technology", "created": "2018-04-17T13:45:53Z"},
"language": "en",
"tables": 1,
"wordCount": 1189,
"billablePages": 5,
"markdown": "# Framework for Improving Critical Infrastructure Cybersecurity\n\nVersion 1.1\n\n...",
"text": null,
"warnings": ["The document has 55 pages; only the first 5 were converted (limit: Max pages per document)."],
"error": null
}

A chunk row adds chunkIndex, chunkCount, headingPath (for example ["Introduction", "Scope"]), pageStart, pageEnd and tokensEstimate. An error row has status: "error", an errorCode (not-found, access-denied, too-large, encrypted, corrupt, unsupported-type, blocked-address, timeout and others) and an error text that says what to do. Rows are written as documents finish; sourceIndex is the position of the link in your input. A summary with totals is stored under the SUMMARY key.

Supported formats

FormatHow it is readNotes
PDFText layer with headings (by font size), lists, tables, two-column pages, repeated headers and footers removed; OCR for pages without textPassword-protected PDFs give a clear error
DOCXHeadings, lists, tables, bold and italic, linksOld .doc files are not supported: save as DOCX
PPTXOne section per slide: title, bullet levels, tables, speaker notesOld .ppt is not supported
XLSXOne Markdown table per sheet, values as last savedAt most 5,000 rows per sheet; old .xls is not supported
CSV / TSVOne Markdown tableDelimiter detected
HTMLMain content as Markdown; navigation, footers, forms and scripts droppedPages that build their text with JavaScript come back empty with a warning
EPUBOne section per chapter
TXT / MarkdownPassed through
PNG, JPEG, TIFF, GIF, BMP, WebPOCRMulti-page TIFF is read frame by frame; needs OCR on

OCR

With auto (the default) a PDF page is read with OCR only when it has no usable text layer. Text PDFs never touch the OCR engine. Images are always read with OCR (with OCR set to off they give an error row). Choose the languages that appear in your documents; each page is read with all languages you select. OCR is Tesseract. Each row reports pagesOcr and ocrConfidence (0 to 100), and a warning appears when the confidence is low. OCR is slower than text extraction: expect a few seconds per page on a full CPU core. Use 4096 MB of memory for OCR-heavy runs: it costs about the same per page and is several times faster.

Chunking for RAG

  • by-headings: one chunk per section under a Markdown heading. A section longer than the chunk size is cut further at paragraph ends.
  • by-size: chunks of about the chunk size, cut at paragraph boundaries, with overlap. Long tables are split by rows and every part repeats the header row.

Every chunk has its heading path and the pages it comes from, so answers can cite where they were found.

Limits and safety

  • Files up to 50 MB by default (200 MB at most). Downloads time out after 3 minutes and are retried with backoff on server errors.
  • Only public http(s) links. Links that point to private or internal network addresses are refused.
  • The file type comes from the file's bytes, not from the link or the Content-Type header.
  • A single document is converted in its own process and stopped when it takes too long, so one bad file cannot stop the others.
  • Default 50 pages per document; raise it up to 2,000.
  • Not supported: password-protected files, scanned handwriting, text inside images embedded in Word or PowerPoint files, JavaScript-rendered web pages, legacy .doc/.xls/.ppt, DRM-protected files.
  • Reading order of complicated layouts (magazines, posters) can be wrong, and tables without ruling lines are found only when their columns line up clearly. Check the output of documents where the layout matters.

Pricing

Free during launch (until 31 October 2026). You pay only Apify's own platform usage for your runs.

From 1 November 2026: pay per event, $1 per 1,000 converted pages, plus $4 per 1,000 pages read with OCR. A page is a PDF page, a slide, a sheet, an image, or for Word, HTML, text, CSV and EPUB files one page per started 500 words. Failed documents are never charged. Every row shows billablePages.

Apify platform usage in our test runs (Free plan compute price): about $0.04 per 1,000 text pages and about $1.30 to $1.60 per 1,000 OCR pages. Higher memory finishes sooner at about the same cost per page.

FAQ

Does it keep tables? Tables in Word, PowerPoint, Excel and HTML are kept as Markdown tables. In PDFs, ruled tables and tables with clearly aligned columns are found; merged header cells are collapsed.

Can I upload files instead of using links? Yes. Use the Uploaded files field in the console; the files are stored in your own account.

Does it send my documents anywhere? Documents are downloaded into the run, converted inside the run and deleted when the run ends. The only output is your dataset. No third-party AI or OCR service is called.

Why did a PDF come back empty? It probably has no text layer and OCR was set to off, or it is a scan of low quality. Check warnings and pagesOcr.

Can I call it from code or an AI agent? Yes, through the Apify API, the client libraries or Apify's MCP server: start the Actor with the input above and read the dataset items.

Disclaimers and legality. Process only documents you have the right to use. You are responsible for the lawful use of the documents you convert and of the output, including copyright, confidentiality and personal data in them. The Actor fetches exactly the links you give it, as a browser would; it does not search, crawl or follow links, and it does not log in to anything. Do not use it to get around access controls or terms of a site. Conversion is automatic and can contain errors: do not rely on it alone for legal, medical or financial decisions.

  • Audio & Podcast Transcriber: turns audio and video links and podcast RSS feeds into text and subtitles, so recordings and documents can go into the same RAG or search pipeline.

Changelog

  • 2026-10-03: added the Related Tinlark Actors section.

Support

Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the document link or file type) and what you expected.