PDF to Markdown & Document to Text for LLM/RAG, with OCR
Pricing
Pay per usage
PDF to Markdown & Document to Text for LLM/RAG, with OCR
Convert PDF, Word, PowerPoint, Excel, CSV, HTML, EPUB and images to clean Markdown or text, with tables, OCR for scans and RAG-ready chunks. Links or uploaded files, one row per document.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Tinlark
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
35 minutes ago
Last modified
Categories
Share
Convert PDF, Word, PowerPoint, Excel, CSV, HTML, EPUB and image files into clean Markdown or plain text. One Actor, one dataset row per document, or one row per chunk when you want ready-to-embed pieces for a RAG pipeline.
It keeps the structure that language models need: headings, lists and tables. Pages that are only a picture (scans, photos, screenshots) are read with OCR, and only those pages, so text PDFs stay fast.
Who it is for
- Teams that feed documents to an LLM, a vector database or an AI agent and need Markdown instead of raw PDF text.
- People who have a pile of reports, contracts, slide decks, spreadsheets and scans and want searchable text.
- Anyone who wants one tool for many file types instead of a separate PDF, Word and OCR service.
What it does
- Downloads each link you give it (or reads files you upload), checks the real file type from the file's own bytes, and converts it.
- Writes Markdown and/or plain text, plus title, metadata, page count, language guess, word count, tables found, pages read with OCR and warnings.
- Optionally splits the text into chunks by headings or by size, with the heading path and page numbers of every chunk.
- A document that cannot be converted (broken link, password-protected PDF, unsupported file) becomes an error row with the reason. It does not stop the run and it is not charged.
Input
| Field | What it does | Default |
|---|---|---|
Document links (sources) | http(s) links to documents. The type is detected from the file, so links without an extension work. | |
Uploaded files (files) | Upload files in the console; they are stored in your account and read from there. | |
Output format (outputFormat) | markdown, text or both | markdown |
OCR (ocr) | auto: only pages with no text layer; off; force: every page | auto |
OCR languages (ocrLanguages) | English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Chinese (simplified), Japanese | English |
Max pages per document (maxPagesPerDocument) | PDF pages, slides, sheets or EPUB chapters converted per document | 50 |
Page range (pageRange) | For example 1-5, 8, 12- | all |
Tables as Markdown tables (includeTables) | Detect tables in PDFs | on |
Include document metadata (includeMetadata) | Author, dates and other properties stored in the file | on |
Chunking (chunking) | off, by-headings or by-size | off |
| Chunk size / overlap | Characters. 2,000 characters are about 500 tokens of English | 2000 / 200 |
Max file size (maxFileSizeMb) | Larger files get an error row | 50 |
Example input:
{"sources": ["https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf","https://example.com/slides.pptx"],"outputFormat": "markdown","chunking": "by-headings","chunkSize": 2000}
Output
One row per document, or one row per chunk with chunking on. A shortened row:
{"recordType": "document","status": "ok","url": "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf","fileName": "NIST.CSWP.04162018.pdf","fileType": "pdf","pages": 55,"pagesConverted": 5,"pagesOcr": 0,"title": "Framework for Improving Critical Infrastructure Cybersecurity, Version 1.1","metadata": {"author": "National Institute of Standards and Technology", "created": "2018-04-17T13:45:53Z"},"language": "en","tables": 1,"wordCount": 1189,"billablePages": 5,"markdown": "# Framework for Improving Critical Infrastructure Cybersecurity\n\nVersion 1.1\n\n...","text": null,"warnings": ["The document has 55 pages; only the first 5 were converted (limit: Max pages per document)."],"error": null}
A chunk row adds chunkIndex, chunkCount, headingPath (for example ["Introduction", "Scope"]), pageStart, pageEnd and tokensEstimate. An error row has status: "error", an errorCode (not-found, access-denied, too-large, encrypted, corrupt, unsupported-type, blocked-address, timeout and others) and an error text that says what to do. Rows are written as documents finish; sourceIndex is the position of the link in your input. A summary with totals is stored under the SUMMARY key.
Supported formats
| Format | How it is read | Notes |
|---|---|---|
| Text layer with headings (by font size), lists, tables, two-column pages, repeated headers and footers removed; OCR for pages without text | Password-protected PDFs give a clear error | |
| DOCX | Headings, lists, tables, bold and italic, links | Old .doc files are not supported: save as DOCX |
| PPTX | One section per slide: title, bullet levels, tables, speaker notes | Old .ppt is not supported |
| XLSX | One Markdown table per sheet, values as last saved | At most 5,000 rows per sheet; old .xls is not supported |
| CSV / TSV | One Markdown table | Delimiter detected |
| HTML | Main content as Markdown; navigation, footers, forms and scripts dropped | Pages that build their text with JavaScript come back empty with a warning |
| EPUB | One section per chapter | |
| TXT / Markdown | Passed through | |
| PNG, JPEG, TIFF, GIF, BMP, WebP | OCR | Multi-page TIFF is read frame by frame; needs OCR on |
OCR
With auto (the default) a PDF page is read with OCR only when it has no usable text layer. Text PDFs never touch the OCR engine. Images are always read with OCR (with OCR set to off they give an error row). Choose the languages that appear in your documents; each page is read with all languages you select. OCR is Tesseract. Each row reports pagesOcr and ocrConfidence (0 to 100), and a warning appears when the confidence is low. OCR is slower than text extraction: expect a few seconds per page on a full CPU core. Use 4096 MB of memory for OCR-heavy runs: it costs about the same per page and is several times faster.
Chunking for RAG
by-headings: one chunk per section under a Markdown heading. A section longer than the chunk size is cut further at paragraph ends.by-size: chunks of about the chunk size, cut at paragraph boundaries, with overlap. Long tables are split by rows and every part repeats the header row.
Every chunk has its heading path and the pages it comes from, so answers can cite where they were found.
Limits and safety
- Files up to 50 MB by default (200 MB at most). Downloads time out after 3 minutes and are retried with backoff on server errors.
- Only public http(s) links. Links that point to private or internal network addresses are refused.
- The file type comes from the file's bytes, not from the link or the Content-Type header.
- A single document is converted in its own process and stopped when it takes too long, so one bad file cannot stop the others.
- Default 50 pages per document; raise it up to 2,000.
- Not supported: password-protected files, scanned handwriting, text inside images embedded in Word or PowerPoint files, JavaScript-rendered web pages, legacy
.doc/.xls/.ppt, DRM-protected files. - Reading order of complicated layouts (magazines, posters) can be wrong, and tables without ruling lines are found only when their columns line up clearly. Check the output of documents where the layout matters.
Pricing
Free during launch (until 31 October 2026). You pay only Apify's own platform usage for your runs.
From 1 November 2026: pay per event, $1 per 1,000 converted pages, plus $4 per 1,000 pages read with OCR. A page is a PDF page, a slide, a sheet, an image, or for Word, HTML, text, CSV and EPUB files one page per started 500 words. Failed documents are never charged. Every row shows billablePages.
Apify platform usage in our test runs (Free plan compute price): about $0.04 per 1,000 text pages and about $1.30 to $1.60 per 1,000 OCR pages. Higher memory finishes sooner at about the same cost per page.
FAQ
Does it keep tables? Tables in Word, PowerPoint, Excel and HTML are kept as Markdown tables. In PDFs, ruled tables and tables with clearly aligned columns are found; merged header cells are collapsed.
Can I upload files instead of using links? Yes. Use the Uploaded files field in the console; the files are stored in your own account.
Does it send my documents anywhere? Documents are downloaded into the run, converted inside the run and deleted when the run ends. The only output is your dataset. No third-party AI or OCR service is called.
Why did a PDF come back empty? It probably has no text layer and OCR was set to off, or it is a scan of low quality. Check warnings and pagesOcr.
Can I call it from code or an AI agent? Yes, through the Apify API, the client libraries or Apify's MCP server: start the Actor with the input above and read the dataset items.
Disclaimers and legality. Process only documents you have the right to use. You are responsible for the lawful use of the documents you convert and of the output, including copyright, confidentiality and personal data in them. The Actor fetches exactly the links you give it, as a browser would; it does not search, crawl or follow links, and it does not log in to anything. Do not use it to get around access controls or terms of a site. Conversion is automatic and can contain errors: do not rely on it alone for legal, medical or financial decisions.
Related Tinlark Actors
- Audio & Podcast Transcriber: turns audio and video links and podcast RSS feeds into text and subtitles, so recordings and documents can go into the same RAG or search pipeline.
Changelog
- 2026-10-03: added the Related Tinlark Actors section.
Support
Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the document link or file type) and what you expected.