PDF Text Extractor & PDF to Markdown - OCR, Word, Excel
Pricing
Pay per event
PDF Text Extractor & PDF to Markdown - OCR, Word, Excel
Document parser and PDF parser for LLMs: convert PDF, Word (DOCX/DOC), PowerPoint, Excel and CSV to clean Markdown text. PDF OCR for scanned files, RAG chunks. URLs, uploads, base64. $2/1,000 docs.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Yukai Lin
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 hours ago
Last modified
Categories
Share
What does PDF Text Extractor & PDF to Markdown do?
Give it documents (links, an uploaded file, or base64 from your code) and get back clean Markdown text for each one: ready for ChatGPT / Claude prompts, vector databases, RAG pipelines, search indexes or content migration.
Supported formats:
| Format | Extensions |
|---|---|
.pdf (text-based PDFs; scanned PDFs with the optional OCR) | |
| Microsoft Word | .docx, .doc, .rtf |
| Microsoft PowerPoint | .pptx, .ppt, .ppsx, .pps (slide text, via PDF) |
| Microsoft Excel | .xlsx, .xls (every sheet becomes a Markdown table) |
| OpenDocument | .odt, .ods, .odp |
| Data | .csv (as a table), .xml (text content only) |
| Web pages | HTML pages work too |
DOC, RTF, PPT/PPTX/PPS/PPSX and ODP files are first converted to PDF with LibreOffice on our own server, then to Markdown (rows show convertedFrom and via: "vm-office"). DOCX, XLSX/XLS and ODT/ODS are converted directly, and fall back to the same LibreOffice path if the direct conversion fails. Same price for every format.
- 📥 Three ways in: public URLs, a file upload, or base64 in the API input (for AI agents, n8n, Make, Zapier)
- 🔗 Cloud share links work: Google Drive, Google Docs/Sheets/Slides, Dropbox and OneDrive/SharePoint links are turned into direct downloads automatically
- 📄 Keeps structure: headings, paragraphs, lists and tables
- 🔍 OCR for scanned PDFs (optional): turn on OCR for scanned PDFs and PDFs without a text layer are transcribed page by page by an AI vision model, tables included, in most languages (tested with English, French, Chinese and Japanese)
- 🚦 Clear status per document:
statusissuccess,no_text,download_blocked,not_found,too_large,unsupportedorfailed; scanned PDFs are flagged withneedsOcr: truewhen OCR is off - 🧹 Cleaner PDF text: words hyphenated at line ends are joined ("repre-sentation" → "representation"), spaces lost between text blocks are put back ("AbstractWe" → "Abstract We") and simple number tables are rebuilt as Markdown tables; each item reports the changes in
pdfCleanup(turn off with Clean up PDF text) - 🧾 PDF facts:
pageCount,author,createdAt, pluswordCount,bytesandfileNamefor every document - ✂️ RAG chunks built in: optional
chunksarray split at paragraph boundaries with overlap, each with aheadingPath; or one row per chunk (outputMode: "chunks") for direct import into a vector database or CSV - ⚡ Fast: no browser needed; most documents are converted in about a second
- 💸 Simple pricing: $2 per 1,000 documents, no start fee, no compute charges
- 🛡️ Pay only for success: broken links, blocked downloads and empty documents are free
How much does it cost?
| Event | Price |
|---|---|
| Converted document | $2.00 / 1,000 documents |
| OCR page (optional, scanned PDFs only) | $5.00 / 1,000 pages |
No start fee. You pay per document, not per page: a 100-page PDF costs the same $0.002 as a one-page file. With per-page pricing, the page fee is multiplied by 100. Failed, blocked and empty (scanned) documents are free. Your maximum charge limit is always respected. On Apify's Scale plan prices are 10% lower, on Business and higher 20% lower.
OCR is charged only when you turn it on and a PDF has no text layer: $0.005 per transcribed page, on top of the $0.002 document price, up to OCR: max pages per document (default 20). Blank pages and failed pages are not charged. A 3-page scan costs $0.002 + 3 × $0.005 = $0.017. For comparison (checked September 2026), memo23 PDF Text Extractor charges $15 and yabanana99 PDF, Word & Excel to Markdown $12 per 1,000 OCR pages, on top of their per-document and per-run fees.
For AI agents that convert one file per call, there is no start fee to add: one 10-page text PDF costs $0.002. Tools that charge a $0.005 start fee per run cost about $0.0085–$0.01 for the same call.
Control your cost
- What is charged: each document converted successfully ($0.002, whatever its page count), and OCR pages when you turn OCR on.
- What is free: broken links (
not_found), blocked downloads, empty or scanned documents with OCR off, unsupported files and invalid input lines. Every row hascharged: trueorfalse. - At the start, the run logs its worst case, e.g.
Plan: 10 documents × $0.002 = at most $0.02(plus the OCR page price when OCR is on), and warns when that is more than your maximum charge per run. - When the maximum charge per run is reached, the run stops and keeps everything converted so far. The status message says so, and the
SUMMARYrecord hasstatus: "LIMIT_REACHED"andnotProcessed(how many documents were not started, and up to 100 of them). - If Apify restarts the run (server migration or Resurrect), items already finished are skipped and not charged again (
SUMMARY.resumedSkipped).
How to use it
- Paste the Document URLs, one per line (or upload a text/CSV file with one link per line in the request-list field under the options), or use Or upload a file for a document on your computer.
- Optional: set a RAG chunk size such as 2000 characters, and Output rows = one row per chunk.
- Optional: turn on OCR for scanned PDFs (and set a language hint for non-Latin scripts).
- Click Start and download the Markdown as JSON, CSV or Excel, or fetch it via API.
Input example
{"urls": ["https://arxiv.org/pdf/1706.03762","https://calibre-ebook.com/downloads/demos/demo.docx"],"chunkSize": 2000}
Google Drive link and base64 file (API)
{"urls": ["https://drive.google.com/file/d/0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx/view"],"base64Documents": [{ "fileName": "hello.csv", "data": "bmFtZSxjaXR5CkFkYSxMb25kb24K" }]}
curl -X POST "https://api.apify.com/v2/acts/tidytools~document-to-markdown/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d "{\"base64Documents\":[{\"fileName\":\"report.pdf\",\"data\":\"$(base64 -w0 report.pdf)\"}]}"
Output example
Real output from a test run (Markdown shortened). The Google Drive share link was converted to a direct download:
{"url": "https://drive.google.com/file/d/0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx/view","resolvedUrl": "https://drive.google.com/uc?export=download&id=0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx&confirm=t","source": "url","finalUrl": "https://drive.usercontent.google.com/download?id=0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx&export=download","fileName": "Sample.pdf","contentType": "application/pdf","bytes": 23567,"title": "Preface","pageCount": 8,"author": "Loren","createdAt": "2010-08-26T10:15:07-05:00","producer": "Acrobat Distiller 9.3.3 (Windows)","success": true,"status": "success","needsOcr": false,"via": "direct","characters": 13129,"wordCount": 2274,"markdown": "# Sample.pdf\n\n## Contents\n### Page 1\n…","convertedAt": "2026-09-29T02:26:12.454Z"}
A base64 CSV (hello.csv) comes back as "source": "base64", "url": null, "bytes": 21 and the table in Markdown:
{ "url": null, "source": "base64", "fileName": "hello.csv", "contentType": "text/csv", "status": "success", "markdown": "# hello.csv\n\n## Sheet1\n\n| name | city |\n|------|--------|\n| Ada | London |\n" }
With Output rows = one row per chunk, each chunk is its own row (the attention paper, https://arxiv.org/pdf/1706.03762, gave 38 chunks of up to 1,500 characters):
{ "url": "https://arxiv.org/pdf/1706.03762", "status": "success", "pageCount": 15, "chunkIndex": 30, "chunkCount": 38, "headingPath": ["document.pdf", "Contents", "Page 11"], "charCount": 1500, "text": "…" }
Scanned PDFs with OCR
{"documentUrls": [{ "url": "https://github.com/ocrmypdf/OCRmyPDF/raw/main/tests/resources/francais.pdf" }],"ocr": true,"ocrMaxPages": 20}
Real output (a scanned French page, no text layer). Charged: 1 document + 1 OCR page:
{"url": "https://github.com/ocrmypdf/OCRmyPDF/raw/main/tests/resources/francais.pdf","fileName": "francais.pdf","pageCount": 1,"producer": "Adobe Photoshop for Macintosh -- Image Conversion Plug-in","success": true,"status": "success","needsOcr": false,"ocr": true,"ocrPages": 1,"ocrPagesCharged": 1,"ocrTruncated": false,"ocrModel": "@cf/meta/llama-4-scout-17b-16e-instruct","characters": 363,"wordCount": 66,"markdown": "# Adobe Photoshop PDF\n\n## Page 1\n\nPortez ce vieux whisky au juge blond qui fume sur son île intérieure, à côté de l'alcôve ovoïde, où les bûches se consument dans l'âtre, …"}
- Each transcribed page starts with
## Page N, followed by the page's own headings, lists and tables. ocrTruncated: truemeans the PDF has more pages than OCR: max pages per document (a 7-page scan with a limit of 6 gave 6 pages and"pageCount": 7).- In our tests, a skewed English scan (OCRmyPDF
skew.pdf) came back word for word with its headings and bullet lists. A Japanese Wikipedia page came back with only a few wrong characters. A dense Traditional Chinese page had about one wrong character in twenty, so check names and figures in CJK scans.
Documents that cannot be converted are reported and not charged. With OCR off, a scanned PDF looks like this:
{ "url": "https://github.com/ocrmypdf/OCRmyPDF/raw/main/tests/resources/skew.pdf", "fileName": "skew.pdf", "success": false, "status": "no_text", "needsOcr": true, "pageCount": 1, "error": "the document contains no extractable text (it looks scanned; turn on \"OCR for scanned PDFs\" to transcribe it). (not charged)", "errorType": "no_text", "charged": false }
{ "url": "https://drive.google.com/file/d/0B9P1L--7Wd2vNm9zMTJWOGxobkU/view", "success": false, "status": "download_blocked", "error": "expected a document but the share link returned a web page (the file is not shared publicly, or it is too large for a direct download). Download the file yourself and send it as base64Documents (or upload it). (not charged)", "errorType": "blocked", "charged": false }
A site that refuses both download routes (real output, https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf, shortened):
{ "input": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf", "inputIndex": 0, "success": false, "status": "download_blocked", "triedRoutes": ["backend", "direct"], "error": "our servers: target returned HTTP 403; Apify network: target returned HTTP 403. Download the file yourself and send it as base64Documents (or upload it). (not charged)", "errorType": "blocked", "charged": false }
A link that does not exist gets "status": "not_found" (HTTP 404 or 410). URL rows carry input (the line as you entered it) and inputIndex (its position in the URL list), so results map back to your input even when documents finish in a different order.
The key-value store's SUMMARY record has status (SUCCESS, PARTIAL_RESULTS, FAILED, NO_RESULTS or LIMIT_REACHED) and counts per document status (plus ocrDocuments and ocrPagesCharged when OCR is on).
Use with AI agents (MCP)
Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/document-to-markdown) to Claude, Cursor or any MCP client, then ask e.g. "Convert https://arxiv.org/pdf/1706.03762 to Markdown and summarize section 3."
Minimal input:
{ "urls": ["https://arxiv.org/pdf/1706.03762"] }
Failed items are not charged and carry an errorType and a status. Agents that already have the file can send it as base64Documents.
Use it from code
curl -X POST "https://api.apify.com/v2/acts/tidytools~document-to-markdown/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"urls":["https://arxiv.org/pdf/1706.03762"]}'
Use the plain urls list from code: the older documentUrls field still works, but Apify rejects the whole run (HTTP 400) when it holds a blank line or a bare domain.
Schedules and integrations: run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.
Limitations
- Scanned PDFs are OCR-ed only when OCR for scanned PDFs is on; otherwise they come back with
status: "no_text"andneedsOcr: trueand are not charged. - PDF clean-up is heuristic: it only uses the document's own words as evidence, so most CamelCase names stay intact, but a rare glued word can remain, and tables are rebuilt only when each row ends with the same number of numeric cells (merged cells, multi-line headers and text tables stay as text).
- OCR applies to PDFs with no text at all. A PDF where only some pages are scanned is converted from its text layer, and the scanned pages stay empty.
- OCR is done by an AI vision model: it is very good on printed text, but it can misread characters (see the Chinese result above) and is not tuned for handwriting. Image files (PNG, JPG) are not supported yet.
- Large Google Drive files (over about 100 MB) show a virus-scan page instead of the file and are reported as
download_blocked; the 15 MB limit applies anyway. - If a site blocks our servers, Advanced settings → Auto retries the download from Apify's network; you can add an Apify proxy (billed to your Apify account).
- Files up to 15 MB; password-protected files are not supported.
- The document must be publicly downloadable (no login).
Related tools
- Website to Markdown Crawler: crawl a whole website or docs portal, including linked documents.
- AI Web Data Extractor: turn pages into structured JSON with the fields you choose.
FAQ
How do I extract text from a PDF? Paste the PDF link (or upload the file, or send base64 from your code) and run. Each document comes back as clean Markdown text with headings, lists and tables, plus pageCount and wordCount.
Does it OCR scanned PDFs? Yes, when you turn on OCR for scanned PDFs. PDFs without a text layer are transcribed page by page by an AI vision model for $5 per 1,000 pages, on top of the document price.
Which file types can it parse? PDF, Word (DOCX, DOC, RTF), PowerPoint (PPTX, PPT), Excel (XLSX, XLS), OpenDocument, CSV, XML and HTML pages, all at the same price.
Can I get chunks for RAG? Yes. Set RAG chunk size (characters) for a chunks array with headingPath, or use outputMode: "chunks" for one row per chunk, ready for a vector database or CSV.
Support
Open an issue in the Issues tab with the document URL. Issues are checked regularly.