PDF Text Extractor: Text, Pages & Metadata from PDF (OCR) avatar

PDF Text Extractor: Text, Pages & Metadata from PDF (OCR)

Pricing

from $3.00 / 1,000 documents

Go to Apify Store
PDF Text Extractor: Text, Pages & Metadata from PDF (OCR)

PDF Text Extractor: Text, Pages & Metadata from PDF (OCR)

Extract the text of any PDF (native or scanned with OCR), Word, PowerPoint, Excel or web page: full text, text per page, page count, title, author and dates as JSON, CSV or Excel. Bulk URLs or uploads, no API key. Pay per document.

Pricing

from $3.00 / 1,000 documents

Rating

0.0

(0)

Developer

Fernando Guiraud

Fernando Guiraud

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does PDF Text Extractor do?

PDF Text Extractor gets the text out of any PDF and gives it back as clean JSON, CSV or Excel: the full text, the text of each page, the page count and the document's title, author, language and dates. It reads native and scanned PDFs (OCR for pages without a text layer) and also Word (DOCX), PowerPoint (PPTX), Excel, web pages and images.

Paste PDF URLs (Google Drive, Dropbox and OneDrive share links work) or upload files, and get one row per document. No API key, no browser, and you pay only per document processed, however many pages it has.

It runs on the Apify platform, so you also get an API, scheduling, integrations (Google Sheets, Make, Zapier, n8n) and access for AI agents through the Apify MCP server.

Why use it?

  • 📄 Bulk PDF to text: contracts, reports, invoices, research papers, government forms, thousands at a time.
  • 🔎 Text per page: search, cite or split documents by page.
  • 🖨️ Scanned PDFs: pages without text are read with OCR (7 languages), and only those pages are billed as OCR.
  • 🏷️ Metadata included: title, author, subject, creator, creation and modification dates, page count and detected language.
  • 🧹 Clean text: repeated headers and footers removed, two-column layouts read in the right order.
  • 🤖 For AI and RAG: switch on markdown or chunks to get LLM-ready Markdown or pieces with page numbers.

How to extract text from a PDF

  1. Click Try for free.
  2. Paste one or more PDF (or DOCX/PPTX/XLSX/HTML) URLs, or upload files.
  3. Click Start.
  4. Read the Documents view, or download the dataset as JSON, CSV or Excel.

Input

FieldDescriptionDefault
sourcesDocument URLs (PDF, DOCX, PPTX, XLSX, HTML, images; share links work)sources or base64Files
base64FilesFiles sent inline (for APIs and AI agents)-
outputstext, pages, markdown, tables, chunkstext, pages
ocrauto (only pages without text), always or neverauto
pageRangee.g. 1-5, 8all pages
pdfPasswordPassword of encrypted PDFs-
followDocumentLinksAlso extract the documents linked on a web pagefalse
{
"sources": [{ "url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf" }]
}

Output

One row per document. Real result for the IRS Form W-9 (October 2026; text and pages shortened):

{
"source": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"status": "ok",
"fileType": "pdf",
"pageCount": 6,
"title": "Form W-9 (Rev. March 2024)",
"author": "SE:W:CAR:MP",
"language": "en",
"characters": 37694,
"ocrPages": 0,
"text": "W-9\nRequest for Taxpayer\nForm Give form to the\n(Rev. March 2024) Identification Number and Certification requester. …",
"pages": [
{ "page": 1, "text": "W-9\nRequest for Taxpayer …", "ocr": false },
{ "page": 2, "text": "Form W-9 (Rev. 3-2024) Page 2\nmust obtain your correct taxpayer identification number (TIN) …", "ocr": false }
],
"metadata": {
"subject": "Request for Taxpayer Identification Number and Certification",
"creator": "Designer 6.5",
"created": "2024-03-06T08:18:13",
"modified": "2024-03-06T08:18:13",
"pageCount": 6
}
}

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Data fields

FieldDescription
textThe whole document as plain text
pages[].page / pages[].text / pages[].ocrText of each page, and whether it was read with OCR
pageCount, charactersSize of the document
title, author, languageFrom the document metadata (language is detected from the text)
metadataSubject, creator, producer, creation and modification dates
ocrPagesPages read with OCR (billed as OCR pages)
status, warningsok or error, and anything worth knowing (for example, no text found)

How much does it cost to extract text from PDFs?

EventPrice
Run start (per GB of memory)$0.001
Document processed (any number of pages)$0.003
OCR page (only scanned pages without text)$0.01

1,000 native PDFs cost about $3, whatever their length. Failed documents are free. Set Max cost per run and the Actor stops cleanly at that limit.

Use it with AI agents (MCP)

Add https://mcp.apify.com?tools=fguiraud/pdf-text-extractor to Claude, Cursor or any MCP client and ask: "Read this contract and list the termination clauses: https://example.com/contract.pdf".

Smallest useful input for an agent: {"sources": [{"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf"}]}.

Agent-friendly input: common field names such as url, urls, startUrls, query or keywords are accepted, and an empty input runs a small sample from the input form (with a note in the results) instead of failing.

FAQ and limitations

  • Scanned PDFs: OCR is automatic for pages without text; ocr: "never" turns it off, "always" forces it.
  • Encrypted PDFs: enter the password in PDF password.
  • Very large documents: results over 9 MB drop the per-page text first and keep the full text, with a warning.
  • Found a problem or need a feature? Open an issue on the Issues tab. Replies within 48 hours.