PDF Text Extractor — Text, Pages & Metadata from Any PDF Link avatar

PDF Text Extractor — Text, Pages & Metadata from Any PDF Link

Pricing

$3.00 / 1,000 pdf delivereds

Go to Apify Store
PDF Text Extractor — Text, Pages & Metadata from Any PDF Link

PDF Text Extractor — Text, Pages & Metadata from Any PDF Link

The text of any PDF at a public link — whole or page by page — with page count, word count and metadata. Scanned PDFs without text are flagged and not charged.

Pricing

$3.00 / 1,000 pdf delivereds

Rating

0.0

(0)

Developer

yestrue

yestrue

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 hours ago

Last modified

Share

Get the text out of PDF files at public links — as one text, or page by page — with page count, word count and the document's metadata: title, subject, keywords, the program that made it, and creation and modification dates. Many PDFs per run.

Why this one

  • 📄 Whole documents, fast. A 75-page research paper — 38,097 words — read in seconds in a live test run.
  • 📑 Page by page if you need it. Switch on Text page by page for a list with one entry per page, for citations or chunking.
  • 🖨️ Scanned PDFs are flagged, not charged. A PDF without a text layer comes back as NO_TEXT — it would need OCR — instead of an empty success you pay for.
  • 🧾 Honest results. A link that returns a web page instead of a PDF, a 404, or a file over 100 MB comes back as a free record with the reason.
  • 🔒 No personal data. The document's Author field is not delivered.

What you get

For every PDF:

FieldWhat it is
url, finalUrlthe link, and where it ended up after redirects
textthe full text
pageswith Text page by page on: the text of each page
pageCount, pagesRead, characters, words, fileBytessize
title, subject, keywordsfrom the document's properties, when set
creator, producer, pdfVersion, languagehow it was made
createdAt, modifiedAtdates from the document's properties
textTruncatedwhen Max text length cut the text
scrapedAt, status, errorwhen it was read, and why a link gave nothing

Example

A real record from a live run, 24 September 2026 (text shortened):

{
"url": "https://arxiv.org/pdf/1706.03762",
"pageCount": 15,
"characters": 39619,
"words": 6120,
"creator": "LaTeX with hyperref",
"producer": "pdfTeX-1.40.25",
"createdAt": "2024-04-10T21:11:43Z",
"text": "Provided proper attribution is provided, Google hereby grants permission to\nreproduce the tables and figures in this paper solely for use in journalistic or\nsch…",
"status": "OK"
}

How to use it

  1. Add PDF links, one per line.
  2. Optionally limit Max pages per PDF, switch on Text page by page, or cap the text with Max text length.
  3. Press Start, and download the results from the Output tab as JSON, CSV or Excel.

Feed the text to an AI model, a search index or a spreadsheet — or call it from the Apify API, Make, Zapier or n8n, or from an AI agent through the Apify MCP server.

Pricing

You are charged per PDF delivered — see the price on this page. Scanned PDFs without text, links that are not PDFs, missing files and files over 100 MB are free.

FAQ

Does it read scanned PDFs? No — it reads the text layer. Scanned pages have none and are reported as NO_TEXT, free.

Are tables kept? Their text is, in reading order; the table layout is not.

Can it read password-protected PDFs? No.

Like it? A short review on the Store page helps other people find this Actor. Something missing or broken? Tell us on the Issues tab — we read every one.