PDF Text Extractor — Text, Pages & Metadata from Any PDF Link
Pricing
$3.00 / 1,000 pdf delivereds
PDF Text Extractor — Text, Pages & Metadata from Any PDF Link
The text of any PDF at a public link — whole or page by page — with page count, word count and metadata. Scanned PDFs without text are flagged and not charged.
Pricing
$3.00 / 1,000 pdf delivereds
Rating
0.0
(0)
Developer
yestrue
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 hours ago
Last modified
Categories
Share
Get the text out of PDF files at public links — as one text, or page by page — with page count, word count and the document's metadata: title, subject, keywords, the program that made it, and creation and modification dates. Many PDFs per run.
Why this one
- 📄 Whole documents, fast. A 75-page research paper — 38,097 words — read in seconds in a live test run.
- 📑 Page by page if you need it. Switch on Text page by page for a list with one entry per page, for citations or chunking.
- 🖨️ Scanned PDFs are flagged, not charged. A PDF without a text layer comes back as
NO_TEXT— it would need OCR — instead of an empty success you pay for. - 🧾 Honest results. A link that returns a web page instead of a PDF, a 404, or a file over 100 MB comes back as a free record with the reason.
- 🔒 No personal data. The document's Author field is not delivered.
What you get
For every PDF:
| Field | What it is |
|---|---|
url, finalUrl | the link, and where it ended up after redirects |
text | the full text |
pages | with Text page by page on: the text of each page |
pageCount, pagesRead, characters, words, fileBytes | size |
title, subject, keywords | from the document's properties, when set |
creator, producer, pdfVersion, language | how it was made |
createdAt, modifiedAt | dates from the document's properties |
textTruncated | when Max text length cut the text |
scrapedAt, status, error | when it was read, and why a link gave nothing |
Example
A real record from a live run, 24 September 2026 (text shortened):
{"url": "https://arxiv.org/pdf/1706.03762","pageCount": 15,"characters": 39619,"words": 6120,"creator": "LaTeX with hyperref","producer": "pdfTeX-1.40.25","createdAt": "2024-04-10T21:11:43Z","text": "Provided proper attribution is provided, Google hereby grants permission to\nreproduce the tables and figures in this paper solely for use in journalistic or\nsch…","status": "OK"}
How to use it
- Add PDF links, one per line.
- Optionally limit Max pages per PDF, switch on Text page by page, or cap the text with Max text length.
- Press Start, and download the results from the Output tab as JSON, CSV or Excel.
Feed the text to an AI model, a search index or a spreadsheet — or call it from the Apify API, Make, Zapier or n8n, or from an AI agent through the Apify MCP server.
Pricing
You are charged per PDF delivered — see the price on this page. Scanned PDFs without text, links that are not PDFs, missing files and files over 100 MB are free.
FAQ
Does it read scanned PDFs? No — it reads the text layer. Scanned pages have none and are reported as NO_TEXT, free.
Are tables kept? Their text is, in reading order; the table layout is not.
Can it read password-protected PDFs? No.
Like it? A short review on the Store page helps other people find this Actor. Something missing or broken? Tell us on the Issues tab — we read every one.