PDF Text Extractor — Text, Pages & Metadata, Scan Detection
Pricing
from $3.00 / 1,000 processed pdfs
PDF Text Extractor — Text, Pages & Metadata, Scan Detection
Extract text and metadata from PDF files by URL, whole or page by page. Detects scanned PDFs that need OCR and does not charge for them. Safe PDF parsing (pdf.js, no script execution), size and time limits, no data in logs.
Pricing
from $3.00 / 1,000 processed pdfs
Rating
0.0
(0)
Developer
San Marino Tools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Extract the text of PDF files from their URLs: the whole document or page by page, with page count and metadata. Scanned PDFs without a text layer are detected and not charged.
What you get
| Field | Meaning |
|---|---|
url | the URL you gave |
status | ok, no_text_layer, encrypted, not_pdf, too_large, download_failed, parse_error |
error | machine-readable detail when status is not ok (e.g. http_404, timeout, password_required, blocked_address), otherwise null |
pageCount | pages in the document |
pagesExtracted | pages actually extracted (see Max pages per PDF) |
charCount | characters of extracted text |
needsOcr | true when at least half of the pages have no text layer (typical of scans) |
text | the text, pages separated by a blank line — or, with Text page by page on, pages: [{ page, text }] |
metadata | title, author, subject, creator, producer, creationDate, modDate (ISO 8601) |
{"url": "https://example.com/invoice.pdf","status": "ok","error": null,"pageCount": 2,"pagesExtracted": 2,"charCount": 77,"needsOcr": false,"text": "Invoice 2026-001\nTotal due: 1,234.56 EUR\n\nPage two\nThank you for your business.","metadata": { "title": "Test invoice", "author": "Test author", "subject": null, "creator": null, "producer": "…", "creationDate": "2026-01-15T10:00:00.000Z", "modDate": null }}
A run summary (counts per status, duplicates, charged events, limits used) is saved in the key-value store as SUMMARY — open it from the Output tab, Run summary.
Very long texts (results over about 8 MB) are stored in the key-value store as a separate record; the result then has text: null and textRecordKey with the record name.
Input
| Field | Default | |
|---|---|---|
pdfUrls | — | public http(s) links, one per line |
outputPerPage | false | page-by-page output |
maxPagesPerPdf | 500 | only the first N pages are extracted and charged (max 2,000) |
maxFileSizeMb | 50 | larger files are not downloaded (max 200) |
Pricing
Pay per event: $0.003 per PDF processed successfully, up to 100 pages. Longer documents: one more event for every started block of 100 pages beyond the first 100.
| Pages extracted | Events | Price |
|---|---|---|
| 1–100 | 1 | $0.003 |
| 101–200 | 2 | $0.006 |
| 201–300 | 3 | $0.009 |
| 500 | 5 | $0.015 |
Not charged: scans without text (no_text_layer), password-protected files, files that are not PDFs, files over the size limit, failed downloads, broken PDFs, duplicate URLs. The Actor respects your spending limit: it stops cleanly before a PDF it could not pay for, and never delivers a result you have not paid for or charges for one you do not receive.
Safety and privacy
- PDFs are parsed with pdf.js with script evaluation disabled (
isEvalSupported: false) and a version patched for CVE-2024-4367. Scripts and XFA forms inside the PDF are never run, and no system fonts are used. - Only
httpandhttps. Internal and private network addresses (loopback, private ranges, link-local, cloud metadata endpoints) are refused, also after redirects (max 5). - Downloads stop as soon as they exceed the size limit; the first bytes must be a PDF header; there is a time limit per file.
- Nothing you submit is written to the logs — no URLs, no text, only counts. Results stay in your own dataset.
Limits, stated plainly
- No OCR in this version: scans are detected (
needsOcr,no_text_layer) and not charged, but their text is not recognized. - Text order follows the PDF's internal structure; complex layouts (columns, tables) may come out in reading order that differs from the visual one.
- Only publicly reachable URLs: files behind a login are not supported.
Licenses
pdf.js (pdfjs-dist) — Apache License 2.0. Apify SDK — Apache License 2.0.