PDF Text Extractor — Text, Pages & Metadata, Scan Detection avatar

PDF Text Extractor — Text, Pages & Metadata, Scan Detection

Pricing

from $3.00 / 1,000 processed pdfs

Go to Apify Store
PDF Text Extractor — Text, Pages & Metadata, Scan Detection

PDF Text Extractor — Text, Pages & Metadata, Scan Detection

Extract text and metadata from PDF files by URL, whole or page by page. Detects scanned PDFs that need OCR and does not charge for them. Safe PDF parsing (pdf.js, no script execution), size and time limits, no data in logs.

Pricing

from $3.00 / 1,000 processed pdfs

Rating

0.0

(0)

Developer

San Marino Tools

San Marino Tools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Extract the text of PDF files from their URLs: the whole document or page by page, with page count and metadata. Scanned PDFs without a text layer are detected and not charged.

What you get

FieldMeaning
urlthe URL you gave
statusok, no_text_layer, encrypted, not_pdf, too_large, download_failed, parse_error
errormachine-readable detail when status is not ok (e.g. http_404, timeout, password_required, blocked_address), otherwise null
pageCountpages in the document
pagesExtractedpages actually extracted (see Max pages per PDF)
charCountcharacters of extracted text
needsOcrtrue when at least half of the pages have no text layer (typical of scans)
textthe text, pages separated by a blank line — or, with Text page by page on, pages: [{ page, text }]
metadatatitle, author, subject, creator, producer, creationDate, modDate (ISO 8601)
{
"url": "https://example.com/invoice.pdf",
"status": "ok",
"error": null,
"pageCount": 2,
"pagesExtracted": 2,
"charCount": 77,
"needsOcr": false,
"text": "Invoice 2026-001\nTotal due: 1,234.56 EUR\n\nPage two\nThank you for your business.",
"metadata": { "title": "Test invoice", "author": "Test author", "subject": null, "creator": null, "producer": "…", "creationDate": "2026-01-15T10:00:00.000Z", "modDate": null }
}

A run summary (counts per status, duplicates, charged events, limits used) is saved in the key-value store as SUMMARY — open it from the Output tab, Run summary.

Very long texts (results over about 8 MB) are stored in the key-value store as a separate record; the result then has text: null and textRecordKey with the record name.

Input

FieldDefault
pdfUrls—public http(s) links, one per line
outputPerPagefalsepage-by-page output
maxPagesPerPdf500only the first N pages are extracted and charged (max 2,000)
maxFileSizeMb50larger files are not downloaded (max 200)

Pricing

Pay per event: $0.003 per PDF processed successfully, up to 100 pages. Longer documents: one more event for every started block of 100 pages beyond the first 100.

Pages extractedEventsPrice
1–1001$0.003
101–2002$0.006
201–3003$0.009
5005$0.015

Not charged: scans without text (no_text_layer), password-protected files, files that are not PDFs, files over the size limit, failed downloads, broken PDFs, duplicate URLs. The Actor respects your spending limit: it stops cleanly before a PDF it could not pay for, and never delivers a result you have not paid for or charges for one you do not receive.

Safety and privacy

  • PDFs are parsed with pdf.js with script evaluation disabled (isEvalSupported: false) and a version patched for CVE-2024-4367. Scripts and XFA forms inside the PDF are never run, and no system fonts are used.
  • Only http and https. Internal and private network addresses (loopback, private ranges, link-local, cloud metadata endpoints) are refused, also after redirects (max 5).
  • Downloads stop as soon as they exceed the size limit; the first bytes must be a PDF header; there is a time limit per file.
  • Nothing you submit is written to the logs — no URLs, no text, only counts. Results stay in your own dataset.

Limits, stated plainly

  • No OCR in this version: scans are detected (needsOcr, no_text_layer) and not charged, but their text is not recognized.
  • Text order follows the PDF's internal structure; complex layouts (columns, tables) may come out in reading order that differs from the visual one.
  • Only publicly reachable URLs: files behind a login are not supported.

Licenses

pdf.js (pdfjs-dist) — Apache License 2.0. Apify SDK — Apache License 2.0.