PDF Text Extractor — Text & Metadata from URLs
Pricing
from $5.00 / 1,000 pdf extracteds
PDF Text Extractor — Text & Metadata from URLs
Extract clean text and metadata from any PDF by URL: full text, page count, title, author, dates as JSON. Perfect for AI pipelines, RAG ingestion, document search and content analysis. No API key needed.
Pricing
from $5.00 / 1,000 pdf extracteds
Rating
0.0
(0)
Developer
Oaida Adrian
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
8 hours ago
Last modified
Categories
Share
PDF Text Extractor — Text & Metadata from Any PDF URL
Extract clean plain text and full document metadata from any PDF — just paste the URLs. No API keys, no login, no proxies, no OCR setup. Point it at a list of PDF links and get back structured, LLM-ready text plus title/author/date/page metadata for every file.
Built for developers and data teams who need to turn scattered PDF documents — research papers, financial reports, whitepapers, manuals, invoices, legal filings — into clean structured data at scale.
Why this Actor
- Zero configuration — no credentials, no proxy budget, no browser. Paste URLs, press run.
- RAG-ready output — optional
includePageTextsmode returns per-page text so you can chunk documents for embeddings without re-splitting. - Resilient by design — unreadable or password-protected files are skipped and logged, never crashing the run. A batch of 500 PDFs with 3 bad links still returns 497 clean items.
- Handles big documents — multi-hundred-page files up to 100 MB each, including "secured" PDFs encrypted with an empty password.
How it works
Give it a list of direct PDF URLs. For each file the Actor downloads it, parses it with a battle-tested PDF engine, and outputs one dataset item per PDF with the full extracted text and all document metadata. Set maxPagesPerPdf to cap long files, or includePageTexts: true to also get an array with each page's text separately.
Input
{"pdfUrls": ["https://arxiv.org/pdf/1706.03762","https://bitcoin.org/bitcoin.pdf"],"maxPagesPerPdf": 0,"includePageTexts": false}
| Field | Type | Description |
|---|---|---|
pdfUrls | array | Direct URLs of PDF files (required) |
maxPagesPerPdf | integer | Max pages to extract per PDF; 0 = all pages |
includePageTexts | boolean | Also output a per-page text array — ideal for RAG chunking |
Output
One item per PDF:
{"url": "https://arxiv.org/pdf/1706.03762","fileName": "1706.03762","fileSizeBytes": 2215244,"numPages": 15,"pagesExtracted": 15,"text": "Attention Is All You Need ...","textLength": 39432,"wordCount": 6120,"title": "Attention Is All You Need","author": null,"creator": "LaTeX with hyperref","producer": "pdfTeX-1.40.25","creationDate": "D:20240410010203Z","modDate": null,"encrypted": false,"pageTexts": null}
With includePageTexts: true, pageTexts becomes ["page 1 text...", "page 2 text...", ...].
Use cases
- 🤖 AI / RAG pipelines — turn PDFs into LLM-ready text; use
pageTextsas natural chunk boundaries for embeddings. - 🔎 Document search & indexing — full text plus title/author/date metadata to power search over a PDF corpus.
- 📊 Content & compliance analysis — word counts, page counts, and document properties across large batches.
- 📚 Bulk research ingestion — pull hundreds of arXiv / SSRN papers or vendor whitepapers in one run.
- 🧾 Report & invoice processing — extract the text layer from financial reports, statements, and structured invoices.
Run it from the API
Trigger the Actor and get results in one call:
curl -X POST "https://api.apify.com/v2/acts/darknezz~pdf-text-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"pdfUrls":["https://bitcoin.org/bitcoin.pdf"],"includePageTexts":true}'
Or from Python with the Apify SDK:
from apify_client import ApifyClientclient = ApifyClient("YOUR-APIFY-TOKEN")run = client.actor("darknezz/pdf-text-extractor").call(run_input={"pdfUrls": ["https://arxiv.org/pdf/1706.03762"], "includePageTexts": True})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["title"], item["wordCount"], "words", "-", item["url"])
Scheduling: attach an Apify Schedule (e.g. hourly) to keep a folder of report URLs continuously ingested into a dataset that feeds your pipeline.
Pricing
Going FREE on 24 Aug 2026. Until then it runs on a small pay-per-event fee (per PDF extracted, plus Apify's standard actor-start and dataset-item micro-fees — roughly $5 per 1,000 PDFs); after the scheduled change this actor is a free tool on the Apify Store — you only pay Apify's standard platform costs for running it, not per file. No monthly minimum.
FAQ
Does it OCR scanned PDFs? No — it extracts the embedded text layer. Scanned image-only PDFs return empty text (the item still includes metadata and page counts, so you can detect and route them to an OCR step).
Password-protected PDFs? Files encrypted with an empty password (common "secured" PDFs) are extracted automatically. User-password-protected files are skipped with a warning rather than failing the run.
What's the file size limit? 100 MB per file.
Can I limit pages for very long documents? Yes — set maxPagesPerPdf to any number; 0 extracts everything.
Do I need a proxy? No. The Actor downloads PDFs directly over HTTPS — no proxy or residential IP budget required.
What if one URL is broken? That file is skipped and logged; every other PDF in the batch still returns normally.
Which engines are supported? Any server-rendered PDF served over HTTPS — arXiv, government filings, corporate reports, manuals. If a URL redirects to a login wall or an HTML page, it's logged as a failed fetch and skipped.
Limitations
- Only the embedded text layer is extracted — scanned (image-only) PDFs yield empty
textand need an OCR step (see FAQ). - Files behind interactive JS viewers (e.g. some document portals) aren't supported — use the direct
.pdfURL. - Extraction quality depends on the PDF's own text layer; poorly generated files may have garbled ordering, which is a property of the source document, not the extractor.
Enjoying the extractor? A quick review on the Apify Store helps others find it.