Batch PDF Text Extractor with Page Numbers
Pricing
Pay per usage
Batch PDF Text Extractor with Page Numbers
Extract existing text from batches of public HTTPS PDFs. Get page-by-page text, page numbers, title and author, plus clear results for files with no text or download errors. Up to 25 PDFs per run; 10 MiB and 100 pages per PDF. No OCR.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Software Mechanics
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Convert a batch of publicly accessible HTTPS PDFs into structured text. Each result contains the source URL, page numbers, per-page text, a combined text field, basic metadata, and extraction status.
What it does
- Reads existing text embedded in a PDF.
- Keeps page boundaries so downstream tools can cite the original page.
- Processes repeated identical URLs once per run.
- Reports download and parsing failures per file.
- Flags pages without extractable text. It cannot distinguish blank pages from scans.
- Supports JSON results and Apify's dataset export interface for spreadsheet downloads.
This version does not perform OCR, interpret diagrams, reconstruct tables, bypass logins, or decrypt documents. Multi-column reading order depends on the PDF. It does not send document contents to an external AI service.
Input
{"urls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"]}
Use direct HTTPS links on port 443. The server must allow ordinary downloads. Private network addresses, embedded URL credentials, and compressed HTTP responses are rejected. Input and results are stored in your Apify run; use documents you are authorized to process and follow your account's retention settings.
Limits and partial batches
- At most 25 input URLs per run.
- At most 10 MiB, 100 pages, and 200,000 extracted characters per PDF.
- Parsing runs in a separate process, with a 20-second timeout and a Linux memory ceiling.
- The batch stops starting new files after 240 seconds, or when the customer's result-charge budget is exhausted.
- Check the
SUMMARYkey-value record for counts, the stopping reason, and any unprocessed URLs. A partial batch is not silently reported as complete.
Results
One dataset row is returned per completed file. status is ok, no_text, or error. pages contains {pageNumber, text} objects. fullText contains page markers for easy export. Error rows contain errorCode and errorMessage.
The Output tab exposes PDF results and Run summary. Select Run summary and open the SUMMARY record to see the batch counts and stopping reason. The output schema also makes the dataset and summary collection links available to API and AI-agent consumers. The dataset schema describes and validates the fields in each result, including the null page count on error rows.
An ok result can still have some blank or image-only pages. Those page numbers are listed in pagesWithoutText. Files that exceed a limit are rejected rather than silently truncated.
Pricing behavior
The publisher configures prices in Apify; no active price is embedded in this source package. When pay-per-event monetization is enabled, a pdf-extracted event is charged for each successfully stored ok result. no_text and error results have no extraction event charge. An automatic run-start fee may still apply, as displayed in the Actor's Pricing tab.
The default dataset-item charge must be removed, so it cannot charge errors or double-charge successful results. The code checks this in paid mode and stops if the configuration is inconsistent. The source's private, unmonetized mode produces results without custom event charges.
Local development
Requires Python 3.12. From this directory:
$python -m venv .venv
Activate that environment (.venv\Scripts\activate on Windows, source .venv/bin/activate on macOS/Linux), then:
python -m pip install -r requirements.txtpython -m unittest discover -s tests -v
For a local SDK run, create storage/key_value_stores/default/INPUT.json using example-input.json, then run:
$python -m src.main
To publish, follow ../LAUNCH.md. The initial package is a development build. Local testing does not establish marketplace demand or replace a cloud acceptance run.