# Batch PDF Text Extractor with Page Numbers (`software_mechanics/page-aware-pdf-text`) Actor

Extract existing text from batches of public HTTPS PDFs. Get page-by-page text, page numbers, title and author, plus clear results for files with no text or download errors. Up to 25 PDFs per run; 10 MiB and 100 pages per PDF. No OCR.

- **URL**: https://apify.com/software\_mechanics/page-aware-pdf-text.md
- **Developed by:** [Software Mechanics](https://apify.com/software_mechanics) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Batch PDF Text Extractor with Page Numbers

Convert a batch of publicly accessible HTTPS PDFs into structured text. Each result contains the source URL, page numbers, per-page text, a combined text field, basic metadata, and extraction status.

### What it does

- Reads existing text embedded in a PDF.
- Keeps page boundaries so downstream tools can cite the original page.
- Processes repeated identical URLs once per run.
- Reports download and parsing failures per file.
- Flags pages without extractable text. It cannot distinguish blank pages from scans.
- Supports JSON results and Apify's dataset export interface for spreadsheet downloads.

This version does not perform OCR, interpret diagrams, reconstruct tables, bypass logins, or decrypt documents. Multi-column reading order depends on the PDF. It does not send document contents to an external AI service.

### Input

```json
{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}
```

Use direct HTTPS links on port 443. The server must allow ordinary downloads. Private network addresses, embedded URL credentials, and compressed HTTP responses are rejected. Input and results are stored in your Apify run; use documents you are authorized to process and follow your account's retention settings.

### Limits and partial batches

- At most 25 input URLs per run.
- At most 10 MiB, 100 pages, and 200,000 extracted characters per PDF.
- Parsing runs in a separate process, with a 20-second timeout and a Linux memory ceiling.
- The batch stops starting new files after 240 seconds, or when the customer's result-charge budget is exhausted.
- Check the `SUMMARY` key-value record for counts, the stopping reason, and any unprocessed URLs. A partial batch is not silently reported as complete.

### Results

One dataset row is returned per completed file. `status` is `ok`, `no_text`, or `error`. `pages` contains `{pageNumber, text}` objects. `fullText` contains page markers for easy export. Error rows contain `errorCode` and `errorMessage`.

The Output tab exposes **PDF results** and **Run summary**. Select Run summary and open the `SUMMARY` record to see the batch counts and stopping reason. The output schema also makes the dataset and summary collection links available to API and AI-agent consumers. The dataset schema describes and validates the fields in each result, including the `null` page count on error rows.

An `ok` result can still have some blank or image-only pages. Those page numbers are listed in `pagesWithoutText`. Files that exceed a limit are rejected rather than silently truncated.

### Pricing behavior

The publisher configures prices in Apify; no active price is embedded in this source package. When pay-per-event monetization is enabled, a `pdf-extracted` event is charged for each successfully stored `ok` result. `no_text` and error results have no extraction event charge. An automatic run-start fee may still apply, as displayed in the Actor's Pricing tab.

The default dataset-item charge must be removed, so it cannot charge errors or double-charge successful results. The code checks this in paid mode and stops if the configuration is inconsistent. The source's private, unmonetized mode produces results without custom event charges.

### Local development

Requires Python 3.12. From this directory:

```bash
python -m venv .venv
```

Activate that environment (`.venv\Scripts\activate` on Windows, `source .venv/bin/activate` on macOS/Linux), then:

```bash
python -m pip install -r requirements.txt
python -m unittest discover -s tests -v
```

For a local SDK run, create `storage/key_value_stores/default/INPUT.json` using `example-input.json`, then run:

```bash
python -m src.main
```

To publish, follow `../LAUNCH.md`. The initial package is a development build. Local testing does not establish marketplace demand or replace a cloud acceptance run.

# Actor input Schema

## `urls` (type: `array`):

1–25 public HTTPS PDF URLs. Each file must be at most 10 MiB and 100 pages. Identical URLs are processed once. Scanned PDFs need a separate OCR tool.

## Actor input object example

```json
{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

One dataset row per completed unique URL. Includes extraction status, page-by-page text, full text with page numbers, basic metadata, warnings, and errors. A completed run can contain error or no\_text rows.

## `summary` (type: `string`):

Open the SUMMARY JSON record for uniqueUrls, processed, successful, noText, errors, stoppedReason, remainingUrls, and monetizationEnabled. Check remainingUrls and stoppedReason before treating a batch as complete.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("software_mechanics/page-aware-pdf-text").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("software_mechanics/page-aware-pdf-text").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}' |
apify call software_mechanics/page-aware-pdf-text --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,software_mechanics/page-aware-pdf-text"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7rqCOLQM8cku6lSZy/builds/PwQzWW1cremcLYCIU/openapi.json
