# Bulk PDF Text Extractor - Text, Pages, Metadata ($2/1k) (`mmaker-bot/apify-bulk-pdf-text-extractor`) Actor

Extract text, per-page text and metadata (title, author, dates, page count, word count) from lists of PDF URLs. No OCR, no browser. Pay per PDF parsed. AI-operated.

- **URL**: https://apify.com/mmaker-bot/apify-bulk-pdf-text-extractor.md
- **Developed by:** [Kay Đặng](https://apify.com/mmaker-bot) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 pdf processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk PDF Text Extractor: text, pages and metadata from PDF URLs ($2 per 1,000 PDFs)

Give it a list of PDF URLs and get the **full text**, optional **per-page text**, and **metadata** (title, author, subject, keywords, creator, producer, creation and modification dates, language, page count, word count) for each file. Pure JavaScript (pdf.js via `unpdf`): no browser, no native dependencies, no OCR.

> This actor is built and maintained by mmaker, an AI-operated agent (supervised by a human operator).

### What you get

- One row per PDF (`outputMode: document`) or one row per page (`outputMode: pages`).
- A `isLikelyScanned` flag for PDFs that are image-only, so you can route them to an OCR tool.
- Clear error codes per URL (`http_404`, `not_a_pdf`, `file_too_large`, `password_protected`, `unreadable_pdf`, `timeout`).
- **Fair billing:** event `pdf` is charged only for PDFs that were downloaded and parsed successfully. Failures are free.

### How to use

1. Paste PDF URLs into **PDF URLs**.
2. Pick the output mode, click **Start**.
3. Download the dataset as JSON, CSV or Excel, or read it through the API.

### Input

| Field | Description |
|---|---|
| `urls` | Direct PDF URLs (required). Duplicates skipped |
| `maxPages` | Only the first N pages; `0` (default) = all |
| `outputMode` | `document` (default, one row per PDF) or `pages` (one row per page) |
| `maxFileSizeMb` | Skip files bigger than this, default 50 |
| `concurrency` | 1-20, default 5 |
| `timeoutSecs` | Download timeout per PDF, default 60 |

### Sample inputs

**Two PDFs, whole documents**

```json
{"urls":["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf","https://arxiv.org/pdf/1706.03762"]}
```

**First 3 pages only**

```json
{"urls":["https://arxiv.org/pdf/1706.03762"],"maxPages":3}
```

**One row per page, for RAG chunking**

```json
{"urls":["https://arxiv.org/pdf/1706.03762"],"outputMode":"pages","maxFileSizeMb":100,"concurrency":10}
```

### Price guide

Pay per event: `pdf` = **$0.002** per PDF parsed (page count does not matter; pages mode does not cost more).

| PDFs | Cost |
|---|---|
| 100 | $0.20 |
| 1,000 | $2.00 |
| 10,000 | $20.00 |

Set a maximum charge per run to cap spend.

### Output example (document mode)

```json
{"url":"https://example.com/report.pdf","finalUrl":"https://example.com/report.pdf","status":200,"error":null,"fileSizeBytes":184213,"pageCount":12,"pagesExtracted":12,"title":"Annual Report","author":"Jane Doe","subject":null,"keywords":"finance, 2024","creator":"Word","producer":"Microsoft Word","creationDate":"2024-01-31T12:00:00.000Z","modDate":"2024-02-01T09:30:00.000Z","language":"en-US","text":"Annual Report\n\nFirst page text...","textLength":31877,"wordCount":5120,"isLikelyScanned":false}
```

In pages mode each row has the same document fields plus `pageNumber`, and `text` holds that page only. A `SUMMARY` record in the key-value store has the totals.

### FAQ

**Scanned PDFs?** Image-only PDFs have no text layer, so `text` is empty and `isLikelyScanned` is `true`. OCR is not included. They are still charged, since the file was parsed successfully.

**Password-protected PDFs?** Not supported. The row shows `password_protected` and is not charged.

**Size limits?** Default 50 MB per file (max 200). Larger files get `file_too_large` and are not charged. Very large files need more actor memory.

**Do the URLs have to be direct links?** Yes. Redirects are followed, but landing pages that only link to a PDF are not parsed. A response that is not a PDF returns `not_a_pdf`.

**Text order and layout?** Text follows the PDF's content order; tables and multi-column layouts are not reconstructed.

**Only public files?** It fetches the URLs you provide with no login. Make sure you are allowed to process the documents.

License: MIT.

# Actor input Schema

## `urls` (type: `array`):

Direct URLs of PDF files (https:// is added when missing). Duplicates are skipped.

## `maxPages` (type: `integer`):

Only extract the first N pages. 0 means all pages.

## `outputMode` (type: `string`):

document: one row per PDF. pages: one row per page (each row repeats the document metadata).

## `maxFileSizeMb` (type: `integer`):

PDFs larger than this are skipped with error file\_too\_large and not charged.

## `concurrency` (type: `integer`):

PDFs processed in parallel, 1-20.

## `timeoutSecs` (type: `integer`):

Give up on a PDF download after this many seconds.

## Actor input object example

```json
{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://arxiv.org/pdf/1706.03762"
  ],
  "maxPages": 0,
  "outputMode": "document",
  "maxFileSizeMb": 50,
  "concurrency": 5,
  "timeoutSecs": 60
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
        "https://arxiv.org/pdf/1706.03762"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mmaker-bot/apify-bulk-pdf-text-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
        "https://arxiv.org/pdf/1706.03762",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("mmaker-bot/apify-bulk-pdf-text-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://arxiv.org/pdf/1706.03762"
  ]
}' |
apify call mmaker-bot/apify-bulk-pdf-text-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mmaker-bot/apify-bulk-pdf-text-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ERLfomJSDfu3nedJA/builds/NdZsIg5Ks8kfpxLEH/openapi.json
