# PDF Text Extractor — Text, Pages & Metadata, Scan Detection (`sanmarino-tools/pdf-text-extractor`) Actor

Extract text and metadata from PDF files by URL, whole or page by page. Detects scanned PDFs that need OCR and does not charge for them. Safe PDF parsing (pdf.js, no script execution), size and time limits, no data in logs.

- **URL**: https://apify.com/sanmarino-tools/pdf-text-extractor.md
- **Developed by:** [San Marino Tools](https://apify.com/sanmarino-tools) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 processed pdfs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF Text Extractor — Text, Pages & Metadata, Scan Detection

Extract the text of PDF files from their URLs: the whole document or page by page, with page count and metadata. Scanned PDFs without a text layer are detected and **not charged**.

### What you get

| Field | Meaning |
|---|---|
| `url` | the URL you gave |
| `status` | `ok`, `no_text_layer`, `encrypted`, `not_pdf`, `too_large`, `download_failed`, `parse_error` |
| `error` | machine-readable detail when `status` is not `ok` (e.g. `http_404`, `timeout`, `password_required`, `blocked_address`), otherwise `null` |
| `pageCount` | pages in the document |
| `pagesExtracted` | pages actually extracted (see *Max pages per PDF*) |
| `charCount` | characters of extracted text |
| `needsOcr` | `true` when at least half of the pages have no text layer (typical of scans) |
| `text` | the text, pages separated by a blank line — or, with *Text page by page* on, `pages: [{ page, text }]` |
| `metadata` | `title`, `author`, `subject`, `creator`, `producer`, `creationDate`, `modDate` (ISO 8601) |

```json
{
  "url": "https://example.com/invoice.pdf",
  "status": "ok",
  "error": null,
  "pageCount": 2,
  "pagesExtracted": 2,
  "charCount": 77,
  "needsOcr": false,
  "text": "Invoice 2026-001\nTotal due: 1,234.56 EUR\n\nPage two\nThank you for your business.",
  "metadata": { "title": "Test invoice", "author": "Test author", "subject": null, "creator": null, "producer": "…", "creationDate": "2026-01-15T10:00:00.000Z", "modDate": null }
}
```

A run summary (counts per status, duplicates, charged events, limits used) is saved in the key-value store as `SUMMARY` — open it from the **Output** tab, **Run summary**.

Very long texts (results over about 8 MB) are stored in the key-value store as a separate record; the result then has `text: null` and `textRecordKey` with the record name.

### Input

| Field | Default | |
|---|---|---|
| `pdfUrls` | — | public `http(s)` links, one per line |
| `outputPerPage` | `false` | page-by-page output |
| `maxPagesPerPdf` | 500 | only the first N pages are extracted and charged (max 2,000) |
| `maxFileSizeMb` | 50 | larger files are not downloaded (max 200) |

### Pricing

Pay per event: **$0.003 per PDF** processed successfully, up to 100 pages. Longer documents: one more event for every started block of 100 pages beyond the first 100.

| Pages extracted | Events | Price |
|---|---|---|
| 1–100 | 1 | $0.003 |
| 101–200 | 2 | $0.006 |
| 201–300 | 3 | $0.009 |
| 500 | 5 | $0.015 |

**Not charged:** scans without text (`no_text_layer`), password-protected files, files that are not PDFs, files over the size limit, failed downloads, broken PDFs, duplicate URLs. The Actor respects your spending limit: it stops cleanly before a PDF it could not pay for, and never delivers a result you have not paid for or charges for one you do not receive.

### Safety and privacy

- PDFs are parsed with **pdf.js** with script evaluation disabled (`isEvalSupported: false`) and a version patched for CVE-2024-4367. Scripts and XFA forms inside the PDF are never run, and no system fonts are used.
- Only `http` and `https`. Internal and private network addresses (loopback, private ranges, link-local, cloud metadata endpoints) are refused, also after redirects (max 5).
- Downloads stop as soon as they exceed the size limit; the first bytes must be a PDF header; there is a time limit per file.
- **Nothing you submit is written to the logs** — no URLs, no text, only counts. Results stay in your own dataset.

### Limits, stated plainly

- **No OCR** in this version: scans are detected (`needsOcr`, `no_text_layer`) and not charged, but their text is not recognized.
- Text order follows the PDF's internal structure; complex layouts (columns, tables) may come out in reading order that differs from the visual one.
- Only publicly reachable URLs: files behind a login are not supported.

### Licenses

pdf.js (`pdfjs-dist`) — Apache License 2.0. Apify SDK — Apache License 2.0.

# Actor input Schema

## `pdfUrls` (type: `array`):

Public http(s) links to PDF files, one per line. Duplicates are processed once and charged once.

## `outputPerPage` (type: `boolean`):

If on, results contain a `pages` array (page number + text) instead of one `text` field.

## `maxPagesPerPdf` (type: `integer`):

Only the first N pages are extracted (and charged). Maximum 2,000.

## `maxFileSizeMb` (type: `integer`):

Larger files are not downloaded (status `too_large`, not charged). Maximum 200.

## Actor input object example

```json
{
  "pdfUrls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "outputPerPage": false,
  "maxPagesPerPdf": 500,
  "maxFileSizeMb": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pdfUrls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("sanmarino-tools/pdf-text-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "pdfUrls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("sanmarino-tools/pdf-text-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pdfUrls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}' |
apify call sanmarino-tools/pdf-text-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sanmarino-tools/pdf-text-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BFQR0rJJKFo8uRfnr/builds/rsL4D5F19ip8hbTME/openapi.json
