# PDF Text, Metadata & Tables Extractor (`springlike_meadowland/digital-pdf-extractor`) Actor

Extract page text, optional tables and PDF metadata from public digital PDF URLs. Get one page per row, document hashes, blank-page markers and coverage summaries.

- **URL**: https://apify.com/springlike\_meadowland/digital-pdf-extractor.md
- **Developed by:** [Akshay Aggarwal](https://apify.com/springlike_meadowland) (community)
- **Categories:** Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 text pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF Text, Metadata & Tables Extractor

Turn public digital PDFs into one dataset row per page for document research, indexing, and downstream automation. Each row has extracted text, optional tables, document metadata, a SHA-256 hash, and a clear marker for pages without a digital text layer. The `OUTPUT` record reports coverage and errors for every document.

### Quick start

```json
{
  "pdfUrls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"],
  "maxPagesPerDocument": 20,
  "maxFileSizeMb": 10,
  "includeTables": false
}
```

For a compact output example, a three-page sample PDF containing a ruled table and one blank page produces these rows:

```json
[
  {"page_number": 1, "text": "Alpha report\nName Value\nBeta 7", "tables": [[ ["Name", "Value"], ["Beta", "7"] ]], "text_layer_present": true, "page_status": "text"},
  {"page_number": 2, "text": "", "tables": [], "text_layer_present": false, "page_status": "no_text_layer"}
]
```

The full rows also include `document_url`, `document_hash`, metadata and truncation flags. This example uses a generated sample PDF; the quick-start URL above is a public W3C PDF.

### Output and billing

The default dataset contains **one saved page per row**: `document_url` (final redirect URL), `document_hash` (SHA-256 of source bytes), `page_number` (1-based), `text`, `tables`, `metadata`, `text_layer_present`, `page_status`, `text_truncated`, and `tables_truncated`. Text includes text inside detected tables. Tables are arrays of rows and cell strings; `null` means the parser found an empty cell. Table detection works best on ruled digital tables and can miss complex layouts. When `includeTables` is false, `tables` is an empty array.

`OUTPUT` is a key-value record, separate from the result dataset. It includes totals for URLs requested, documents fetched/skipped/with errors, pages processed/saved, text and blank pages saved, at least how many pages were truncated, and document-level status, errors, and limit reasons. A page limit stops after the chosen number; the Actor detects at least one extra page but does not count all omitted pages. Failed PDFs have no paid result rows. A page parse error is recorded for that page and later pages are attempted.

The billing unit is **one saved page with a digital text layer** (`text-page` event). Blank/image-only page rows have no event, and there is no separate document fee.

### Scope and limits

- Public direct HTTPS PDF URLs only, 1–10 per run. No login, private/local address, IP literal, custom port, or HTTP link. Redirects are checked before connection. Downloads require a PDF content type and `%PDF-` signature, and are streamed with a byte cap.
- `maxPagesPerDocument`: 1–200, default 20. `maxFileSizeMb`: 1–20 MiB, default 10. `includeTables`: default false.
- Extracted text is capped at 100,000 characters per page. Tables are capped at 20 per page, 100 rows per table, 20 columns per row, and 2,000 characters per cell; truncation flags are set in the row. Table discovery itself can be costly on complex pages, so use a small page limit first.
- No OCR, LLM calls, scanned-page transcription, encrypted-PDF unlocking, form filling, JavaScript execution, or external enrichment. A blank/image-only page is emitted with empty text and `no_text_layer` status. Embedded images and attachments are omitted.
- Source hosts may block anonymous cloud requests or change their files. Check `OUTPUT` for partial results and failed URLs; saved dataset rows remain available.

# Actor input Schema

## `pdfUrls` (type: `array`):

1–10 direct public HTTPS PDF URLs. Pages behind a login are unsupported.

## `maxPagesPerDocument` (type: `integer`):

Stop after this many pages per PDF. The output summary marks a document partial if more pages remain.

## `maxFileSizeMb` (type: `integer`):

Files larger than this are rejected before or during download.

## `includeTables` (type: `boolean`):

Find tables drawn in digital PDFs. The best results are on ruled tables; no OCR is performed.

## Actor input object example

```json
{
  "pdfUrls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "maxPagesPerDocument": 20,
  "maxFileSizeMb": 10,
  "includeTables": false
}
```

# Actor output Schema

## `pages` (type: `string`):

One saved page per row, including blank or image-only pages.

## `summary` (type: `string`):

Fetched, saved and skipped counts, partial documents, limits and errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pdfUrls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("springlike_meadowland/digital-pdf-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "pdfUrls": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("springlike_meadowland/digital-pdf-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pdfUrls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}' |
apify call springlike_meadowland/digital-pdf-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,springlike_meadowland/digital-pdf-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fd0QigYbzTbONCRxa/builds/1XrEo9zNlhVKip0ps/openapi.json
