# PDF to Excel Extractor (`automation-lab/schema-guided-pdf-to-excel-extractor`) Actor

Extract user-defined fields from repeated-layout text PDFs into XLSX with page provenance and missing-field warnings.

- **URL**: https://apify.com/automation-lab/schema-guided-pdf-to-excel-extractor.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.80 / 1,000 document extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF to Excel Extractor

Turn repeated-layout, text-based PDFs into a consistent Excel workbook. Define the fields you need with regular expressions, process up to 50 supplied PDFs, and receive both typed dataset rows and a downloadable XLSX file.

Unlike a generic PDF text dump, this Actor maps only your explicit fields, records the source page and matched text for every field, and warns when required values are missing. It can create a workbook or populate specific columns in your supplied XLSX template.

### What does this PDF to Excel Actor do?

The Actor:

1. downloads a public PDF URL or decodes an inline base64 PDF;
2. reads the existing text layer page by page;
3. applies your reusable field mapping to every document;
4. converts captured values to strings, numbers, or ISO dates;
5. records page-level provenance and validation warnings;
6. writes one row per PDF to an Excel workbook;
7. stores the workbook as `OUTPUT.xlsx` in the run key-value store.

It is designed for repeated forms, invoices, statements, reports, and other PDFs whose labels and layout stay consistent.

### Who is it for?

- Finance teams mapping repeated invoice fields into an import workbook
- Operations teams processing recurring forms or reports
- Analysts who need auditable PDF-to-spreadsheet extraction
- Data engineers building scheduled document pipelines
- QA teams checking that required PDF fields are present

### Why use schema-guided extraction?

A full-text converter makes downstream users search the text again. This Actor applies a field contract at extraction time. Each result includes normalized values, the page where each match occurred, the exact matched text, and warnings for missing required fields.

You control the mapping. The Actor does not send document content to an AI provider and does not guess unsupported values.

### Supported PDFs and limits

- Text-based PDFs with an embedded text layer
- Public HTTP(S) URLs or base64 file content
- Up to 50 PDFs per run
- Up to 20 MB per PDF
- Up to 100 field definitions
- A 30-second download timeout per file
- New workbooks or existing XLSX templates

Scanned image-only PDFs are not OCRed. Convert them with OCR before using this Actor. Password-protected, malformed, or inaccessible PDFs fail the run with a clear error.

### Input fields

| Field | Type | Required | Purpose |
| --- | --- | --- | --- |
| `documents` | array | Yes | PDFs supplied by `url` or `base64`, with an optional `name` |
| `fields` | array | Yes | Reusable extraction schema |
| `templateUrl` | string | No | Public URL of an XLSX template |
| `templateBase64` | string | No | Base64-encoded XLSX template |
| `worksheetName` | string | No | Worksheet to update or create |
| `startRow` | integer | No | First row for extracted data |

Each field mapping supports:

| Property | Purpose |
| --- | --- |
| `key` | Stable key in the dataset `values` object |
| `label` | Human-readable Excel heading |
| `pattern` | Regular expression; capture group 1 becomes the value |
| `type` | `string`, `number`, or `date` |
| `required` | Adds a warning when no value is found |
| `page` | Searches only this one-based PDF page |
| `column` | Writes to this Excel column when using a template |

### Getting started

1. Open the Actor input.
2. Add one or more text-based PDFs.
3. Define a regex for every field you want.
4. Mark fields as required when missing data should be visible.
5. Optionally provide an XLSX template and target columns.
6. Run the Actor.
7. Download `OUTPUT.xlsx` from the Output tab and inspect dataset provenance.

### Example input

```json
{
  "documents": [
    {
      "url": "https://pdfobject.com/pdf/sample.pdf",
      "name": "Sample PDF"
    }
  ],
  "fields": [
    {
      "key": "title",
      "label": "Title",
      "pattern": "^(Sample PDF)",
      "required": true
    },
    {
      "key": "opening",
      "label": "Opening sentence",
      "pattern": "(This is a simple PDF file\\.)",
      "required": true
    }
  ]
}
```

### Example output

```json
{
  "source": "https://pdfobject.com/pdf/sample.pdf",
  "documentName": "Sample PDF",
  "status": "complete",
  "pageCount": 1,
  "values": {
    "title": "Sample PDF",
    "opening": "This is a simple PDF file."
  },
  "provenance": [
    {"field": "title", "page": 1, "matchedText": "Sample PDF"},
    {"field": "opening", "page": 1, "matchedText": "This is a simple PDF file."}
  ],
  "warnings": [],
  "workbookKey": "OUTPUT.xlsx",
  "processedAt": "2026-01-15T12:00:00.000Z"
}
```

There is one dataset item and one workbook row per processed PDF. Dynamic extracted fields live under `values`; validation details stay under `warnings` and `provenance`.

### Populate your own Excel template

Provide either `templateUrl` or `templateBase64`. Set `column` on each field mapping, such as `B`, `D`, or `AA`, and choose `startRow`. The Actor loads the workbook, writes one document per row, preserves other cells, and returns the completed workbook.

If no template is supplied, it creates an `Extracted data` worksheet with field labels, source, and warnings columns.

### Validation and failure behavior

A missing optional field becomes `null`. A missing required field also adds a warning and changes the document status to `warning`. This keeps partial but auditable records available.

Invalid input, an unreachable file, a non-PDF response, a PDF larger than 20 MB, or an unreadable workbook template fails the run. Failed documents are not silently converted into successful rows.

### Tips for reliable field patterns

- Anchor patterns to nearby labels, for example `Invoice(?: No)?:\s*(\S+)`.
- Use the first capture group for the clean value.
- Add `page` when the same label appears on multiple pages.
- Escape backslashes in JSON, so `\s` becomes `\\s`.
- Start with one sample PDF, then test several layout variants.
- Keep patterns specific enough to avoid capturing neighboring fields.
- Mark business-critical fields as required.

### How much does it cost to extract PDFs into Excel?

Pay-per-event pricing has a **$0.005 run start** and **$0.008 per successfully processed PDF** at the BRONZE tier. Spend tiers reduce the per-document price automatically.

| PDFs | BRONZE calculation |
| ---: | --- |
| 1 | One start event plus one document event |
| 10 | One start event plus ten document events |
| 50 | One start event plus fifty document events |

Multiply the active $0.008 document price by the PDF count and add the $0.005 start price. These examples use current Actor event prices. Actual billing can be affected by Apify refunds, fraud controls, disputes, taxes, corrections, or clawbacks. Missing-field warnings do not create a separate charge.

### Automation workflows

#### Recurring invoice batch

Schedule the Actor with the same field schema and a changing document list. Send the workbook to cloud storage or an accounting import step.

#### Document completeness audit

Mark contract identifiers, dates, and totals as required. Filter dataset items where `status` is `warning` for manual follow-up.

#### Data pipeline ingestion

Read normalized dataset rows through the Apify API while retaining `OUTPUT.xlsx` for business users.

#### Template population

Supply a workbook with formulas, formatting, and fixed headers. Map fields to columns and begin writing below the template header.

### Use the Apify API

Replace `APIFY_TOKEN` with your token.

#### cURL

```bash
curl -X POST \
  "https://api.apify.com/v2/acts/automation-lab~schema-guided-pdf-to-excel-extractor/runs?token=APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d @input.json
```

#### JavaScript

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/schema-guided-pdf-to-excel-extractor').call(input);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

#### Python

```python
from apify_client import ApifyClient

client = ApifyClient(token="APIFY_TOKEN")
run = client.actor("automation-lab/schema-guided-pdf-to-excel-extractor").call(run_input=input_data)
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(items)
```

### Use with MCP and AI assistants

Add the Actor to Claude Code:

```bash
claude mcp add --transport http apify \
  "https://mcp.apify.com?tools=automation-lab/schema-guided-pdf-to-excel-extractor"
```

**Claude Desktop setup:** add this HTTP server in the desktop MCP configuration. **Cursor setup:** add the same server URL in Cursor MCP settings. **VS Code setup:** add it to the MCP servers configuration used by your extension.

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=automation-lab/schema-guided-pdf-to-excel-extractor"
    }
  }
}
```

Example prompts:

- “Extract the invoice number, invoice date, and total from these PDF URLs into Excel.”
- “Run this field schema over my monthly reports and show documents with missing required values.”
- “Populate columns B, D, and F in my workbook template from these repeated forms.”

### Integrations

Use the dataset and workbook with:

- Google Sheets or Microsoft Excel workflows
- Make and Zapier
- Webhooks and scheduled Apify Tasks
- Python, JavaScript, or BI pipelines
- Cloud storage and accounting import automations

### Legality and responsible use

Only process documents you are authorized to access. Do not expose confidential files through public URLs unnecessarily; base64 input can be used for controlled automation. Follow applicable privacy, copyright, retention, and contractual requirements. Review extracted values before using them for financial, legal, medical, or other high-impact decisions.

### FAQ

#### Does it OCR scanned PDFs?

No. The Actor intentionally supports text-layer PDFs only. Image-only pages usually produce missing-field warnings; OCR the document first.

#### Why is a value null?

The pattern did not match the searched page text, or a numeric capture could not be converted. Inspect `provenance`, test the regex against extracted text, and add or remove a page restriction.

#### Can one PDF produce multiple rows?

No. The current contract is one row per document. It is intended for document-level fields, not repeating line-item table extraction.

#### Can I use my formatted workbook?

Yes. Supply one XLSX template, map fields to columns, and choose the first data row. Existing workbook content is preserved except for cells the Actor writes.

#### Are warning rows charged?

Yes, when a PDF is successfully parsed and delivered as a useful auditable row. Warnings are part of the result, not a failed run. Downloads or parse failures produce no document result.

### Related Actors

- [PDF Text Extractor](https://apify.com/automation-lab/pdf-text-extractor) for full page text rather than schema-guided fields
- [PDF to Structured Markdown Converter](https://apify.com/automation-lab/pdf-to-structured-markdown-converter) for RAG and publishing workflows
- [CSV & Excel Data Quality Cleaner](https://apify.com/automation-lab/csv-excel-data-quality-cleaner) for normalizing the resulting tabular data

### Changelog

See the Store changelog for user-visible release notes and supported limitations.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/schema-guided-pdf-to-excel-extractor/changelog.md

# Actor input Schema

## `documents` (type: `array`):

One to 50 text-based PDF files. Provide exactly one public URL or base64 payload per document.

## `fields` (type: `array`):

Regex mappings. The first capture group becomes the value; otherwise the full match is used.

## `templateUrl` (type: `string`):

Optional public URL of an XLSX template. Use field column mappings to place values.

## `templateBase64` (type: `string`):

Optional base64-encoded XLSX template. Do not combine with templateUrl.

## `worksheetName` (type: `string`):

Worksheet to update or create.

## `startRow` (type: `integer`):

Optional one-based row where extracted records begin.

## Actor input object example

```json
{
  "documents": [
    {
      "url": "https://pdfobject.com/pdf/sample.pdf",
      "name": "sample.pdf"
    }
  ],
  "fields": [
    {
      "key": "title",
      "label": "Document title",
      "pattern": "^(Sample PDF)",
      "type": "string",
      "required": true
    },
    {
      "key": "opening",
      "label": "Opening sentence",
      "pattern": "(This is a simple PDF file\\.)",
      "type": "string",
      "required": true
    }
  ],
  "worksheetName": "Extracted data",
  "startRow": 2
}
```

# Actor output Schema

## `dataset` (type: `string`):

One record per processed PDF with values, provenance, and warnings.

## `workbook` (type: `string`):

Generated XLSX workbook ready to download.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/schema-guided-pdf-to-excel-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/schema-guided-pdf-to-excel-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call automation-lab/schema-guided-pdf-to-excel-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/schema-guided-pdf-to-excel-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Ssdgb2arAid5POgzG/builds/AMYvQ7aXmVeH12toh/openapi.json
