PDF to Excel Extractor avatar

PDF to Excel Extractor

Pricing

from $4.80 / 1,000 document extracteds

Go to Apify Store
PDF to Excel Extractor

PDF to Excel Extractor

Extract user-defined fields from repeated-layout text PDFs into XLSX with page provenance and missing-field warnings.

Pricing

from $4.80 / 1,000 document extracteds

Rating

0.0

(0)

Developer

Automation Lab

Automation Lab

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

10 days ago

Last modified

Categories

Share

Turn repeated-layout, text-based PDFs into a consistent Excel workbook. Define the fields you need with regular expressions, process up to 50 supplied PDFs, and receive both typed dataset rows and a downloadable XLSX file.

Unlike a generic PDF text dump, this Actor maps only your explicit fields, records the source page and matched text for every field, and warns when required values are missing. It can create a workbook or populate specific columns in your supplied XLSX template.

What does this PDF to Excel Actor do?

The Actor:

  1. downloads a public PDF URL or decodes an inline base64 PDF;
  2. reads the existing text layer page by page;
  3. applies your reusable field mapping to every document;
  4. converts captured values to strings, numbers, or ISO dates;
  5. records page-level provenance and validation warnings;
  6. writes one row per PDF to an Excel workbook;
  7. stores the workbook as OUTPUT.xlsx in the run key-value store.

It is designed for repeated forms, invoices, statements, reports, and other PDFs whose labels and layout stay consistent.

Who is it for?

  • Finance teams mapping repeated invoice fields into an import workbook
  • Operations teams processing recurring forms or reports
  • Analysts who need auditable PDF-to-spreadsheet extraction
  • Data engineers building scheduled document pipelines
  • QA teams checking that required PDF fields are present

Why use schema-guided extraction?

A full-text converter makes downstream users search the text again. This Actor applies a field contract at extraction time. Each result includes normalized values, the page where each match occurred, the exact matched text, and warnings for missing required fields.

You control the mapping. The Actor does not send document content to an AI provider and does not guess unsupported values.

Supported PDFs and limits

  • Text-based PDFs with an embedded text layer
  • Public HTTP(S) URLs or base64 file content
  • Up to 50 PDFs per run
  • Up to 20 MB per PDF
  • Up to 100 field definitions
  • A 30-second download timeout per file
  • New workbooks or existing XLSX templates

Scanned image-only PDFs are not OCRed. Convert them with OCR before using this Actor. Password-protected, malformed, or inaccessible PDFs fail the run with a clear error.

Input fields

FieldTypeRequiredPurpose
documentsarrayYesPDFs supplied by url or base64, with an optional name
fieldsarrayYesReusable extraction schema
templateUrlstringNoPublic URL of an XLSX template
templateBase64stringNoBase64-encoded XLSX template
worksheetNamestringNoWorksheet to update or create
startRowintegerNoFirst row for extracted data

Each field mapping supports:

PropertyPurpose
keyStable key in the dataset values object
labelHuman-readable Excel heading
patternRegular expression; capture group 1 becomes the value
typestring, number, or date
requiredAdds a warning when no value is found
pageSearches only this one-based PDF page
columnWrites to this Excel column when using a template

Getting started

  1. Open the Actor input.
  2. Add one or more text-based PDFs.
  3. Define a regex for every field you want.
  4. Mark fields as required when missing data should be visible.
  5. Optionally provide an XLSX template and target columns.
  6. Run the Actor.
  7. Download OUTPUT.xlsx from the Output tab and inspect dataset provenance.

Example input

{
"documents": [
{
"url": "https://pdfobject.com/pdf/sample.pdf",
"name": "Sample PDF"
}
],
"fields": [
{
"key": "title",
"label": "Title",
"pattern": "^(Sample PDF)",
"required": true
},
{
"key": "opening",
"label": "Opening sentence",
"pattern": "(This is a simple PDF file\\.)",
"required": true
}
]
}

Example output

{
"source": "https://pdfobject.com/pdf/sample.pdf",
"documentName": "Sample PDF",
"status": "complete",
"pageCount": 1,
"values": {
"title": "Sample PDF",
"opening": "This is a simple PDF file."
},
"provenance": [
{"field": "title", "page": 1, "matchedText": "Sample PDF"},
{"field": "opening", "page": 1, "matchedText": "This is a simple PDF file."}
],
"warnings": [],
"workbookKey": "OUTPUT.xlsx",
"processedAt": "2026-01-15T12:00:00.000Z"
}

There is one dataset item and one workbook row per processed PDF. Dynamic extracted fields live under values; validation details stay under warnings and provenance.

Populate your own Excel template

Provide either templateUrl or templateBase64. Set column on each field mapping, such as B, D, or AA, and choose startRow. The Actor loads the workbook, writes one document per row, preserves other cells, and returns the completed workbook.

If no template is supplied, it creates an Extracted data worksheet with field labels, source, and warnings columns.

Validation and failure behavior

A missing optional field becomes null. A missing required field also adds a warning and changes the document status to warning. This keeps partial but auditable records available.

Invalid input, an unreachable file, a non-PDF response, a PDF larger than 20 MB, or an unreadable workbook template fails the run. Failed documents are not silently converted into successful rows.

Tips for reliable field patterns

  • Anchor patterns to nearby labels, for example Invoice(?: No)?:\s*(\S+).
  • Use the first capture group for the clean value.
  • Add page when the same label appears on multiple pages.
  • Escape backslashes in JSON, so \s becomes \\s.
  • Start with one sample PDF, then test several layout variants.
  • Keep patterns specific enough to avoid capturing neighboring fields.
  • Mark business-critical fields as required.

How much does it cost to extract PDFs into Excel?

Pay-per-event pricing has a $0.005 run start and $0.008 per successfully processed PDF at the BRONZE tier. Spend tiers reduce the per-document price automatically.

PDFsBRONZE calculation
1One start event plus one document event
10One start event plus ten document events
50One start event plus fifty document events

Multiply the active $0.008 document price by the PDF count and add the $0.005 start price. These examples use current Actor event prices. Actual billing can be affected by Apify refunds, fraud controls, disputes, taxes, corrections, or clawbacks. Missing-field warnings do not create a separate charge.

Automation workflows

Recurring invoice batch

Schedule the Actor with the same field schema and a changing document list. Send the workbook to cloud storage or an accounting import step.

Document completeness audit

Mark contract identifiers, dates, and totals as required. Filter dataset items where status is warning for manual follow-up.

Data pipeline ingestion

Read normalized dataset rows through the Apify API while retaining OUTPUT.xlsx for business users.

Template population

Supply a workbook with formulas, formatting, and fixed headers. Map fields to columns and begin writing below the template header.

Use the Apify API

Replace APIFY_TOKEN with your token.

cURL

curl -X POST \
"https://api.apify.com/v2/acts/automation-lab~schema-guided-pdf-to-excel-extractor/runs?token=APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d @input.json

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/schema-guided-pdf-to-excel-extractor').call(input);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Python

from apify_client import ApifyClient
client = ApifyClient(token="APIFY_TOKEN")
run = client.actor("automation-lab/schema-guided-pdf-to-excel-extractor").call(run_input=input_data)
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(items)

Use with MCP and AI assistants

Add the Actor to Claude Code:

claude mcp add --transport http apify \
"https://mcp.apify.com?tools=automation-lab/schema-guided-pdf-to-excel-extractor"

Claude Desktop setup: add this HTTP server in the desktop MCP configuration. Cursor setup: add the same server URL in Cursor MCP settings. VS Code setup: add it to the MCP servers configuration used by your extension.

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=automation-lab/schema-guided-pdf-to-excel-extractor"
}
}
}

Example prompts:

  • “Extract the invoice number, invoice date, and total from these PDF URLs into Excel.”
  • “Run this field schema over my monthly reports and show documents with missing required values.”
  • “Populate columns B, D, and F in my workbook template from these repeated forms.”

Integrations

Use the dataset and workbook with:

  • Google Sheets or Microsoft Excel workflows
  • Make and Zapier
  • Webhooks and scheduled Apify Tasks
  • Python, JavaScript, or BI pipelines
  • Cloud storage and accounting import automations

Legality and responsible use

Only process documents you are authorized to access. Do not expose confidential files through public URLs unnecessarily; base64 input can be used for controlled automation. Follow applicable privacy, copyright, retention, and contractual requirements. Review extracted values before using them for financial, legal, medical, or other high-impact decisions.

FAQ

Does it OCR scanned PDFs?

No. The Actor intentionally supports text-layer PDFs only. Image-only pages usually produce missing-field warnings; OCR the document first.

Why is a value null?

The pattern did not match the searched page text, or a numeric capture could not be converted. Inspect provenance, test the regex against extracted text, and add or remove a page restriction.

Can one PDF produce multiple rows?

No. The current contract is one row per document. It is intended for document-level fields, not repeating line-item table extraction.

Can I use my formatted workbook?

Yes. Supply one XLSX template, map fields to columns, and choose the first data row. Existing workbook content is preserved except for cells the Actor writes.

Are warning rows charged?

Yes, when a PDF is successfully parsed and delivered as a useful auditable row. Warnings are part of the result, not a failed run. Downloads or parse failures produce no document result.

Changelog

See the Store changelog for user-visible release notes and supported limitations.