PDF to Excel Extractor
Pricing
from $4.80 / 1,000 document extracteds
PDF to Excel Extractor
Extract user-defined fields from repeated-layout text PDFs into XLSX with page provenance and missing-field warnings.
Pricing
from $4.80 / 1,000 document extracteds
Rating
0.0
(0)
Developer
Automation Lab
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
10 days ago
Last modified
Categories
Share
Turn repeated-layout, text-based PDFs into a consistent Excel workbook. Define the fields you need with regular expressions, process up to 50 supplied PDFs, and receive both typed dataset rows and a downloadable XLSX file.
Unlike a generic PDF text dump, this Actor maps only your explicit fields, records the source page and matched text for every field, and warns when required values are missing. It can create a workbook or populate specific columns in your supplied XLSX template.
What does this PDF to Excel Actor do?
The Actor:
- downloads a public PDF URL or decodes an inline base64 PDF;
- reads the existing text layer page by page;
- applies your reusable field mapping to every document;
- converts captured values to strings, numbers, or ISO dates;
- records page-level provenance and validation warnings;
- writes one row per PDF to an Excel workbook;
- stores the workbook as
OUTPUT.xlsxin the run key-value store.
It is designed for repeated forms, invoices, statements, reports, and other PDFs whose labels and layout stay consistent.
Who is it for?
- Finance teams mapping repeated invoice fields into an import workbook
- Operations teams processing recurring forms or reports
- Analysts who need auditable PDF-to-spreadsheet extraction
- Data engineers building scheduled document pipelines
- QA teams checking that required PDF fields are present
Why use schema-guided extraction?
A full-text converter makes downstream users search the text again. This Actor applies a field contract at extraction time. Each result includes normalized values, the page where each match occurred, the exact matched text, and warnings for missing required fields.
You control the mapping. The Actor does not send document content to an AI provider and does not guess unsupported values.
Supported PDFs and limits
- Text-based PDFs with an embedded text layer
- Public HTTP(S) URLs or base64 file content
- Up to 50 PDFs per run
- Up to 20 MB per PDF
- Up to 100 field definitions
- A 30-second download timeout per file
- New workbooks or existing XLSX templates
Scanned image-only PDFs are not OCRed. Convert them with OCR before using this Actor. Password-protected, malformed, or inaccessible PDFs fail the run with a clear error.
Input fields
| Field | Type | Required | Purpose |
|---|---|---|---|
documents | array | Yes | PDFs supplied by url or base64, with an optional name |
fields | array | Yes | Reusable extraction schema |
templateUrl | string | No | Public URL of an XLSX template |
templateBase64 | string | No | Base64-encoded XLSX template |
worksheetName | string | No | Worksheet to update or create |
startRow | integer | No | First row for extracted data |
Each field mapping supports:
| Property | Purpose |
|---|---|
key | Stable key in the dataset values object |
label | Human-readable Excel heading |
pattern | Regular expression; capture group 1 becomes the value |
type | string, number, or date |
required | Adds a warning when no value is found |
page | Searches only this one-based PDF page |
column | Writes to this Excel column when using a template |
Getting started
- Open the Actor input.
- Add one or more text-based PDFs.
- Define a regex for every field you want.
- Mark fields as required when missing data should be visible.
- Optionally provide an XLSX template and target columns.
- Run the Actor.
- Download
OUTPUT.xlsxfrom the Output tab and inspect dataset provenance.
Example input
{"documents": [{"url": "https://pdfobject.com/pdf/sample.pdf","name": "Sample PDF"}],"fields": [{"key": "title","label": "Title","pattern": "^(Sample PDF)","required": true},{"key": "opening","label": "Opening sentence","pattern": "(This is a simple PDF file\\.)","required": true}]}
Example output
{"source": "https://pdfobject.com/pdf/sample.pdf","documentName": "Sample PDF","status": "complete","pageCount": 1,"values": {"title": "Sample PDF","opening": "This is a simple PDF file."},"provenance": [{"field": "title", "page": 1, "matchedText": "Sample PDF"},{"field": "opening", "page": 1, "matchedText": "This is a simple PDF file."}],"warnings": [],"workbookKey": "OUTPUT.xlsx","processedAt": "2026-01-15T12:00:00.000Z"}
There is one dataset item and one workbook row per processed PDF. Dynamic extracted fields live under values; validation details stay under warnings and provenance.
Populate your own Excel template
Provide either templateUrl or templateBase64. Set column on each field mapping, such as B, D, or AA, and choose startRow. The Actor loads the workbook, writes one document per row, preserves other cells, and returns the completed workbook.
If no template is supplied, it creates an Extracted data worksheet with field labels, source, and warnings columns.
Validation and failure behavior
A missing optional field becomes null. A missing required field also adds a warning and changes the document status to warning. This keeps partial but auditable records available.
Invalid input, an unreachable file, a non-PDF response, a PDF larger than 20 MB, or an unreadable workbook template fails the run. Failed documents are not silently converted into successful rows.
Tips for reliable field patterns
- Anchor patterns to nearby labels, for example
Invoice(?: No)?:\s*(\S+). - Use the first capture group for the clean value.
- Add
pagewhen the same label appears on multiple pages. - Escape backslashes in JSON, so
\sbecomes\\s. - Start with one sample PDF, then test several layout variants.
- Keep patterns specific enough to avoid capturing neighboring fields.
- Mark business-critical fields as required.
How much does it cost to extract PDFs into Excel?
Pay-per-event pricing has a $0.005 run start and $0.008 per successfully processed PDF at the BRONZE tier. Spend tiers reduce the per-document price automatically.
| PDFs | BRONZE calculation |
|---|---|
| 1 | One start event plus one document event |
| 10 | One start event plus ten document events |
| 50 | One start event plus fifty document events |
Multiply the active $0.008 document price by the PDF count and add the $0.005 start price. These examples use current Actor event prices. Actual billing can be affected by Apify refunds, fraud controls, disputes, taxes, corrections, or clawbacks. Missing-field warnings do not create a separate charge.
Automation workflows
Recurring invoice batch
Schedule the Actor with the same field schema and a changing document list. Send the workbook to cloud storage or an accounting import step.
Document completeness audit
Mark contract identifiers, dates, and totals as required. Filter dataset items where status is warning for manual follow-up.
Data pipeline ingestion
Read normalized dataset rows through the Apify API while retaining OUTPUT.xlsx for business users.
Template population
Supply a workbook with formulas, formatting, and fixed headers. Map fields to columns and begin writing below the template header.
Use the Apify API
Replace APIFY_TOKEN with your token.
cURL
curl -X POST \"https://api.apify.com/v2/acts/automation-lab~schema-guided-pdf-to-excel-extractor/runs?token=APIFY_TOKEN" \-H "Content-Type: application/json" \-d @input.json
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('automation-lab/schema-guided-pdf-to-excel-extractor').call(input);const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
Python
from apify_client import ApifyClientclient = ApifyClient(token="APIFY_TOKEN")run = client.actor("automation-lab/schema-guided-pdf-to-excel-extractor").call(run_input=input_data)items = client.dataset(run["defaultDatasetId"]).list_items().itemsprint(items)
Use with MCP and AI assistants
Add the Actor to Claude Code:
claude mcp add --transport http apify \"https://mcp.apify.com?tools=automation-lab/schema-guided-pdf-to-excel-extractor"
Claude Desktop setup: add this HTTP server in the desktop MCP configuration. Cursor setup: add the same server URL in Cursor MCP settings. VS Code setup: add it to the MCP servers configuration used by your extension.
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=automation-lab/schema-guided-pdf-to-excel-extractor"}}}
Example prompts:
- “Extract the invoice number, invoice date, and total from these PDF URLs into Excel.”
- “Run this field schema over my monthly reports and show documents with missing required values.”
- “Populate columns B, D, and F in my workbook template from these repeated forms.”
Integrations
Use the dataset and workbook with:
- Google Sheets or Microsoft Excel workflows
- Make and Zapier
- Webhooks and scheduled Apify Tasks
- Python, JavaScript, or BI pipelines
- Cloud storage and accounting import automations
Legality and responsible use
Only process documents you are authorized to access. Do not expose confidential files through public URLs unnecessarily; base64 input can be used for controlled automation. Follow applicable privacy, copyright, retention, and contractual requirements. Review extracted values before using them for financial, legal, medical, or other high-impact decisions.
FAQ
Does it OCR scanned PDFs?
No. The Actor intentionally supports text-layer PDFs only. Image-only pages usually produce missing-field warnings; OCR the document first.
Why is a value null?
The pattern did not match the searched page text, or a numeric capture could not be converted. Inspect provenance, test the regex against extracted text, and add or remove a page restriction.
Can one PDF produce multiple rows?
No. The current contract is one row per document. It is intended for document-level fields, not repeating line-item table extraction.
Can I use my formatted workbook?
Yes. Supply one XLSX template, map fields to columns, and choose the first data row. Existing workbook content is preserved except for cells the Actor writes.
Are warning rows charged?
Yes, when a PDF is successfully parsed and delivered as a useful auditable row. Warnings are part of the result, not a failed run. Downloads or parse failures produce no document result.
Related Actors
- PDF Text Extractor for full page text rather than schema-guided fields
- PDF to Structured Markdown Converter for RAG and publishing workflows
- CSV & Excel Data Quality Cleaner for normalizing the resulting tabular data
Changelog
See the Store changelog for user-visible release notes and supported limitations.