PDF Table Extractor - PDF to Excel, CSV & JSON (validated rows) avatar

PDF Table Extractor - PDF to Excel, CSV & JSON (validated rows)

Pricing

from $2.00 / 1,000 accepted rows

Go to Apify Store
PDF Table Extractor - PDF to Excel, CSV & JSON (validated rows)

PDF Table Extractor - PDF to Excel, CSV & JSON (validated rows)

Extract tables from PDF price lists, invoices and order confirmations into clean rows for Excel, CSV or JSON. Maps your columns, reads EU and US numbers, checks totals and duplicates, returns rejected rows with reasons. Pay only for accepted rows.

Pricing

from $2.00 / 1,000 accepted rows

Rating

0.0

(0)

Developer

Jordan Nabbe

Jordan Nabbe

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

PDF Table to Validated Rows (CSV, Excel, JSON)

Turn the table in a supplier PDF (price lists, order confirmations, stock reports, statements) into clean spreadsheet rows that you can import without re-checking them by hand. You describe the columns you expect once; every run finds the table, maps header variants such as Qty, QTY or Aantal to your names, converts European and US numbers and dates, and checks required fields, duplicates and line totals. Rows that pass go to the dataset. Rows that fail are listed with the reason, the page and the row number. You only pay for rows that passed.

Built for operations, purchasing and finance teams and for the people who automate their imports (Make, n8n, Zapier, scripts), when the same kind of PDF arrives every week or month.

Try it in one click

Press Start. The prefilled input contains a real one-page supplier price list PDF. The result shows five accepted rows. It also shows one rejected row: the supplier's line total is wrong (3 × 42.00 was printed as 162.00). A wrapped two-line description is joined back into one cell. Then replace the sample PDF with your own and adjust the columns.

What you get

Each accepted row is one flat dataset item with your column names first, ready for CSV or Excel export:

skudescriptionquantityunit_priceline_totalsourceDocumentsourcePagesourceRow
HC-2032Hose clamp 20-32 mm, stainless500.8442northwind-price-list-2026-1011
PG-0010Pressure gauge 0-10 bar, glycerine filled418.975.6northwind-price-list-2026-1012

Every row also carries extraction: how it was read (pdf-text-layer, csv, …), how many text cells were found, how many fragments were joined, how many wrapped lines were merged and how many cells straddled a column boundary. These are measured facts about the extraction, not a guess at accuracy. Rows with merged or boundary-straddling cells are the ones to spot-check.

The OUTPUT record holds:

  • outcome: SUCCEEDED, PARTIAL (some rows or documents failed), FAILED (nothing usable; the run also ends as failed) or LIMIT_REACHED, plus counts;
  • documents: per document the header mapping, warnings, the detected number and date notation and an extraction report (pages read, repeated headers skipped, lines that could not be assigned to any column, with samples);
  • rejectedRows: every rejected row with its values and error codes (MISSING_REQUIRED_FIELD, INVALID_TYPE, AMBIGUOUS_NUMBER, AMBIGUOUS_DATE, DUPLICATE_KEY, ARITHMETIC_CHECK_FAILED, UNEXPECTED_COLUMN).

How to use it with your own files

  1. Put each PDF in documents: { "id": "march-prices", "base64Document": "<base64>" } or { "id": "march-prices", "url": "https://…/march.pdf" }. Tables you already have can be passed as csv, text, json, html or records.
  2. List the columns you want in targetSchema.columns: name, type (string, number, integer, boolean, date), required and the header spellings your suppliers use in aliases.
  3. Optional checks: uniqueBy (for example the SKU) and arithmeticChecks (for example quantity × unit price = line total, with a tolerance).
  4. Run, then export the dataset as CSV/Excel or read it through the API.

Numbers such as 1,234.56, 1.234,56, € 12,50, 12.50 EUR, (45.00) and 45,00- are read automatically. The notation is decided per document from values that can only be read one way. A value like 1.250 that could mean 1.25 or 1250 is never guessed. It is rejected as AMBIGUOUS_NUMBER unless the document shows the notation elsewhere or you set options.numberFormat to en or eu. Dates work the same way: 25-10-2026, 10/25/2026, 03.10.2026, 3 Oct 2026 and 2026-10-03 all become 2026-10-03. A date like 03/10/2026 is only accepted when the document or options.dateFormat settles the order.

Strict mode (options.failOnAnyRejectedRow: true) delivers nothing unless every row passes, for imports that must be all-or-nothing.

Automate it

Call the Actor from Make, n8n, Zapier or a script with the Apify endpoint run-sync-get-dataset-items: it waits for the run and returns the accepted rows. Minimal Python (standard library only):

import base64, json, os, urllib.request
body = {"documents": [{"id": "march", "base64Document": base64.b64encode(open("march.pdf", "rb").read()).decode()}],
"targetSchema": {"columns": [{"name": "sku", "type": "string", "required": True, "aliases": ["SKU"]},
{"name": "line_total", "type": "number", "required": True, "aliases": ["Total"]}]}}
req = urllib.request.Request("https://api.apify.com/v2/acts/first_watch~pdf-table-extraction/run-sync-get-dataset-items?maxTotalChargeUsd=1",
data=json.dumps(body).encode(), method="POST",
headers={"Authorization": "Bearer " + os.environ["APIFY_TOKEN"], "Content-Type": "application/json"})
rows = json.load(urllib.request.urlopen(req, timeout=330))

maxTotalChargeUsd caps what one run can cost. Keep the same targetSchema per supplier so the column mapping stays stable when their headers change slightly.

Pricing

Pay per event. The main charge is per accepted table row; everything else listed below is free.

EventPrice (USD)Charged when
Run start (apify-actor-start)$0.005once per run; the default 512 MB memory is one event
Accepted row (apify-default-dataset-item)$0.002per accepted row delivered to the dataset

Not charged:

  • rejected rows with their error codes (in the OUTPUT record)
  • documents that could not be read (corrupt, scanned, encrypted, no matching table)
  • the run summary, header mappings and extraction-quality report
  • rows withheld because strict mode found a rejected row

Examples (default 512 MB memory, one start event):

  • One supplier price list with 40 accepted rows: $0.005 start + 40 × $0.002 = $0.085
  • Monthly batch of 12 PDFs, 600 accepted rows: $0.005 start + 600 × $0.002 = $1.205
  • A scanned PDF that cannot be read: $0.005 start + 0 × $0.002 = $0.005

Spending limit: when you set a maximum cost per run, the Actor stops before it would exceed it, keeps every accepted table row it already delivered and reports limitReached in the OUTPUT record. It never delivers results beyond the limit.

The Store's Pricing tab shows the prices in force. If this section and the Pricing tab ever differ, the Pricing tab applies.

Limits (read these before relying on it)

  • Only PDFs with a text layer. Scanned or photographed PDFs are refused with SCANNED_PDF_REQUIRES_OCR. They are not charged. There is no OCR.
  • One table per document: the header line that best matches your columns. Rows below it are read on every page. Repeated headers on later pages are skipped.
  • Headers must be on one line. Multi-line headers such as "Unit" above "price" are not recognised.
  • Columns are assigned by position under the header. Unusual layouts (merged cells, sub-totals inside the table, two tables side by side) can put a value in the wrong column. The schema checks usually catch this as a rejected row. Review the first runs of a new supplier template.
  • A line with text only in free-text columns directly below a row is treated as a wrapped cell of that row and counted in continuationLines.
  • Password-protected and truncated PDFs are refused.
  • Up to 50 documents per run, 250 pages and 5,000 rows per document, 10 MiB per inline document and 25 MiB per downloaded PDF. Downloads are HTTPS only, public addresses only, at most 3 redirects and 30 seconds each.

FAQ

Is my data stored or logged? Only in your own Apify run storage. Logs contain counts and status codes, never cell values.

What happens if one PDF fails? It is reported in documents with an error code and is not charged. The other documents are still processed.

Why was a row rejected that looks fine? Check rejectedRows[].errors. Common causes are an ambiguous number or date (set numberFormat or dateFormat), or a column the schema does not list (add an alias or set allowUnexpectedColumns).

Can I keep extra columns? Yes. Set targetSchema.allowUnexpectedColumns to true and they are passed through under their header text.