PDF Table to CSV with QA Receipts avatar

PDF Table to CSV with QA Receipts

Pricing

$15.00 / 1,000 page processed through table qas

Go to Apify Store
PDF Table to CSV with QA Receipts

PDF Table to CSV with QA Receipts

Pricing

$15.00 / 1,000 page processed through table qas

Rating

0.0

(0)

Developer

muazah

muazah

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Batch-extract tables from born-digital PDFs and get receipts you can audit, not just a blob of cells.

What it does

  • Pulls up to 5 public HTTPS PDFs per run, 10 selected pages per PDF, 50 pages per run, 100 tables per run.
  • Detects ruled tables (lines), borderless aligned columns (text), or auto, which tries each deterministic strategy once and records which one it used.
  • Every table is separated with page number, table index, and bounding box. Unrelated side-by-side tables are never merged, and tables are not joined across pages unless you opt in (matching headers and column counts only, page provenance preserved).
  • Keeps raw cell strings exactly as printed: leading zeros, decimal commas, quotes, newlines. Blank vs. missing cells stay distinct, and every row keeps its source row index.
  • QA receipts per page and per table: column-count consistency, expected-header match, duplicate headers, empty-cell density, suspected merged cells. These are structural heuristics, not a guarantee that cell values are semantically correct.
  • Optional repeated-header removal, only when you supply expected headers: the first header is kept, later exact repeats are removed, and every removal is recorded as evidence.
  • One UTF-8 quoted CSV per table in the run's key-value store. Numbers and dates stay strings. CSV export is formula-injection guarded by default (a leading apostrophe is added in CSV only); the dataset JSON always keeps the unmodified strings.
  • Optional debug images (up to 10 per run): the real page with detected table and cell boundaries drawn on it, so you can see exactly what was read. Stored privately in your run storage.

Honest limits

  • Born-digital PDFs only. Pages without a usable text layer come back as UNSUPPORTED_OR_INSUFFICIENT_TEXT, uncharged. There is no OCR and no AI in this actor.
  • A valid page with no detected table is processed and billable, labeled NO_TABLE_DETECTED. Failed downloads, invalid input, encrypted/malformed files, and unsupported pages are uncharged.
  • Documents must be publicly reachable HTTPS URLs. Redirects are validated, private-network targets are refused, and signed query strings are never logged or stored.

Input

{
"documents": [{"id": "report-1", "url": "https://example.com/report.pdf", "pages": "1-3"}],
"strategy": "auto",
"expectedHeaders": ["Item", "Quantity", "Amount"],
"removeExactRepeatedHeaders": true,
"debugImages": true
}

Output

Dataset rows with recordType of page_receipt, table_receipt, or table_row, plus per-table CSVs, summary.json, and optional debug PNGs in the run's key-value store. Each row carries the document ID, sanitized source URL, source SHA-256, page, table ID, bounding box, strategy, status, and warning codes.

Billing

Pay per selected page successfully processed through table QA. A 10-page selection costs $0.15. Unsupported, failed, or invalid pages are free.