PDF Table to CSV with QA Receipts
Pricing
$15.00 / 1,000 page processed through table qas
PDF Table to CSV with QA Receipts
Pricing
$15.00 / 1,000 page processed through table qas
Rating
0.0
(0)
Developer
muazah
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Batch-extract tables from born-digital PDFs and get receipts you can audit, not just a blob of cells.
What it does
- Pulls up to 5 public HTTPS PDFs per run, 10 selected pages per PDF, 50 pages per run, 100 tables per run.
- Detects ruled tables (
lines), borderless aligned columns (text), orauto, which tries each deterministic strategy once and records which one it used. - Every table is separated with page number, table index, and bounding box. Unrelated side-by-side tables are never merged, and tables are not joined across pages unless you opt in (matching headers and column counts only, page provenance preserved).
- Keeps raw cell strings exactly as printed: leading zeros, decimal commas, quotes, newlines. Blank vs. missing cells stay distinct, and every row keeps its source row index.
- QA receipts per page and per table: column-count consistency, expected-header match, duplicate headers, empty-cell density, suspected merged cells. These are structural heuristics, not a guarantee that cell values are semantically correct.
- Optional repeated-header removal, only when you supply expected headers: the first header is kept, later exact repeats are removed, and every removal is recorded as evidence.
- One UTF-8 quoted CSV per table in the run's key-value store. Numbers and dates stay strings. CSV export is formula-injection guarded by default (a leading apostrophe is added in CSV only); the dataset JSON always keeps the unmodified strings.
- Optional debug images (up to 10 per run): the real page with detected table and cell boundaries drawn on it, so you can see exactly what was read. Stored privately in your run storage.
Honest limits
- Born-digital PDFs only. Pages without a usable text layer come back as
UNSUPPORTED_OR_INSUFFICIENT_TEXT, uncharged. There is no OCR and no AI in this actor. - A valid page with no detected table is processed and billable, labeled
NO_TABLE_DETECTED. Failed downloads, invalid input, encrypted/malformed files, and unsupported pages are uncharged. - Documents must be publicly reachable HTTPS URLs. Redirects are validated, private-network targets are refused, and signed query strings are never logged or stored.
Input
{"documents": [{"id": "report-1", "url": "https://example.com/report.pdf", "pages": "1-3"}],"strategy": "auto","expectedHeaders": ["Item", "Quantity", "Amount"],"removeExactRepeatedHeaders": true,"debugImages": true}
Output
Dataset rows with recordType of page_receipt, table_receipt, or table_row, plus per-table CSVs, summary.json, and optional debug PNGs in the run's key-value store. Each row carries the document ID, sanitized source URL, source SHA-256, page, table ID, bounding box, strategy, status, and warning codes.
Billing
Pay per selected page successfully processed through table QA. A 10-page selection costs $0.15. Unsupported, failed, or invalid pages are free.