Invoice to Excel Extractor
Pricing
from $50.00 / 1,000 results
Invoice to Excel Extractor
Extracts vendor, dates, line items, subtotal, tax and total from uploaded invoice PDFs/images (OCR + regex/table parsing, no AI dependency) and pushes structured rows to the dataset, plus an Excel workbook per invoice.
Pricing
from $50.00 / 1,000 results
Rating
0.0
(0)
Developer
MST MORIUM AKTHER MAYA
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
2 days ago
Last modified
Categories
Share
Extracts structured data from invoices (PDF, PNG, JPG, JPEG, TIF, TIFF, BMP, WEBP) and turns each one into a clean Excel workbook — no AI/LLM dependency, so pricing and behavior stay predictable.
What it extracts
- Header fields: vendor, invoice number, invoice date, subtotal, tax, total
- Line items: description, quantity, unit price, and line total (supports both classic 4-column tables and simpler description+amount rows)
Each field comes with a confidence score and a source (native_text for a PDF's real text layer, ocr for scanned/image invoices). A field is only ever filled in when there's direct evidence for it in the document — if OCR confidence is too low to trust, the field is left blank and flagged rather than guessed. This keeps the output honest: what you see is what the document actually says, not an AI's best guess.
How it works
- For PDFs with a real text layer, the native text and table structure are read directly (fast, high accuracy).
- For scanned documents or images, Tesseract OCR extracts text with per-word confidence scores.
- Regex and layout-aware rules pull out header fields and line items from the extracted text.
- Results are written to the dataset (as structured JSON) and to a per-invoice Excel workbook (
.xlsx) with a Header sheet and a Line Items sheet.
Input
| Field | Description |
|---|---|
| Invoice files (required) | Upload files directly, or paste direct URLs to PDF/image files. |
| Source storage | Only needed if you uploaded files directly above — select that same key-value store so the Actor can read them. |
| OCR language | Tesseract language code for scanned pages (default: eng). |
| OCR confidence threshold | Minimum OCR word confidence (0–100) required to trust a value (default: 60). |
Output
One row per invoice in the dataset, plus an Excel file per invoice available for download via the excel_url field in that row. Each field/line item carries value, source, confidence, flagged, and flag_reason so you can see exactly how confident the extraction is and why anything was left blank.
Known limitations
- Line-item and vendor extraction on scanned images depends on OCR quality; very low-resolution or heavily multi-column layouts can occasionally merge text from unrelated parts of the page.
- Date parsing does not attempt to disambiguate DD/MM vs. MM/DD formats — an ambiguous date is left blank rather than guessed.