Invoice to Excel Extractor avatar

Invoice to Excel Extractor

Pricing

from $50.00 / 1,000 results

Go to Apify Store
Invoice to Excel Extractor

Invoice to Excel Extractor

Extracts vendor, dates, line items, subtotal, tax and total from uploaded invoice PDFs/images (OCR + regex/table parsing, no AI dependency) and pushes structured rows to the dataset, plus an Excel workbook per invoice.

Pricing

from $50.00 / 1,000 results

Rating

0.0

(0)

Developer

MST MORIUM AKTHER MAYA

MST MORIUM AKTHER MAYA

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

2 days ago

Last modified

Categories

Share

Extracts structured data from invoices (PDF, PNG, JPG, JPEG, TIF, TIFF, BMP, WEBP) and turns each one into a clean Excel workbook — no AI/LLM dependency, so pricing and behavior stay predictable.

What it extracts

  • Header fields: vendor, invoice number, invoice date, subtotal, tax, total
  • Line items: description, quantity, unit price, and line total (supports both classic 4-column tables and simpler description+amount rows)

Each field comes with a confidence score and a source (native_text for a PDF's real text layer, ocr for scanned/image invoices). A field is only ever filled in when there's direct evidence for it in the document — if OCR confidence is too low to trust, the field is left blank and flagged rather than guessed. This keeps the output honest: what you see is what the document actually says, not an AI's best guess.

How it works

  1. For PDFs with a real text layer, the native text and table structure are read directly (fast, high accuracy).
  2. For scanned documents or images, Tesseract OCR extracts text with per-word confidence scores.
  3. Regex and layout-aware rules pull out header fields and line items from the extracted text.
  4. Results are written to the dataset (as structured JSON) and to a per-invoice Excel workbook (.xlsx) with a Header sheet and a Line Items sheet.

Input

FieldDescription
Invoice files (required)Upload files directly, or paste direct URLs to PDF/image files.
Source storageOnly needed if you uploaded files directly above — select that same key-value store so the Actor can read them.
OCR languageTesseract language code for scanned pages (default: eng).
OCR confidence thresholdMinimum OCR word confidence (0–100) required to trust a value (default: 60).

Output

One row per invoice in the dataset, plus an Excel file per invoice available for download via the excel_url field in that row. Each field/line item carries value, source, confidence, flagged, and flag_reason so you can see exactly how confident the extraction is and why anything was left blank.

Known limitations

  • Line-item and vendor extraction on scanned images depends on OCR quality; very low-resolution or heavily multi-column layouts can occasionally merge text from unrelated parts of the page.
  • Date parsing does not attempt to disambiguate DD/MM vs. MM/DD formats — an ambiguous date is left blank rather than guessed.