Invoice & PDF to JSON Parser (OCR) – DOCX to Markdown avatar

Invoice & PDF to JSON Parser (OCR) – DOCX to Markdown

Pricing

from $2.00 / 1,000 document processeds

Go to Apify Store
Invoice & PDF to JSON Parser (OCR) – DOCX to Markdown

Invoice & PDF to JSON Parser (OCR) – DOCX to Markdown

Convert PDF, DOCX and scanned documents to clean Markdown, tables and structured JSON. Extracts invoice & receipt fields (vendor, dates, totals, tax, line items, IBAN, VAT ID) with no AI key required. Fast, batch, pay per document.

Pricing

from $2.00 / 1,000 document processeds

Rating

0.0

(0)

Developer

Cemal Atakli

Cemal Atakli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Turn invoices, receipts, PDFs, Word (DOCX) files, scans and photos into structured JSON, clean Markdown and tables. For invoices and receipts it returns vendor, invoice number, dates, subtotal, tax, total, currency, line items, IBAN and VAT IDs. It needs no OpenAI or other AI key, and it checks every total against the others.

  • Invoice fields without an AI key. Labels are recognized in 8 languages and in US and European number formats. Totals are cross-checked, and each result has a confidence score. You can optionally refine results with your own Claude or OpenAI key.
  • Any document to Markdown and JSON. It handles text PDFs, scanned PDFs and phone photos (Tesseract OCR), DOCX and images, and returns ruled tables as header + rows, ready for RAG and AI agents.
  • Cheap and fast. $2.30 per 1,000 one-page documents, or $9.30 per 1,000 one-page invoices with extracted fields. A 100-page text PDF takes seconds, and failed files are free.

Quick start

This is the prefilled input: one sample invoice, done in a few seconds, for about $0.009.

{ "documentUrls": ["https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf"] }

Sample output

invoiceNumberinvoiceDatevendorNamesubtotaltaxAmounttotalAmountcurrencyconfidence
INV-33372016-01-25DEMO - Sliced Invoices85.008.5093.50USD1.0

Price comparison (Apify Store, September 2026)

ActorPriceMonthly users
Invoice & PDF to JSON Parser (this Actor)$2.30 per 1,000 one-page docs; $5.00 per 1,000 ten-page PDFs; $9.30 per 1,000 invoices with fieldsnew
automation-lab/pdf-text-extractor$3.45 per 1,00067
opportunity-biz/document-to-json-mcp$0.01 per invoice ($10 per 1,000)1

Related Actors: SEC EDGAR MCP Server & API for SEC filings and XBRL financials, and Tenders Scraper & Alerts for public procurement tenders.


What does it extract?

For every documentFor invoices & receipts (invoice object)
Markdown (headings, paragraphs, lists, tables)vendorName, customerName
Tables as JSON (header + rows, page, bounding box)invoiceNumber, invoiceDate, dueDate (ISO 8601)
Plain text (optional)currency (ISO 4217)
Per-page text / Markdown / tables (optional)subtotal, taxAmount, taxRate, totalAmount, amountDue
Metadata (title, author, creation date, producer)lineItems[]: description, quantity, unit, unitPrice, amount, taxRate
Page count, OCR page count, warningstaxIds (VAT, USt-IdNr, GSTIN, VKN, ABN, SIRET…), iban (checksum-validated), bic, emails, websites
validation (subtotal + tax = total, line items sum) and a confidence score from 0 to 1

Supported inputs: PDF (text-based and scanned), DOCX, PNG, JPG, TIFF (multi-page), WEBP, GIF, BMP and TXT. Share links from Google Drive, Dropbox and GitHub are converted to direct downloads automatically. You can also upload files directly in the Console.

Use cases

  • Accounts payable automation: pull invoice data into your ERP, spreadsheet or accounting tool (QuickBooks, Xero, DATEV, Logo, Paraşüt...).
  • Expense management: read receipts and scanned bills with OCR.
  • RAG and LLM pipelines: convert PDFs and Word files to Markdown before chunking and embedding.
  • AI agents: give Claude, ChatGPT, Cursor or your own agent a tool that reads any document URL.
  • Data extraction from reports: get tables from government, financial and statistical PDFs as JSON.
  • Document archiving and search: index text and metadata from large document collections.

How to use it

  1. Paste one or more document URLs into Document URLs, or upload files.
  2. Keep Document type = Auto-detect. Invoices and receipts are recognized automatically, or choose Invoice to force invoice extraction.
  3. Click Start. Results appear in the Output tab: Overview, Invoices (a flat table you can export to Excel/CSV), and the full JSON.

Input example

{
"documentUrls": [
"https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf",
"https://arxiv.org/pdf/1706.03762"
],
"documentType": "auto",
"includeMarkdown": true,
"includeTables": true,
"ocrMode": "auto",
"ocrLanguages": "eng"
}

Output example (invoice)

{
"url": "https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf",
"status": "ok",
"fileType": "pdf",
"documentType": "invoice",
"pageCount": 1,
"invoice": {
"invoiceNumber": "INV-3337",
"invoiceDate": "2016-01-25",
"dueDate": "2016-01-31",
"vendorName": "DEMO - Sliced Invoices",
"customerName": "Test Business",
"currency": "USD",
"subtotal": 85.0,
"taxAmount": 8.5,
"taxRate": 10.0,
"totalAmount": 93.5,
"amountDue": 93.5,
"lineItems": [
{ "description": "Web Design This is a sample description...", "quantity": 1.0, "unitPrice": 85.0, "amount": 85.0, "taxRate": null }
],
"iban": [],
"taxIds": [],
"emails": ["admin@slicedinvoices.com", "test@test.com"],
"validation": { "subtotalPlusTaxEqualsTotal": true, "lineItemsSumMatches": true },
"confidence": 1.0,
"extractionMethod": "heuristic"
},
"markdown": "# Invoice\n\n| Invoice Number | INV-3337 |\n| --- | --- |\n| Order Number | 12345 |\n...",
"tables": [
{ "page": 1, "index": 0, "header": ["Hrs/Qty", "Service", "Rate/Price", "Adjust", "Sub Total"],
"rows": [["1.00", "Web Design This is a sample description...", "$85.00", "0.00%", "$85.00"]] }
],
"processingTimeMs": 310
}

A failed document produces a row like {"url": "...", "status": "error", "error": "HTTP 404 when downloading the document"}, and you are not charged for it.

Pricing

Pay only for what you process. There is no monthly fee.

EventPrice
Document processed$0.002
Page processed$0.0003
OCR page (scans / photos only)+ $0.002
Invoice / receipt fields extracted+ $0.007

Examples: 1,000 one-page invoices cost $9.30. 1,000 ten-page PDFs converted to Markdown cost $5.00. A scanned receipt costs $0.0113.

You are never charged for failed downloads, unsupported or empty files, or pages over your maxPagesPerDocument limit. Set a maximum cost per run in the run options and the Actor stops before exceeding it.

Accuracy

  • Text PDFs: the text layer is read directly, so the text itself is exact.
  • Invoices: built-in rules find each field from its label (in 8 languages), from table columns, and from label rows followed by value rows. Totals are then cross-checked. On the public invoice2data test invoices (AWS, Flipkart, Coolblue, QualityHosting, Free, OYO, Saeco...), the invoice number, date, total and currency were correct in 36 of 36 checks. Note that the rules were tuned while testing on these same invoices.
  • Check the result: look at confidence and validation. If subtotalPlusTaxEqualsTotal is true, the amounts are consistent with each other.
  • Higher accuracy on unusual layouts: set LLM refinement to Anthropic Claude or OpenAI and add your own API key. The rule-based result is sent along as a draft, the LLM fills in missing fields, and the validation runs again. Your LLM provider bills you for this directly.

Using it from AI agents (MCP) and code

MCP (Claude Desktop, Cursor, VS Code, any MCP client). Add the Apify MCP server with this Actor as a tool:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com/?actors=gazidev/document-invoice-to-json",
"headers": { "Authorization": "Bearer <APIFY_TOKEN>" }
}
}
}

Your agent can then say "extract the totals from this invoice URL" or "read this PDF as Markdown". Output is one compact JSON row per document, so it fits well in an LLM context. Turn off includeTables or includeMarkdown to make it smaller.

REST API (synchronous, returns the results directly):

curl -X POST "https://api.apify.com/v2/acts/gazidev~document-invoice-to-json/run-sync-get-dataset-items?token=<APIFY_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"documentUrls": ["https://example.com/invoice.pdf"], "documentType": "invoice"}'

Python:

from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("gazidev/document-invoice-to-json").call(
run_input={"documentUrls": ["https://example.com/invoice.pdf"]}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["invoice"]["totalAmount"] if item.get("invoice") else item["markdown"][:200])

No-code: use the Apify integrations for n8n, Make, Zapier, Google Sheets and webhooks. For example: new email attachment → this Actor → a row in Google Sheets.

Input options

FieldDefaultDescription
documentUrls—List of document URLs (PDF, DOCX, images, TXT)
uploadedFiles—Files uploaded in the Console
documentTypeautoauto, invoice or generic
includeMarkdown / includeTables / includeText / includePageson / on / off / offWhich outputs to include
ocrModeautoauto (only scanned pages), force or off
ocrLanguagesenge.g. eng+deu+tur (eng, deu, fra, spa, ita, por, nld, tur)
tableStrategylineslines (ruled tables), text (experimental, whitespace-aligned) or none
maxPagesPerDocument300Page limit per document
maxFileSizeMb50Size limit per file
maxConcurrency4Documents processed in parallel
pdfPassword—Password for encrypted PDFs
llmProvider / llmApiKey / llmModel / llmBaseUrlnoneOptional LLM refinement with your own key

FAQ

Do I need an OpenAI or Claude API key? No. Invoice fields are extracted with built-in rules and table parsing. An LLM key is optional and only used if you add one.

Does it work with scanned PDFs and phone photos of receipts? Yes. With ocrMode: auto, pages without a text layer and image files go through Tesseract OCR. Photos should be reasonably sharp and not rotated by more than a few degrees. Only OCR pages cost the extra $0.002.

How good is table extraction? Ruled (bordered) tables are detected precisely, including merged cells, and returned as header + rows. Tables that are only aligned with whitespace, such as academic papers, stay readable in the Markdown text. You can also try tableStrategy: text.

Is my data stored? Documents are processed in memory during the run. Results are kept only in your own Apify dataset, under your account's data retention settings. If you use LLM refinement, the document text is sent to the LLM provider you chose.

Which languages and number formats does it support? Invoice labels in English, German, French, Spanish, Italian, Dutch, Portuguese and Turkish. Numbers like 1,234.56, 1.234,56, 1'234.56 and 1 234,56. Dates like 2024-03-15, 15.03.2024, 03/15/2024, 15 March 2024, 7. Mai 2014 and 12 Nisan 2024. All dates are returned as ISO YYYY-MM-DD.

What happens with very large PDFs? Pages beyond maxPagesPerDocument are skipped and not charged. A warning in the result tells you this happened. A 100-page text PDF takes a few seconds and needs about 120 MB of RAM.

Can I process files from Google Drive or Dropbox? Yes, as long as the file is shared publicly. Paste the share link and it is converted to a direct download automatically.

Is XLSX or PPTX supported? Not yet. If you need these formats, open an issue in the Issues tab.

Limitations

  • Right-to-left scripts and CJK OCR are not included (OCR languages: eng, deu, fra, spa, ita, por, nld, tur).
  • Invoice extraction uses rules. On unusual layouts some fields can be null. Check confidence and turn on LLM refinement if you need more accuracy.
  • Multi-column layouts, such as scientific papers, are read line by line.

Feedback

Found a document that isn't parsed well? Open an issue with the (public) URL. Parser improvements ship regularly.