# Invoice & PDF to JSON Parser (OCR) – DOCX to Markdown (`gazidev/document-invoice-to-json`) Actor

Convert PDF, DOCX and scanned documents to clean Markdown, tables and structured JSON. Extracts invoice & receipt fields (vendor, dates, totals, tax, line items, IBAN, VAT ID) with no AI key required. Fast, batch, pay per document.

- **URL**: https://apify.com/gazidev/document-invoice-to-json.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 document processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Invoice & PDF to JSON Parser (OCR) – DOCX to Markdown

Turn **invoices, receipts, PDFs, Word (DOCX) files, scans and photos** into **structured JSON, clean Markdown and tables**. For invoices and receipts it returns vendor, invoice number, dates, subtotal, tax, total, currency, line items, IBAN and VAT IDs. It needs **no OpenAI or other AI key**, and it checks every total against the others.

- **Invoice fields without an AI key.** Labels are recognized in 8 languages and in US and European number formats. Totals are cross-checked, and each result has a `confidence` score. You can optionally refine results with your own Claude or OpenAI key.
- **Any document to Markdown and JSON.** It handles text PDFs, scanned PDFs and phone photos (Tesseract OCR), DOCX and images, and returns ruled tables as `header` + `rows`, ready for RAG and AI agents.
- **Cheap and fast.** $2.30 per 1,000 one-page documents, or $9.30 per 1,000 one-page invoices with extracted fields. A 100-page text PDF takes seconds, and failed files are free.

### Quick start

This is the prefilled input: one sample invoice, done in a few seconds, for about $0.009.

```json
{ "documentUrls": ["https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf"] }
```

#### Sample output

| invoiceNumber | invoiceDate | vendorName | subtotal | taxAmount | totalAmount | currency | confidence |
|---|---|---|---|---|---|---|---|
| INV-3337 | 2016-01-25 | DEMO - Sliced Invoices | 85.00 | 8.50 | 93.50 | USD | 1.0 |

#### Price comparison (Apify Store, September 2026)

| Actor | Price | Monthly users |
|---|---|---|
| **Invoice & PDF to JSON Parser (this Actor)** | **$2.30 per 1,000 one-page docs; $5.00 per 1,000 ten-page PDFs; $9.30 per 1,000 invoices with fields** | new |
| automation-lab/pdf-text-extractor | $3.45 per 1,000 | 67 |
| opportunity-biz/document-to-json-mcp | $0.01 per invoice ($10 per 1,000) | 1 |

**Related Actors:** [SEC EDGAR MCP Server & API](https://apify.com/gazidev/sec-edgar-mcp) for SEC filings and XBRL financials, and [Tenders Scraper & Alerts](https://apify.com/gazidev/global-tender-alerts) for public procurement tenders.

***

### What does it extract?

| For every document | For invoices & receipts (`invoice` object) |
|---|---|
| Markdown (headings, paragraphs, lists, tables) | `vendorName`, `customerName` |
| Tables as JSON (`header` + `rows`, page, bounding box) | `invoiceNumber`, `invoiceDate`, `dueDate` (ISO 8601) |
| Plain text (optional) | `currency` (ISO 4217) |
| Per-page text / Markdown / tables (optional) | `subtotal`, `taxAmount`, `taxRate`, `totalAmount`, `amountDue` |
| Metadata (title, author, creation date, producer) | `lineItems[]`: description, quantity, unit, unitPrice, amount, taxRate |
| Page count, OCR page count, warnings | `taxIds` (VAT, USt-IdNr, GSTIN, VKN, ABN, SIRET…), `iban` (checksum-validated), `bic`, emails, websites |
| | `validation` (subtotal + tax = total, line items sum) and a `confidence` score from 0 to 1 |

**Supported inputs:** PDF (text-based and scanned), DOCX, PNG, JPG, TIFF (multi-page), WEBP, GIF, BMP and TXT.
Share links from **Google Drive, Dropbox and GitHub** are converted to direct downloads automatically.
You can also upload files directly in the Console.

### Use cases

- **Accounts payable automation:** pull invoice data into your ERP, spreadsheet or accounting tool (QuickBooks, Xero, DATEV, Logo, Paraşüt...).
- **Expense management:** read receipts and scanned bills with OCR.
- **RAG and LLM pipelines:** convert PDFs and Word files to Markdown before chunking and embedding.
- **AI agents:** give Claude, ChatGPT, Cursor or your own agent a tool that reads any document URL.
- **Data extraction from reports:** get tables from government, financial and statistical PDFs as JSON.
- **Document archiving and search:** index text and metadata from large document collections.

### How to use it

1. Paste one or more document URLs into **Document URLs**, or upload files.
2. Keep **Document type = Auto-detect**. Invoices and receipts are recognized automatically, or choose *Invoice* to force invoice extraction.
3. Click **Start**. Results appear in the **Output** tab: *Overview*, *Invoices* (a flat table you can export to Excel/CSV), and the full JSON.

#### Input example

```json
{
  "documentUrls": [
    "https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf",
    "https://arxiv.org/pdf/1706.03762"
  ],
  "documentType": "auto",
  "includeMarkdown": true,
  "includeTables": true,
  "ocrMode": "auto",
  "ocrLanguages": "eng"
}
```

#### Output example (invoice)

```json
{
  "url": "https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf",
  "status": "ok",
  "fileType": "pdf",
  "documentType": "invoice",
  "pageCount": 1,
  "invoice": {
    "invoiceNumber": "INV-3337",
    "invoiceDate": "2016-01-25",
    "dueDate": "2016-01-31",
    "vendorName": "DEMO - Sliced Invoices",
    "customerName": "Test Business",
    "currency": "USD",
    "subtotal": 85.0,
    "taxAmount": 8.5,
    "taxRate": 10.0,
    "totalAmount": 93.5,
    "amountDue": 93.5,
    "lineItems": [
      { "description": "Web Design This is a sample description...", "quantity": 1.0, "unitPrice": 85.0, "amount": 85.0, "taxRate": null }
    ],
    "iban": [],
    "taxIds": [],
    "emails": ["admin@slicedinvoices.com", "test@test.com"],
    "validation": { "subtotalPlusTaxEqualsTotal": true, "lineItemsSumMatches": true },
    "confidence": 1.0,
    "extractionMethod": "heuristic"
  },
  "markdown": "# Invoice\n\n| Invoice Number | INV-3337 |\n| --- | --- |\n| Order Number | 12345 |\n...",
  "tables": [
    { "page": 1, "index": 0, "header": ["Hrs/Qty", "Service", "Rate/Price", "Adjust", "Sub Total"],
      "rows": [["1.00", "Web Design This is a sample description...", "$85.00", "0.00%", "$85.00"]] }
  ],
  "processingTimeMs": 310
}
```

A failed document produces a row like
`{"url": "...", "status": "error", "error": "HTTP 404 when downloading the document"}`,
and **you are not charged for it**.

### Pricing

Pay only for what you process. There is no monthly fee.

| Event | Price |
|---|---|
| Document processed | $0.002 |
| Page processed | $0.0003 |
| OCR page (scans / photos only) | + $0.002 |
| Invoice / receipt fields extracted | + $0.007 |

**Examples:** 1,000 one-page invoices cost **$9.30**. 1,000 ten-page PDFs converted to Markdown cost **$5.00**.
A scanned receipt costs $0.0113.

You are never charged for failed downloads, unsupported or empty files, or pages over your
`maxPagesPerDocument` limit. Set a **maximum cost per run** in the run options and the Actor stops
before exceeding it.

### Accuracy

- **Text PDFs:** the text layer is read directly, so the text itself is exact.
- **Invoices:** built-in rules find each field from its label (in 8 languages), from table columns, and from label rows followed by value rows. Totals are then cross-checked. On the public
  [invoice2data](https://github.com/invoice-x/invoice2data) test invoices (AWS, Flipkart, Coolblue, QualityHosting, Free, OYO, Saeco...), the invoice number, date, total and currency were **correct in 36 of 36 checks**. Note that the rules were tuned while testing on these same invoices.
- **Check the result:** look at `confidence` and `validation`. If `subtotalPlusTaxEqualsTotal` is `true`, the amounts are consistent with each other.
- **Higher accuracy on unusual layouts:** set **LLM refinement** to Anthropic Claude or OpenAI and add *your own* API key. The rule-based result is sent along as a draft, the LLM fills in missing fields, and the validation runs again. Your LLM provider bills you for this directly.

### Using it from AI agents (MCP) and code

**MCP (Claude Desktop, Cursor, VS Code, any MCP client).** Add the Apify MCP server with this Actor as a tool:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com/?actors=gazidev/document-invoice-to-json",
      "headers": { "Authorization": "Bearer <APIFY_TOKEN>" }
    }
  }
}
```

Your agent can then say "extract the totals from this invoice URL" or "read this PDF as Markdown".
Output is one compact JSON row per document, so it fits well in an LLM context. Turn off
`includeTables` or `includeMarkdown` to make it smaller.

**REST API (synchronous, returns the results directly):**

```bash
curl -X POST "https://api.apify.com/v2/acts/gazidev~document-invoice-to-json/run-sync-get-dataset-items?token=<APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"documentUrls": ["https://example.com/invoice.pdf"], "documentType": "invoice"}'
```

**Python:**

```python
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("gazidev/document-invoice-to-json").call(
    run_input={"documentUrls": ["https://example.com/invoice.pdf"]}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["invoice"]["totalAmount"] if item.get("invoice") else item["markdown"][:200])
```

**No-code:** use the Apify integrations for **n8n, Make, Zapier, Google Sheets and webhooks**.
For example: new email attachment → this Actor → a row in Google Sheets.

### Input options

| Field | Default | Description |
|---|---|---|
| `documentUrls` | — | List of document URLs (PDF, DOCX, images, TXT) |
| `uploadedFiles` | — | Files uploaded in the Console |
| `documentType` | `auto` | `auto`, `invoice` or `generic` |
| `includeMarkdown` / `includeTables` / `includeText` / `includePages` | on / on / off / off | Which outputs to include |
| `ocrMode` | `auto` | `auto` (only scanned pages), `force` or `off` |
| `ocrLanguages` | `eng` | e.g. `eng+deu+tur` (eng, deu, fra, spa, ita, por, nld, tur) |
| `tableStrategy` | `lines` | `lines` (ruled tables), `text` (experimental, whitespace-aligned) or `none` |
| `maxPagesPerDocument` | 300 | Page limit per document |
| `maxFileSizeMb` | 50 | Size limit per file |
| `maxConcurrency` | 4 | Documents processed in parallel |
| `pdfPassword` | — | Password for encrypted PDFs |
| `llmProvider` / `llmApiKey` / `llmModel` / `llmBaseUrl` | none | Optional LLM refinement with your own key |

### FAQ

**Do I need an OpenAI or Claude API key?**
No. Invoice fields are extracted with built-in rules and table parsing. An LLM key is optional and
only used if you add one.

**Does it work with scanned PDFs and phone photos of receipts?**
Yes. With `ocrMode: auto`, pages without a text layer and image files go through Tesseract OCR.
Photos should be reasonably sharp and not rotated by more than a few degrees. Only OCR pages cost
the extra $0.002.

**How good is table extraction?**
Ruled (bordered) tables are detected precisely, including merged cells, and returned as
`header` + `rows`. Tables that are only aligned with whitespace, such as academic papers, stay
readable in the Markdown text. You can also try `tableStrategy: text`.

**Is my data stored?**
Documents are processed in memory during the run. Results are kept only in your own Apify dataset,
under your account's data retention settings. If you use LLM refinement, the document text is sent
to the LLM provider you chose.

**Which languages and number formats does it support?**
Invoice labels in English, German, French, Spanish, Italian, Dutch, Portuguese and Turkish.
Numbers like `1,234.56`, `1.234,56`, `1'234.56` and `1 234,56`. Dates like `2024-03-15`, `15.03.2024`,
`03/15/2024`, `15 March 2024`, `7. Mai 2014` and `12 Nisan 2024`. All dates are returned as ISO `YYYY-MM-DD`.

**What happens with very large PDFs?**
Pages beyond `maxPagesPerDocument` are skipped and not charged. A warning in the result tells you
this happened. A 100-page text PDF takes a few seconds and needs about 120 MB of RAM.

**Can I process files from Google Drive or Dropbox?**
Yes, as long as the file is shared publicly. Paste the share link and it is converted to a direct
download automatically.

**Is XLSX or PPTX supported?**
Not yet. If you need these formats, open an issue in the **Issues** tab.

### Limitations

- Right-to-left scripts and CJK OCR are not included (OCR languages: eng, deu, fra, spa, ita, por, nld, tur).
- Invoice extraction uses rules. On unusual layouts some fields can be `null`. Check `confidence` and turn on LLM refinement if you need more accuracy.
- Multi-column layouts, such as scientific papers, are read line by line.

### Feedback

Found a document that isn't parsed well? Open an issue with the (public) URL. Parser improvements
ship regularly.

# Actor input Schema

## `documentUrls` (type: `array`):

Public URLs of PDF, DOCX, image (PNG/JPG/TIFF/WEBP) or TXT files. Google Drive, Dropbox and GitHub share links are converted to direct downloads automatically. One result row is produced per document.

## `uploadedFiles` (type: `array`):

Upload documents directly instead of providing URLs.

## `documentType` (type: `string`):

'auto' detects invoices/receipts and extracts invoice fields only for them. 'invoice' always runs invoice extraction. 'generic' only returns Markdown, text and tables.

## `includeMarkdown` (type: `boolean`):

Clean Markdown with headings, paragraphs, lists and tables - ideal for LLMs, RAG and AI agents.

## `includeTables` (type: `boolean`):

Every detected table as {header, rows} arrays (page number and bounding box included).

## `includeText` (type: `boolean`):

Full plain text of the document.

## `includePages` (type: `boolean`):

Adds a 'pages' array with text, Markdown and tables for each page (useful for citations / page-level RAG).

## `ocrMode` (type: `string`):

'auto' runs OCR only on pages without a text layer (scans, photos) and on image files. 'force' OCRs every page. 'off' never runs OCR (fastest, cheapest). OCR pages are billed with the extra 'ocr-page' event.

## `ocrLanguages` (type: `string`):

Tesseract language codes joined with '+'. Available: eng, deu, fra, spa, ita, por, nld, tur. Example: 'eng+deu'.

## `tableStrategy` (type: `string`):

'lines' detects ruled/bordered tables (accurate, recommended). 'text' additionally tries whitespace-aligned tables when no ruled table is found on a page (can produce noisy tables). 'none' disables table detection.

## `maxPagesPerDocument` (type: `integer`):

Pages beyond this limit are skipped (and not billed).

## `maxFileSizeMb` (type: `integer`):

Larger files are rejected (and not billed).

## `maxConcurrency` (type: `integer`):

How many documents are downloaded/processed in parallel.

## `pdfPassword` (type: `string`):

Password for encrypted PDFs (applied to all documents).

## `llmProvider` (type: `string`):

Invoice fields are extracted with built-in rules and need no API key. Optionally let an LLM refine them using YOUR key - billed by your LLM provider, not by this Actor. 'openai' works with any OpenAI-compatible endpoint (OpenAI, OpenRouter, Groq, Together...).

## `llmApiKey` (type: `string`):

Your Anthropic or OpenAI(-compatible) API key. Stored encrypted.

## `llmModel` (type: `string`):

Model ID. Defaults: 'claude-opus-5' (Anthropic), 'gpt-4o-mini' (OpenAI). Use a smaller model such as 'claude-haiku-4-5' for lower LLM cost.

## `llmBaseUrl` (type: `string`):

Only for provider 'openai'. Default https://api.openai.com/v1 (e.g. https://openrouter.ai/api/v1).

## Actor input object example

```json
{
  "documentUrls": [
    "https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf"
  ],
  "documentType": "auto",
  "includeMarkdown": true,
  "includeTables": true,
  "includeText": false,
  "includePages": false,
  "ocrMode": "auto",
  "ocrLanguages": "eng",
  "tableStrategy": "lines",
  "maxPagesPerDocument": 300,
  "maxFileSizeMb": 50,
  "maxConcurrency": 4,
  "llmProvider": "none"
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documentUrls": [
        "https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/document-invoice-to-json").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "documentUrls": ["https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("gazidev/document-invoice-to-json").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documentUrls": [
    "https://slicedinvoices.com/pdf/wordpress-pdf-invoice-plugin-sample.pdf"
  ]
}' |
apify call gazidev/document-invoice-to-json --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/document-invoice-to-json"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iNgLr2Igw0Na9ZkpY/builds/clly8BeN8OgigzpzA/openapi.json
