# Document → Structured Data Extractor (`warant/document-structured-data-extractor`) Actor

PDF or image documents in — full text, detected tables and business fields (dates, amounts, totals, reference numbers, emails, phones) out as JSON/CSV. Deterministic parsing with Tesseract OCR for scans. No LLM, no login.

- **URL**: https://apify.com/warant/document-structured-data-extractor.md
- **Developed by:** [Waran T](https://apify.com/warant) (community)
- **Categories:**
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 document processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Document → Structured Data Extractor

**PDF or image documents in — structured data out.** Feed it document URLs; get back the full text, **detected tables as rows you can export to CSV**, and **pattern-detected business fields** — dates, amounts, totals, reference/invoice numbers, emails, phones — as clean JSON.

Most PDF extractors dump raw text and stop. This one gives your workflow something it can actually *use* downstream: rows, fields, and per-page provenance.

### What you get per document

| Output | Description |
|---|---|
| `text` | Full extracted text (per-page, layout-aware line reconstruction) |
| `tables[]` | Detected tables as `rows[][]` — export straight to CSV/Excel |
| `fields{}` | `dates`, `amounts`, `totals`, `referenceNumbers`, `emails`, `phones` |
| `pages[]` | Per-page character counts + extraction method (text-layer vs OCR) |
| `kind`, `pagesProcessed`, `error` | Type detection and honest failure reporting |

### Pricing — pay only for work actually performed

| Event | Price |
|---|---|
| Document processed (tables + fields extraction) | **$0.01** |
| Text-layer page | **$0.001** |
| OCR'd page (images / scans) | **$0.02** |

Examples: a 20-page digital PDF ≈ **$0.03**. A 200-page digital PDF ≈ **$0.21**. 50 scanned images ≈ **$1.01**. Failed downloads and unparseable documents are **not charged**.

### What it does NOT do (honesty section)

- **No LLM, no AI guessing** — everything is deterministically parsed or OCR'd; fields are regex/pattern-detected and should be spot-checked for critical use.
- **Scanned PDF pages** (image-only, no text layer) are **flagged `needsOcr: true` and not charged** in this version — direct image files (PNG/JPG/TIFF) *are* OCR'd. Send scans as images for OCR.
- No login-protected documents, no DRM circumvention.
- Table detection is conservative — it prefers missing an ambiguous table over fabricating rows.

### Input

```json
{
  "documents": ["https://example.com/invoice.pdf", "https://example.com/scan.png"],
  "ocr": true,
  "language": "eng",
  "maxPages": 50
}
```

Up to 200 documents per run, 500 pages per document, 25 MB per file. OCR supports Tesseract language codes (`eng`, `deu`, `fra`, `spa`, …).

### Typical uses

- Invoices / receipts → amounts, dates, reference numbers for bookkeeping automation
- Product catalogs / spec sheets → tables to CSV
- Reports and filings → searchable text + key figures
- Contact-bearing documents → emails and phones with document provenance

### FAQ

**Why not a flat per-page price?** Because a text-layer page costs ~40× less to process than an OCR page. You shouldn't subsidize other people's scans.

**What about scanned multi-page PDFs?** v1 flags those pages (`needsOcr`) without charging. Convert to images for OCR, or watch for the next version.

**Is my document stored?** Documents are fetched, processed in the run's container, and the extracted results are written to your dataset. The Actor keeps no copy outside your run's storage.

# Actor input Schema

## `documents` (type: `array`):

Direct http(s) URLs to PDF or image (PNG/JPG/TIFF) documents. Max 200 per run.

## `ocr` (type: `boolean`):

Run Tesseract OCR on images and scanned PDF pages that have no text layer.

## `language` (type: `string`):

Tesseract language code (e.g. eng, deu, fra, spa).

## `maxPages` (type: `integer`):

Cap pages processed per document (1-500).

## Actor input object example

```json
{
  "documents": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "ocr": true,
  "language": "eng",
  "maxPages": 50
}
```

# Actor output Schema

## `documents` (type: `string`):

The dataset of extracted document records, one item per input document.

## `overview` (type: `string`):

Human-friendly table view of the extraction results.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documents": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("warant/document-structured-data-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "documents": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("warant/document-structured-data-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documents": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}' |
apify call warant/document-structured-data-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,warant/document-structured-data-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pN59NCfJMHJrfneVs/builds/xzWWbsXHX9HZub9gc/openapi.json
