# PDF Table to CSV with QA Receipts (`muazah/pdf-table-receipts`) Actor

- **URL**: https://apify.com/muazah/pdf-table-receipts.md
- **Developed by:** [muazah](https://apify.com/muazah) (community)
- **Categories:** Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$15.00 / 1,000 page processed through table qas

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF Table to CSV with QA Receipts

Batch-extract tables from born-digital PDFs and get receipts you can audit, not just a blob of cells.

### What it does

- Pulls up to 5 public HTTPS PDFs per run, 10 selected pages per PDF, 50 pages per run, 100 tables per run.
- Detects ruled tables (`lines`), borderless aligned columns (`text`), or `auto`, which tries each deterministic strategy once and records which one it used.
- Every table is separated with page number, table index, and bounding box. Unrelated side-by-side tables are never merged, and tables are not joined across pages unless you opt in (matching headers and column counts only, page provenance preserved).
- Keeps raw cell strings exactly as printed: leading zeros, decimal commas, quotes, newlines. Blank vs. missing cells stay distinct, and every row keeps its source row index.
- QA receipts per page and per table: column-count consistency, expected-header match, duplicate headers, empty-cell density, suspected merged cells. These are structural heuristics, not a guarantee that cell values are semantically correct.
- Optional repeated-header removal, only when you supply expected headers: the first header is kept, later exact repeats are removed, and every removal is recorded as evidence.
- One UTF-8 quoted CSV per table in the run's key-value store. Numbers and dates stay strings. CSV export is formula-injection guarded by default (a leading apostrophe is added in CSV only); the dataset JSON always keeps the unmodified strings.
- Optional debug images (up to 10 per run): the real page with detected table and cell boundaries drawn on it, so you can see exactly what was read. Stored privately in your run storage.

### Honest limits

- Born-digital PDFs only. Pages without a usable text layer come back as `UNSUPPORTED_OR_INSUFFICIENT_TEXT`, uncharged. There is no OCR and no AI in this actor.
- A valid page with no detected table is processed and billable, labeled `NO_TABLE_DETECTED`. Failed downloads, invalid input, encrypted/malformed files, and unsupported pages are uncharged.
- Documents must be publicly reachable HTTPS URLs. Redirects are validated, private-network targets are refused, and signed query strings are never logged or stored.

### Input

```json
{
  "documents": [{"id": "report-1", "url": "https://example.com/report.pdf", "pages": "1-3"}],
  "strategy": "auto",
  "expectedHeaders": ["Item", "Quantity", "Amount"],
  "removeExactRepeatedHeaders": true,
  "debugImages": true
}
```

### Output

Dataset rows with `recordType` of `page_receipt`, `table_receipt`, or `table_row`, plus per-table CSVs, `summary.json`, and optional debug PNGs in the run's key-value store. Each row carries the document ID, sanitized source URL, source SHA-256, page, table ID, bounding box, strategy, status, and warning codes.

### Billing

Pay per selected page successfully processed through table QA. A 10-page selection costs $0.15. Unsupported, failed, or invalid pages are free.

# Actor input Schema

## `documents` (type: `array`):

Up to 5 public HTTPS PDFs per run. Each: id, url, pages (one-based, e.g. "1-3,5"; explicit selection required, max 10 per PDF). Optional crops: \[{"page": 2, "bbox": \[x0, top, x1, bottom]}] in PDF points from the top-left.

## `strategy` (type: `string`):

lines: ruled tables. text: borderless aligned columns. auto tries lines once, then text once, and records which it used.

## `expectedHeaders` (type: `array`):

Optional exact column headers. Enables header matching, mismatch warnings, stable column fields, and repeated-header removal.

## `removeExactRepeatedHeaders` (type: `boolean`):

Only with Expected headers. Removes later exact header repeats (e.g. across pages) and keeps evidence of every removed row.

## `joinMatchingTablesAcrossPages` (type: `boolean`):

Opt-in. Off by default: tables on different pages stay separate. When on, only tables with matching normalized headers and column counts join, and page provenance is preserved.

## `debugImages` (type: `boolean`):

Render up to 10 page images per run with detected table and cell boundaries drawn on the real page. Stored privately in the run's key-value store.

## `spreadsheetSafeCsv` (type: `boolean`):

On by default: cells starting with = + - @ are prefixed with an apostrophe in CSV files (formula-injection guard). Dataset JSON always keeps the unmodified strings.

## Actor input object example

```json
{
  "documents": [
    {
      "id": "demo-1",
      "url": "https://raw.githubusercontent.com/jsvine/pdfplumber/stable/tests/pdfs/nics-background-checks-2015-11.pdf",
      "pages": "1"
    }
  ],
  "strategy": "auto",
  "debugImages": true,
  "spreadsheetSafeCsv": true
}
```

# Actor output Schema

## `receiptRows` (type: `string`):

page\_receipt / table\_receipt / table\_row records: status, warning codes, bounding boxes, strategies, source row indexes and cells.

## `summary` (type: `string`):

Billable pages, unsupported pages, tables, warnings and artifact references.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documents": [
        {
            "id": "demo-1",
            "url": "https://raw.githubusercontent.com/jsvine/pdfplumber/stable/tests/pdfs/nics-background-checks-2015-11.pdf",
            "pages": "1"
        }
    ],
    "strategy": "auto",
    "removeExactRepeatedHeaders": false,
    "joinMatchingTablesAcrossPages": false,
    "debugImages": true,
    "spreadsheetSafeCsv": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("muazah/pdf-table-receipts").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "documents": [{
            "id": "demo-1",
            "url": "https://raw.githubusercontent.com/jsvine/pdfplumber/stable/tests/pdfs/nics-background-checks-2015-11.pdf",
            "pages": "1",
        }],
    "strategy": "auto",
    "removeExactRepeatedHeaders": False,
    "joinMatchingTablesAcrossPages": False,
    "debugImages": True,
    "spreadsheetSafeCsv": True,
}

# Run the Actor and wait for it to finish
run = client.actor("muazah/pdf-table-receipts").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documents": [
    {
      "id": "demo-1",
      "url": "https://raw.githubusercontent.com/jsvine/pdfplumber/stable/tests/pdfs/nics-background-checks-2015-11.pdf",
      "pages": "1"
    }
  ],
  "strategy": "auto",
  "removeExactRepeatedHeaders": false,
  "joinMatchingTablesAcrossPages": false,
  "debugImages": true,
  "spreadsheetSafeCsv": true
}' |
apify call muazah/pdf-table-receipts --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,muazah/pdf-table-receipts"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qfSFxFpELgNAFovNh/builds/or5PqNhwKh998yP4x/openapi.json
