# PDF Text, Tables & OCR Extractor (Markdown output) (`everyotherfriday/pdf-extractor`) Actor

Turn PDFs into usable data: per-page text, tables as JSON, document metadata and a Markdown version with headings, plus OCR only where a page has no text layer. Handles page ranges, size caps and encrypted files gracefully. Built for RAG pipelines and document processing.

- **URL**: https://apify.com/everyotherfriday/pdf-extractor.md
- **Developed by:** [Paul Vasquez](https://apify.com/everyotherfriday) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 pdf processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF Text and Tables Extractor with OCR

Turn public PDFs or uploaded files into document rows, full text, Markdown, and
structured table JSON. This Python 3.12 actor uses PyMuPDF for text, metadata and
rendering, pdfplumber for tables, and local Tesseract through pytesseract for OCR.
It needs no API keys, external OCR service, or separate website. Documents are
processed sequentially in the actor container. Use it for report indexing,
document search, research collections, and preparing material for downstream
analysis. Extraction preserves source wording rather than summarizing documents.

### Quick start

Use the included `INPUT.json` to process three public PDFs: W3C's sample, an
arXiv paper, and a Census retail ecommerce report. The local validation completed
well below two minutes; network availability and hosted OCR can change timings.
For local development, create a Python 3.12 virtual environment, install
`requirements.txt`, and run:

```powershell
.venv/Scripts/python.exe -m unittest discover -s tests -v
apify validate-schema .actor/input_schema.json
powershell -NoProfile -ExecutionPolicy Bypass -File validation/run_live.ps1
```

The validation script copies INPUT into a fresh local store, runs the actual
Apify SDK entry point, saves counts and timings, and checks persisted output.
For normal execution use `python -m src` with INPUT in the default KVS. Uploaded
PDFs must be binary records with `application/pdf` content type in that same
run's default key-value store. Set `keyValueStoreKeys` to their record names;
no separate storage ID or account token belongs in actor input.

### Input options

`pdfUrls` accepts HTTP or HTTPS URLs; `keyValueStoreKeys` is the optional upload
alternative. Either list can be omitted. Exact duplicates within each list are
processed once. URLs must not contain embedded username/password credentials.
`pages` accepts one-based selections such as `1-5,8`; duplicate selections are
merged and results use ascending document order. Out-of-range pages are ignored.
A selection containing no actual pages produces an uncharged summary row.

`extractText` and `extractTables` default to true. Disabling text also disables
OCR. `ocrMode` defaults to `auto`, which attempts OCR only when a selected page
has fewer than 20 stripped native characters. `always` replaces each selected
page's native text with recognized text; `never` disables recognition.
`ocrLanguage` defaults to `eng`. The Docker image includes English; other
Tesseract language packs require a custom image. Multiple installed languages
can be requested with codes such as `eng+deu`.

`outputMarkdown` defaults to true. `maxPages`, default 200, rejects an entire
document above the limit, even when only a subset was requested. `maxFileMb`,
default 50, means MiB and applies to downloaded or uploaded bytes. Downloads
check both declared length and streamed content. `timeoutSecs` defaults to 30
and bounds each download attempt and each OCR subprocess, not total extraction
CPU time. The example uses 15 seconds. HTTP 429, server failures and transport
failures receive two retries after one and two seconds; other HTTP errors do not.

### Results and files

One dataset row represents each document. It includes `source`, `fileName`,
`fileSizeBytes`, total `pageCount`, `pagesProcessed`, metadata, `textPages`,
`ocrPages`, `tables`, `warnings`, and nullable `error`. Metadata includes title,
author, subject, creator, producer, createdAt and modifiedAt. Dates retain their
original PDF representation; missing values are null.

`fullText` contains at most 50,000 characters. The complete extracted text is
saved as `TEXT-<hash>.txt`, identified by the additional `textKey` field.
`markdownKey` points to `MARKDOWN-<hash>.md`; `tablesKey` points to
`TABLES-<hash>.json`. Hashes incorporate source, document bytes and extraction
settings. Tables contain `{page, index, rows}` with one-based page/table numbers
and nested cell arrays, including null cells when applicable. Each `pages`
entry contains page number, complete character count, OCR status, and at most
2,000 characters of text. `SUMMARY` stores document totals, eligible event
requests and per-document timings. Empty source lists produce one free summary.

### Pricing and extraction limits

The custom events are `pdf-processed` at $0.003 per successfully opened,
accepted document with selected pages; `text-page` at $0.0005 per page with
nonempty native extracted text; and `ocr-page` at $0.008 per completed OCR page,
including recognition yielding empty text. A page with native text that also
undergoes OCR incurs both page events. Failed downloads, malformed or encrypted
PDFs, and over-limit documents produce uncharged error rows. Failed OCR attempts
are uncharged and retain native text with a warning. Missing Tesseract also
produces warnings. Local SDK runs request events but do not bill.

Artifacts are stored before charging, then the dataset row is written. Budget
checks reject insufficient document budgets before charges start. Billing and
storage are not transactional: a later charge or dataset failure cannot undo
earlier charges. Configure all three custom events and disable synthetic charges
before publication. Docker and real hosted billing remain unverified locally.

Markdown headings are font-size guesses relative to the dominant character
font; OCR text has page headings only. Multi-column reading order, mathematical
notation, merged cells, and borderless tables can be imperfect. Table detection
uses native PDF geometry and does not reconstruct scanned tables from OCR.
OCR accuracy depends on scan quality. Windows tests skip real OCR when
Tesseract is absent; no system software is installed by validation. See
`VALIDATION.md` for measured coverage and the exact remaining limitations.

### Example output

Recorded local validation output from [storage/live-20260926-064724/datasets/default/000000001.json](storage/live-20260926-064724/datasets/default/000000001.json), dataset row 1. Fields are omitted for brevity; retained values are unchanged. This is a historical example, not a current-source claim.

```json
{
    "source":  "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "fileName":  "dummy.pdf",
    "fileSizeBytes":  13264,
    "pageCount":  1,
    "pagesProcessed":  1,
    "textPages":  1,
    "ocrPages":  0,
    "tables":  0,
    "fullText":  "Dummy PDF file",
    "warnings":  [
                     "Page 1: OCR unavailable; Tesseract not installed"
                 ],
    "error":  null
}
```

### Use cases

- A research librarian extracts native text from public papers and uses the full-text artifacts to populate a document index.
- A retail analyst extracts tables from government reports, then checks cell alignment against the original PDFs before spreadsheet analysis.
- A records digitization contractor processes scanned pages with installed OCR languages and reviews warnings and recognized text for quality.
- A knowledge-base administrator converts selected report pages into Markdown and retains source links for later editorial verification.

### Pricing example

100 accepted PDFs, 1,000 nonempty native-text pages, and 50 completed OCR pages cost (100 x $0.003) + (1,000 x $0.0005) + (50 x $0.008) = **$1.20** in declared events. These are event counts: a page can contribute to both page categories. Rates come from [the local event declaration](.actor/pay_per_event.json). This calculation is an event subtotal, not a measured invoice; local validation does not bill.

### Limitations

This saved row retains its actual missing-Tesseract warning; it demonstrates native extraction, not successful OCR. Extraction does not establish factual accuracy of the source document. Review page order and table cells before downstream use, and retrieve the text artifact when the dataset text cap would omit content.

# Actor input Schema

## `pdfUrls` (type: `array`):

Public HTTP(S) PDF URLs; no credentials needed. May be omitted for KVS-only input.

## `keyValueStoreKeys` (type: `array`):

Binary PDF keys uploaded to this run default key-value store.

## `pages` (type: `string`):

Optional one-based page ranges, for example 1-5,8; absent means all pages.

## `extractText` (type: `boolean`):

Extract native text and enable the selected OCR policy.

## `extractTables` (type: `boolean`):

Extract native PDF tables with pdfplumber; scanned table reconstruction is not supported.

## `ocrMode` (type: `string`):

auto: OCR pages below 20 native characters; always: OCR all selected pages; never: disable OCR.

## `ocrLanguage` (type: `string`):

Tesseract language code; Docker includes eng only. Additional languages need a custom image.

## `outputMarkdown` (type: `boolean`):

Save per-document Markdown; font-size heading guesses apply to native text.

## `maxPages` (type: `integer`):

Reject documents whose total page count exceeds this limit, even with a page selection.

## `maxFileMb` (type: `integer`):

Reject files above this many MiB, checking headers and streamed bytes.

## `timeoutSecs` (type: `integer`):

Total timeout for each download attempt and each Tesseract call; not a whole-run CPU timeout.

## Actor input object example

```json
{
  "extractText": true,
  "extractTables": true,
  "ocrMode": "auto",
  "ocrLanguage": "eng",
  "outputMarkdown": true,
  "maxPages": 200,
  "maxFileMb": 50,
  "timeoutSecs": 30
}
```

# Actor output Schema

## `documents` (type: `string`):

JSON document and uncharged error rows.

## `files` (type: `string`):

TEXT, MARKDOWN, TABLES and SUMMARY records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("everyotherfriday/pdf-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("everyotherfriday/pdf-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call everyotherfriday/pdf-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,everyotherfriday/pdf-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0LlQ8zTGBeWka1yGg/builds/YzcikbkvGzHmbtitX/openapi.json
