# PDF Text Extractor: OCR & Metadata (`maximedupre/pdf-to-text-ocr`) Actor

Extract text from public PDF URLs, including scanned pages with OCR. Choose page ranges and get plain text, Markdown, page details, OCR information, metadata, bookmarks, and page-tagged chunks when available in your dataset.

- **URL**: https://apify.com/maximedupre/pdf-to-text-ocr.md
- **Developed by:** [Maxime Dupré](https://apify.com/maximedupre) (community)
- **Categories:** Developer tools, Automation, Education
- **Stats:** 3 total users, 2 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.70 / 1,000 pdf texts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### 📄 Turn PDFs into text with OCR

For developers, researchers, and teams looking for PDF to text OCR, this Actor reads public PDF URLs and returns plain text, Markdown, page details, OCR details, metadata, bookmarks, and text chunks in your dataset. Choose pages when you need part of a document, or use OCR for scanned pages.

- Read a scanned document with [**Scanned PDF to Text**](https://apify.com/maximedupre/pdf-to-text-ocr/examples/scanned-pdf-to-text) and save its recovered text.
- [**Extract Text from PDF**](https://apify.com/maximedupre/pdf-to-text-ocr/examples/extract-text-from-pdf) returns readable text from a public PDF URL.
- Turn a public document into plain text with [**PDF to Text**](https://apify.com/maximedupre/pdf-to-text-ocr/examples/pdf-to-text).
- Read scanned PDF pages with [**Online OCR**](https://apify.com/maximedupre/pdf-to-text-ocr/examples/online-ocr) and review the returned text.
- Recover text from image-only pages with [**OCR PDF**](https://apify.com/maximedupre/pdf-to-text-ocr/examples/ocr-pdf).

#### 📚 PDF text, pages, and document details

Each saved dataset row describes one submitted PDF. It can include the source URL, file name, media type, file size, page count, extracted text, Markdown, text counts, page-level text, OCR details, PDF metadata, bookmarks, and page-tagged chunks.

#### ▶️ Run PDF text extraction

1. Add one or more public PDF URLs.
2. Optionally enter page numbers or ranges and choose whether to use OCR for scanned pages.
3. Run the Actor and open the PDF results dataset.

#### ⚙️ Input

**Input fields**

| Field | Type | What it does |
| --- | --- | --- |
| `pdfSources` | array of objects | Required. Adds one or more public PDF sources. |
| `pdfSources[].url` | string (URL) | Public URL of the PDF to read. |
| `pages` | string | Optional page numbers or ranges, separated by commas, such as `1,3-5`. Leave empty to read every page. |
| `useOcr` | boolean | Uses OCR for image-only or scanned pages when `true`. Set it to `false` to skip OCR and use normal PDF text extraction. Defaults to `true`. |

**Default input**

This is the smallest common input from a successful current-beta run:

```json
{
  "pdfSources": [
    {
      "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    }
  ],
  "useOcr": true
}
```

#### 🧾 Output

**Dataset fields**

| Field | Type | What it does |
| --- | --- | --- |
| `sourceUrl` | string (URL) | Public URL of the PDF read for this row. |
| `fileName` | string | File name found for the PDF, when available. |
| `mimeType` | string | Media type reported for the source file, when available. |
| `fileSizeBytes` | integer | Source PDF size in bytes, when available. |
| `pageCount` | integer | Number of pages in the source PDF. |
| `text` | string | Text read from the requested pages, including OCR text when OCR is used. |
| `markdown` | string | Markdown version of the extracted text with layout cues when available. |
| `characterCount` | integer | Character count for the requested content. |
| `wordCount` | integer | Word count for the requested content. |
| `pages` | array of objects | Text and counts for each requested page. |
| `pages[].pageNumber` | integer | Number of the page in the source PDF. |
| `pages[].text` | string | Text read from that page. |
| `pages[].characterCount` | integer | Character count for that page. |
| `pages[].wordCount` | integer | Word count for that page. |
| `ocr` | object | OCR coverage and signals for the row. |
| `ocr.pagesProcessed` | integer | Number of pages read with OCR. |
| `ocr.language` | string | OCR language, when the source reports it. |
| `ocr.confidence` | number | OCR confidence reported by the source, when available. |
| `ocr.remainingScannedPages` | integer | Scanned pages left without OCR text, when reported. |
| `metadata` | object | Metadata stored in the source PDF, when available. |
| `metadata.title` | string | Title stored in the PDF metadata. |
| `metadata.author` | string | Author stored in the PDF metadata. |
| `metadata.subject` | string | Subject stored in the PDF metadata. |
| `metadata.keywords` | string | Keywords stored in the PDF metadata. |
| `metadata.creator` | string | Application named as the PDF creator, when available. |
| `metadata.producer` | string | Application named as the PDF producer, when available. |
| `metadata.creationDate` | string | Creation date stored in the PDF metadata. |
| `metadata.modificationDate` | string | Modification date stored in the PDF metadata. |
| `bookmarks` | array of objects | Bookmark or table-of-contents entries found in the PDF. |
| `bookmarks[].title` | string | Title of the bookmark entry. |
| `bookmarks[].pageNumber` | integer | Page opened by the bookmark, when it has a destination. |
| `bookmarks[].level` | integer | Nesting level of the bookmark in the PDF outline. |
| `chunks` | array of objects | Page-tagged text chunks for embedding and citation work, when produced. |
| `chunks[].text` | string | Text in the chunk. |
| `chunks[].pageNumber` | integer | Page that the chunk belongs to. |

**Standard PDF row**

This complete row is from a successful current-beta run:

```json
{
  "sourceUrl": "https://www.w3.org/WAI/WCAG22/working-examples/pdf-bookmarks/bookmarks.pdf",
  "fileName": "bookmarks.pdf",
  "mimeType": "application/pdf",
  "fileSizeBytes": 106567,
  "pageCount": 2,
  "text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
  "markdown": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
  "characterCount": 776,
  "wordCount": 28,
  "pages": [
    {
      "pageNumber": 1,
      "text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
      "characterCount": 776,
      "wordCount": 28
    }
  ],
  "ocr": {
    "pagesProcessed": 0
  },
  "metadata": {
    "title": "Working example of creating bookmarks in PDF documents",
    "creator": "Acrobat PDFMaker 9.0 for Word",
    "producer": "Adobe PDF Library 9.0",
    "creationDate": "D:20110202105905-05'00'",
    "modificationDate": "D:20230718095527-07'00'"
  },
  "bookmarks": [
    {
      "title": "Header One",
      "pageNumber": 2,
      "level": 0
    },
    {
      "title": "Header Two for List",
      "pageNumber": 2,
      "level": 1
    },
    {
      "title": "Header Three Text",
      "pageNumber": 2,
      "level": 2
    },
    {
      "title": "Header Four",
      "pageNumber": 2,
      "level": 3
    },
    {
      "title": "Another Header One",
      "pageNumber": 2,
      "level": 0
    },
    {
      "title": "Another Header Two for List",
      "pageNumber": 2,
      "level": 1
    },
    {
      "title": "Links\r",
      "level": 0
    },
    {
      "title": "WCAG2.0",
      "pageNumber": 2,
      "level": 1
    }
  ],
  "chunks": [
    {
      "text": "1 Contents Header One ................................................................................................................................................... 2 Header Two for List ................................................................................................................................... 2 Header Three Text ................................................................................................................................ 2 Another Header One ..................................................................................................................................... 2 Another Header Two for List..................................................................................................................... 2",
      "pageNumber": 1
    }
  ]
}
```

**OCR PDF row**

This complete row is from a successful current-beta OCR run:

```json
{
  "sourceUrl": "https://raw.githubusercontent.com/tinytoolkit-org/pdf-sample-files/4fd733c47ead5446b06372ab910cb78d1c7da55e/out/sample-scanned.pdf",
  "fileName": "sample-scanned.pdf",
  "mimeType": "application/octet-stream",
  "fileSizeBytes": 432160,
  "pageCount": 3,
  "text": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data",
  "markdown": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data",
  "characterCount": 1219,
  "wordCount": 203,
  "pages": [
    {
      "pageNumber": 1,
      "text": "Scanned-style sample (image-only pages)\n\nPage 1 of 3\n\nThis document is a machine-generated sample file. It contains no real data: every name, number, and\nfigure on this page is placeholder content produced for software testing.\n\nUse it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without\nworrying about licensing or privacy. The file is dedicated to the public domain under CCO.\n\nText on this page is real, extractable text set in a standard font, not an image. A text-extraction tool\nshould recover these paragraphs verbatim, in reading order, with no OCR involved.\n\nPage boundaries matter for testing. Each page carries a heading with its own page number so that split,\nextract, reorder, and delete operations can be verified against what the output claims.\n\nIf a tool you are testing reports a different page count, a different order, or drops one of these\nparagraphs, the tool — not this file — is the thing to investigate next.\n\nA reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a\npage rotated sideways. The files in this collection cover each of those separately.\n\n©CO / public domain — generated sample, no real data",
      "characterCount": 1219,
      "wordCount": 203
    }
  ],
  "ocr": {
    "pagesProcessed": 1,
    "language": "eng",
    "confidence": 94,
    "remainingScannedPages": 0
  },
  "metadata": {
    "title": "Scanned-style sample",
    "author": "tinytoolkit sample files",
    "subject": "Image-only pages — no text layer, use for OCR testing",
    "keywords": "sample test pdf cc0 pdftoolskit.org",
    "creator": "https://pdftoolskit.org/sample-pdfs",
    "producer": "pdf-sample-files generator (pdf-lib)",
    "creationDate": "D:20260719000000Z",
    "modificationDate": "D:20260719000000Z"
  },
  "chunks": [
    {
      "text": "Scanned-style sample (image-only pages) Page 1 of 3 This document is a machine-generated sample file. It contains no real data: every name, number, and figure on this page is placeholder content produced for software testing. Use it to exercise PDF tooling — text extraction, page manipulation, rendering, compression — without worrying about licensing or privacy. The file is dedicated to the public domain under CCO. Text on this page is real, extractable text set in a standard font, not an image. A text-extraction tool should recover these paragraphs verbatim, in reading order, with no OCR involved. Page boundaries matter for testing. Each page carries a heading with its own page number so that split, extract, reorder, and delete operations can",
      "pageNumber": 1
    },
    {
      "text": "be verified against what the output claims. If a tool you are testing reports a different page count, a different order, or drops one of these paragraphs, the tool — not this file — is the thing to investigate next. A reasonable test suite checks the boring cases first: one page, ten pages, a page with nothing on it, a page rotated sideways. The files in this collection cover each of those separately. ©CO / public domain — generated sample, no real data",
      "pageNumber": 1
    }
  ]
}
```

#### 💳 Pricing

**How charges work**

The `PDF text` event is charged once for each PDF that is processed and returns extracted text. The `OCR page` event is charged for each scanned or image-only page successfully read with OCR and returned as text. Current tiered rates are shown in the Pricing tab.

#### 🔌 Integrations

Start runs in Apify Console or through the Apify API, then read the PDF results dataset in JSON, CSV, or Excel. Use the dataset in your own script or document workflow.

https://www.youtube.com/watch?v=bNACk1\_S\_6w\&list=PLObrtcm1Kw6MUrlLNDbK9QRg8VDJg0gOW\&index=4

#### ❓ FAQ

##### Can this read scanned PDFs?

Yes. Set `useOcr` to `true` to read image-only or scanned pages with OCR. The OCR details in each row show the pages processed and any other signals the source reports.

##### Can I extract only some pages?

Yes. Enter page numbers or ranges in `pages`, such as `1,3-5`. Leave it empty to read every page in each submitted PDF.

##### Can I submit more than one PDF?

Yes. Add one or more public PDF URL objects to `pdfSources`. The Actor returns a document result for each submitted source.

##### Can I upload a local or private PDF?

No. This input accepts public PDF URLs. It does not bypass login pages, paywalls, DRM, or other access restrictions.

##### Does it return tables or images as separate files?

No. The Actor returns extracted text and Markdown, plus document and page details. It does not extract structured table cells or image files.

### 📝 Changelog

**v0.0** (30-09-2026)

- Initial release.

### 🆘 Support

For issues, questions, or feature requests, [file a ticket](https://console.apify.com/actors/maximedupre~pdf-to-text-ocr/issues) and I'll fix or implement it in less than 24h 🫡

### 🔗 Related Actors

- [Webpage Text Extractor](https://apify.com/maximedupre/webpage-text-extractor) - Extract public webpage text and Markdown before or alongside PDF work.
- [Markdown to HTML Converter](https://apify.com/maximedupre/markdown-to-html-converter) - Turn returned Markdown into HTML for a page or document.
- [URL to BibTeX Converter](https://apify.com/maximedupre/url-to-bibtex-converter) - Create citation data from public paper and article URLs for research workflows.
- [PDF Text Extractor — Markdown, OCR & Metadata](https://apify.com/memo23/pdf-text-extractor) - Compare another PDF text workflow with Markdown, OCR, and metadata.
- [PDF Text Extractor - Multi-Column Layout, OCR & Metadata](https://apify.com/webdata_labs/pdf-text-extractor) - Compare a PDF text workflow with multi-column layout, OCR, and metadata.

**Made with ❤️ by Maxime Dupré**

# Actor input Schema

## `pdfSources` (type: `array`):

Add one or more public PDF URLs. The Actor returns extracted text and document details for each URL.

## `pages` (type: `string`):

Optional page numbers or ranges, separated by commas, such as 1,3-5. Leave empty to extract every page.

## `useOcr` (type: `boolean`):

When enabled, read image-only or scanned pages with OCR. Turn it off to skip OCR and return only text found by normal PDF extraction.

## Actor input object example

```json
{
  "pdfSources": [
    {
      "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    }
  ],
  "useOcr": true
}
```

# Actor output Schema

## `dataset` (type: `string`):

Open the extracted PDF results.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pdfSources": [
        {
            "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("maximedupre/pdf-to-text-ocr").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "pdfSources": [{ "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf" }] }

# Run the Actor and wait for it to finish
run = client.actor("maximedupre/pdf-to-text-ocr").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pdfSources": [
    {
      "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    }
  ]
}' |
apify call maximedupre/pdf-to-text-ocr --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,maximedupre/pdf-to-text-ocr"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kyLLx2SwGzkou4q8P/builds/hxY4YZcwLYNeDq27z/openapi.json
