# PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots (`spokentext/pdf-image-to-text-ocr`) Actor

Extract text from PDFs and images. Digital PDF pages are read exactly; scanned pages, photos and screenshots go through OCR in 16 languages. Paste a link, a Google Drive or Dropbox share link, or upload a file. Pay per page, from $1 per 1,000 pages.

- **URL**: https://apify.com/spokentext/pdf-image-to-text-ocr.md
- **Developed by:** [clement](https://apify.com/spokentext) (community)
- **Categories:** AI, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 ocr pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots

Extract clean text from **PDFs and images**. Paste a file link, a **Google Drive** or **Dropbox** share link, or **upload a file**.

- **Smart per page.** Digital PDF pages already contain their text, so it is read exactly, with no OCR errors. Only scanned pages, photos and screenshots go through OCR.
- **You pay less for digital pages.** $1 per 1,000 digital pages, $4 per 1,000 OCR pages. No subscription.
- **Mixed documents just work.** A PDF with typed pages and scanned pages is handled page by page.
- **16 languages for OCR:** English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Swedish, Danish, Norwegian, Finnish, Czech, Turkish, Russian and Ukrainian.
- **Private by design.** Files are processed inside your run. Nothing is sent to an outside AI service.
- **Fast on digital PDFs.** A digital PDF takes a second or two. OCR takes a few seconds per page, depending on how dense the page is.

### What you can provide

| Input | Example |
|---|---|
| PDF link | `https://example.com/report.pdf` |
| Image link (PNG, JPEG, TIFF, WEBP, BMP, GIF) | `https://example.com/receipt.jpg` |
| Google Drive share link | `https://drive.google.com/file/d/FILE_ID/view` |
| Dropbox share link | `https://www.dropbox.com/scl/fi/.../contract.pdf?rlkey=...` |
| File upload | Use the **Or upload a file** field |

Share links must be set to "anyone with the link can view". You can mix several links in one run.

### How to use

1. Get a link to your PDF or image: a direct file URL, or a Google Drive or Dropbox link shared with "anyone with the link". Or use **Or upload a file**.
2. Paste the link into **PDF or image URLs**. Add as many as you like.
3. If your documents are scanned and not in English, pick their languages in **Languages for OCR**.
4. Click **Start**. When the run finishes, open the **Output** tab and download the text as JSON, CSV or Excel.

### Output

One dataset item per document:

```json
{
    "inputUrl": "https://example.com/report.pdf",
    "status": "ok",
    "fileName": "report.pdf",
    "fileType": "pdf",
    "pageCount": 2,
    "textPages": 1,
    "ocrPages": 1,
    "truncated": false,
    "characters": 1164,
    "text": "Address at Rice University on the Nation's Space Effort...",
    "pages": [
        { "page": 1, "method": "text", "text": "Address at Rice University on the Nation's Space Effort..." },
        { "page": 2, "method": "ocr", "text": "Invoice No. 2026-0417..." }
    ]
}
```

`method` tells you how each page was read: `text` for the PDF's own text layer, `ocr` for recognised text. Files that cannot be processed produce an item with `"status": "error"` and an `error` message explaining why.

### Options

| Option | Default | What it does |
|---|---|---|
| Languages for OCR | English | Pick the languages printed in your scans. Fewer languages means faster, more accurate OCR. |
| When to use OCR | Automatic | **Automatic** runs OCR only on pages without a text layer. **Always** runs OCR on every page. **Never** reads text layers only. |
| Keep the page layout | Off | Keeps columns and tables aligned with spaces, for digital PDF pages. |
| Maximum pages per document | 500 | Caps how much of a long document is processed. |

### Pricing

- **Digital page** (text read from the PDF): $0.001
- **OCR page** (scanned page or image): $0.004

Set **Maximum pages per document** or the run's maximum charge to cap spending. The Actor stops before exceeding either and marks the document with `"truncated": true`.

### Use cases

- Feed contracts, reports, invoices and research papers to an LLM or a search index (RAG).
- Digitise scanned archives, receipts and letters.
- Pull text out of screenshots and photos of documents.
- Batch-convert a folder of PDFs to text through the API.

### Limits

- **Printed text only.** Handwriting is not recognised reliably.
- OCR quality depends on the scan: sharp, straight pages at normal size work best; blurry or skewed photos give errors.
- In Automatic mode, a page that has some typed text **and** a scanned image is read from its text layer only. Use **Always** for such pages.
- Tables come out as text lines, not as structured rows and columns.
- Password-protected PDFs cannot be read. Remove the password first.
- Asian and right-to-left scripts are not supported for OCR yet.

### FAQ

#### How do I extract text from a PDF?

Paste a link to the PDF or upload it, then start the run. You get the full text of the document and the text of each page.

#### How do I convert a scanned PDF to text?

The same way. The Actor notices that a page has no text of its own and runs OCR on it automatically. Pick the document's language in **Languages for OCR** for the best result.

#### How do I get text from an image, photo or screenshot?

Paste a link to the image (PNG, JPEG, TIFF, WEBP, BMP or GIF) or upload it. Images always go through OCR.

#### What is the difference between a digital page and an OCR page?

A digital page was created on a computer, and its text can be read from the file exactly. An OCR page is a picture of text, such as a scan or a photo, and its text has to be recognised, which is slower and can contain mistakes. The `method` field tells you which one each page was.

#### How accurate is the OCR?

Clean, straight scans of printed text come out very accurately. Small print, blurry photos, skewed pages, stamps and handwriting cause mistakes. Digital pages are always exact.

#### How long does it take and how much does it cost?

A 15-page digital PDF takes a few seconds and costs $0.015. OCR takes about 4 to 5 seconds for a dense page and costs $0.004 per page.

#### Does it extract tables?

Tables come out as lines of text. Turn on **Keep the page layout** to keep the columns aligned with spaces. The Actor does not return tables as structured rows and columns.

#### Can I use it from my own code, Make, Zapier or n8n?

Yes. Call it through the Apify API (see the example below), or connect it with Apify's integrations for Make, Zapier and n8n. Every run returns the same JSON.

#### What happens to my files?

They are downloaded and read inside your run, then deleted when it ends. Nothing is sent to an outside service, and the extracted text is stored only in your own Apify account.

### Run it from the API

```bash
curl -X POST "https://api.apify.com/v2/acts/spokentext~pdf-image-to-text-ocr/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "urls": ["https://example.com/report.pdf"], "languages": ["en", "fr"] }'
```

### Related

For audio and video, use [Audio & Video to Text](https://apify.com/spokentext/audio-video-to-text).

# Actor input Schema

## `urls` (type: `array`):

Direct links to PDF or image files (PNG, JPEG, TIFF, WEBP, BMP, GIF), or Google Drive and Dropbox share links. Share links must be open to anyone with the link.

## `file` (type: `string`):

Upload one PDF or image from your computer instead of pasting a link.

## `languages` (type: `array`):

Languages printed in your scanned pages and images. Pick only the ones you need: each extra language makes OCR slower. Not used for digital PDF pages.

## `ocrMode` (type: `string`):

Digital PDF pages already contain their text, which is read exactly and costs less. OCR is needed for scans, photos and screenshots.

## `preserveLayout` (type: `boolean`):

Keeps columns and tables aligned with spaces, as they appear on the page. Leave off for clean, flowing text. Applies to digital PDF pages.

## `includePages` (type: `boolean`):

Adds a list with the text of each page and how it was read (text layer or OCR), besides the full text.

## `maxPagesPerDocument` (type: `integer`):

Longer documents are cut at this many pages, so a single run can never cost more than you expect.

## Actor input object example

```json
{
  "urls": [
    "https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/sample-text-and-scanned.pdf?signature=1oxjqoBQ1rS8GvxN9TuBV"
  ],
  "languages": [
    "en"
  ],
  "ocrMode": "auto",
  "preserveLayout": false,
  "includePages": true,
  "maxPagesPerDocument": 500
}
```

# Actor output Schema

## `documents` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/sample-text-and-scanned.pdf?signature=1oxjqoBQ1rS8GvxN9TuBV"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("spokentext/pdf-image-to-text-ocr").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/sample-text-and-scanned.pdf?signature=1oxjqoBQ1rS8GvxN9TuBV"] }

# Run the Actor and wait for it to finish
run = client.actor("spokentext/pdf-image-to-text-ocr").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://api.apify.com/v2/key-value-stores/kyR7mGkrVQ6H3xQe8/records/sample-text-and-scanned.pdf?signature=1oxjqoBQ1rS8GvxN9TuBV"
  ]
}' |
apify call spokentext/pdf-image-to-text-ocr --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,spokentext/pdf-image-to-text-ocr"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QOrzMD1n3mIZwRM67/builds/TXR4c8m9dqSgKd6pQ/openapi.json
