PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots avatar

PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots

Pricing

from $4.00 / 1,000 ocr pages

Go to Apify Store
PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots

PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots

Extract text from PDFs and images. Digital PDF pages are read exactly; scanned pages, photos and screenshots go through OCR in 16 languages. Paste a link, a Google Drive or Dropbox share link, or upload a file. Pay per page, from $1 per 1,000 pages.

Pricing

from $4.00 / 1,000 ocr pages

Rating

0.0

(0)

Developer

clement

clement

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 hours ago

Last modified

Share

Extract clean text from PDFs and images. Paste a file link, a Google Drive or Dropbox share link, or upload a file.

  • Smart per page. Digital PDF pages already contain their text, so it is read exactly, with no OCR errors. Only scanned pages, photos and screenshots go through OCR.
  • You pay less for digital pages. $1 per 1,000 digital pages, $4 per 1,000 OCR pages. No subscription.
  • Mixed documents just work. A PDF with typed pages and scanned pages is handled page by page.
  • 16 languages for OCR: English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Swedish, Danish, Norwegian, Finnish, Czech, Turkish, Russian and Ukrainian.
  • Private by design. Files are processed inside your run. Nothing is sent to an outside AI service.
  • Fast on digital PDFs. A digital PDF takes a second or two. OCR takes a few seconds per page, depending on how dense the page is.

What you can provide

InputExample
PDF linkhttps://example.com/report.pdf
Image link (PNG, JPEG, TIFF, WEBP, BMP, GIF)https://example.com/receipt.jpg
Google Drive share linkhttps://drive.google.com/file/d/FILE_ID/view
Dropbox share linkhttps://www.dropbox.com/scl/fi/.../contract.pdf?rlkey=...
File uploadUse the Or upload a file field

Share links must be set to "anyone with the link can view". You can mix several links in one run.

How to use

  1. Get a link to your PDF or image: a direct file URL, or a Google Drive or Dropbox link shared with "anyone with the link". Or use Or upload a file.
  2. Paste the link into PDF or image URLs. Add as many as you like.
  3. If your documents are scanned and not in English, pick their languages in Languages for OCR.
  4. Click Start. When the run finishes, open the Output tab and download the text as JSON, CSV or Excel.

Output

One dataset item per document:

{
"inputUrl": "https://example.com/report.pdf",
"status": "ok",
"fileName": "report.pdf",
"fileType": "pdf",
"pageCount": 2,
"textPages": 1,
"ocrPages": 1,
"truncated": false,
"characters": 1164,
"text": "Address at Rice University on the Nation's Space Effort...",
"pages": [
{ "page": 1, "method": "text", "text": "Address at Rice University on the Nation's Space Effort..." },
{ "page": 2, "method": "ocr", "text": "Invoice No. 2026-0417..." }
]
}

method tells you how each page was read: text for the PDF's own text layer, ocr for recognised text. Files that cannot be processed produce an item with "status": "error" and an error message explaining why.

Options

OptionDefaultWhat it does
Languages for OCREnglishPick the languages printed in your scans. Fewer languages means faster, more accurate OCR.
When to use OCRAutomaticAutomatic runs OCR only on pages without a text layer. Always runs OCR on every page. Never reads text layers only.
Keep the page layoutOffKeeps columns and tables aligned with spaces, for digital PDF pages.
Maximum pages per document500Caps how much of a long document is processed.

Pricing

  • Digital page (text read from the PDF): $0.001
  • OCR page (scanned page or image): $0.004

Set Maximum pages per document or the run's maximum charge to cap spending. The Actor stops before exceeding either and marks the document with "truncated": true.

Use cases

  • Feed contracts, reports, invoices and research papers to an LLM or a search index (RAG).
  • Digitise scanned archives, receipts and letters.
  • Pull text out of screenshots and photos of documents.
  • Batch-convert a folder of PDFs to text through the API.

Limits

  • Printed text only. Handwriting is not recognised reliably.
  • OCR quality depends on the scan: sharp, straight pages at normal size work best; blurry or skewed photos give errors.
  • In Automatic mode, a page that has some typed text and a scanned image is read from its text layer only. Use Always for such pages.
  • Tables come out as text lines, not as structured rows and columns.
  • Password-protected PDFs cannot be read. Remove the password first.
  • Asian and right-to-left scripts are not supported for OCR yet.

FAQ

How do I extract text from a PDF?

Paste a link to the PDF or upload it, then start the run. You get the full text of the document and the text of each page.

How do I convert a scanned PDF to text?

The same way. The Actor notices that a page has no text of its own and runs OCR on it automatically. Pick the document's language in Languages for OCR for the best result.

How do I get text from an image, photo or screenshot?

Paste a link to the image (PNG, JPEG, TIFF, WEBP, BMP or GIF) or upload it. Images always go through OCR.

What is the difference between a digital page and an OCR page?

A digital page was created on a computer, and its text can be read from the file exactly. An OCR page is a picture of text, such as a scan or a photo, and its text has to be recognised, which is slower and can contain mistakes. The method field tells you which one each page was.

How accurate is the OCR?

Clean, straight scans of printed text come out very accurately. Small print, blurry photos, skewed pages, stamps and handwriting cause mistakes. Digital pages are always exact.

How long does it take and how much does it cost?

A 15-page digital PDF takes a few seconds and costs $0.015. OCR takes about 4 to 5 seconds for a dense page and costs $0.004 per page.

Does it extract tables?

Tables come out as lines of text. Turn on Keep the page layout to keep the columns aligned with spaces. The Actor does not return tables as structured rows and columns.

Can I use it from my own code, Make, Zapier or n8n?

Yes. Call it through the Apify API (see the example below), or connect it with Apify's integrations for Make, Zapier and n8n. Every run returns the same JSON.

What happens to my files?

They are downloaded and read inside your run, then deleted when it ends. Nothing is sent to an outside service, and the extracted text is stored only in your own Apify account.

Run it from the API

curl -X POST "https://api.apify.com/v2/acts/spokentext~pdf-image-to-text-ocr/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "urls": ["https://example.com/report.pdf"], "languages": ["en", "fr"] }'

For audio and video, use Audio & Video to Text.