PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots
Pricing
from $4.00 / 1,000 ocr pages
PDF & Image to Text: OCR for Scanned PDFs, Photos & Screenshots
Extract text from PDFs and images. Digital PDF pages are read exactly; scanned pages, photos and screenshots go through OCR in 16 languages. Paste a link, a Google Drive or Dropbox share link, or upload a file. Pay per page, from $1 per 1,000 pages.
Pricing
from $4.00 / 1,000 ocr pages
Rating
0.0
(0)
Developer
clement
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 hours ago
Last modified
Categories
Share
Extract clean text from PDFs and images. Paste a file link, a Google Drive or Dropbox share link, or upload a file.
- Smart per page. Digital PDF pages already contain their text, so it is read exactly, with no OCR errors. Only scanned pages, photos and screenshots go through OCR.
- You pay less for digital pages. $1 per 1,000 digital pages, $4 per 1,000 OCR pages. No subscription.
- Mixed documents just work. A PDF with typed pages and scanned pages is handled page by page.
- 16 languages for OCR: English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Swedish, Danish, Norwegian, Finnish, Czech, Turkish, Russian and Ukrainian.
- Private by design. Files are processed inside your run. Nothing is sent to an outside AI service.
- Fast on digital PDFs. A digital PDF takes a second or two. OCR takes a few seconds per page, depending on how dense the page is.
What you can provide
| Input | Example |
|---|---|
| PDF link | https://example.com/report.pdf |
| Image link (PNG, JPEG, TIFF, WEBP, BMP, GIF) | https://example.com/receipt.jpg |
| Google Drive share link | https://drive.google.com/file/d/FILE_ID/view |
| Dropbox share link | https://www.dropbox.com/scl/fi/.../contract.pdf?rlkey=... |
| File upload | Use the Or upload a file field |
Share links must be set to "anyone with the link can view". You can mix several links in one run.
How to use
- Get a link to your PDF or image: a direct file URL, or a Google Drive or Dropbox link shared with "anyone with the link". Or use Or upload a file.
- Paste the link into PDF or image URLs. Add as many as you like.
- If your documents are scanned and not in English, pick their languages in Languages for OCR.
- Click Start. When the run finishes, open the Output tab and download the text as JSON, CSV or Excel.
Output
One dataset item per document:
{"inputUrl": "https://example.com/report.pdf","status": "ok","fileName": "report.pdf","fileType": "pdf","pageCount": 2,"textPages": 1,"ocrPages": 1,"truncated": false,"characters": 1164,"text": "Address at Rice University on the Nation's Space Effort...","pages": [{ "page": 1, "method": "text", "text": "Address at Rice University on the Nation's Space Effort..." },{ "page": 2, "method": "ocr", "text": "Invoice No. 2026-0417..." }]}
method tells you how each page was read: text for the PDF's own text layer, ocr for recognised text. Files that cannot be processed produce an item with "status": "error" and an error message explaining why.
Options
| Option | Default | What it does |
|---|---|---|
| Languages for OCR | English | Pick the languages printed in your scans. Fewer languages means faster, more accurate OCR. |
| When to use OCR | Automatic | Automatic runs OCR only on pages without a text layer. Always runs OCR on every page. Never reads text layers only. |
| Keep the page layout | Off | Keeps columns and tables aligned with spaces, for digital PDF pages. |
| Maximum pages per document | 500 | Caps how much of a long document is processed. |
Pricing
- Digital page (text read from the PDF): $0.001
- OCR page (scanned page or image): $0.004
Set Maximum pages per document or the run's maximum charge to cap spending. The Actor stops before exceeding either and marks the document with "truncated": true.
Use cases
- Feed contracts, reports, invoices and research papers to an LLM or a search index (RAG).
- Digitise scanned archives, receipts and letters.
- Pull text out of screenshots and photos of documents.
- Batch-convert a folder of PDFs to text through the API.
Limits
- Printed text only. Handwriting is not recognised reliably.
- OCR quality depends on the scan: sharp, straight pages at normal size work best; blurry or skewed photos give errors.
- In Automatic mode, a page that has some typed text and a scanned image is read from its text layer only. Use Always for such pages.
- Tables come out as text lines, not as structured rows and columns.
- Password-protected PDFs cannot be read. Remove the password first.
- Asian and right-to-left scripts are not supported for OCR yet.
FAQ
How do I extract text from a PDF?
Paste a link to the PDF or upload it, then start the run. You get the full text of the document and the text of each page.
How do I convert a scanned PDF to text?
The same way. The Actor notices that a page has no text of its own and runs OCR on it automatically. Pick the document's language in Languages for OCR for the best result.
How do I get text from an image, photo or screenshot?
Paste a link to the image (PNG, JPEG, TIFF, WEBP, BMP or GIF) or upload it. Images always go through OCR.
What is the difference between a digital page and an OCR page?
A digital page was created on a computer, and its text can be read from the file exactly. An OCR page is a picture of text, such as a scan or a photo, and its text has to be recognised, which is slower and can contain mistakes. The method field tells you which one each page was.
How accurate is the OCR?
Clean, straight scans of printed text come out very accurately. Small print, blurry photos, skewed pages, stamps and handwriting cause mistakes. Digital pages are always exact.
How long does it take and how much does it cost?
A 15-page digital PDF takes a few seconds and costs $0.015. OCR takes about 4 to 5 seconds for a dense page and costs $0.004 per page.
Does it extract tables?
Tables come out as lines of text. Turn on Keep the page layout to keep the columns aligned with spaces. The Actor does not return tables as structured rows and columns.
Can I use it from my own code, Make, Zapier or n8n?
Yes. Call it through the Apify API (see the example below), or connect it with Apify's integrations for Make, Zapier and n8n. Every run returns the same JSON.
What happens to my files?
They are downloaded and read inside your run, then deleted when it ends. Nothing is sent to an outside service, and the extracted text is stored only in your own Apify account.
Run it from the API
curl -X POST "https://api.apify.com/v2/acts/spokentext~pdf-image-to-text-ocr/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{ "urls": ["https://example.com/report.pdf"], "languages": ["en", "fr"] }'
Related
For audio and video, use Audio & Video to Text.