Image & PDF OCR — Image to Text and Scanned PDF to Text
Pricing
from $2.00 / 1,000 page ocreds
Image & PDF OCR — Image to Text and Scanned PDF to Text
Extract text from images and scanned PDFs with Tesseract OCR running inside the Actor: your files are never sent to a third-party OCR service. 16 languages, text per page with confidence scores, optional word boxes. $2 per 1,000 pages; blank pages and failed files are free.
Pricing
from $2.00 / 1,000 page ocreds
Rating
0.0
(0)
Developer
KeyMan98
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Image & PDF OCR
Extract text from images and PDFs — including scanned, image-only documents with no text layer at all — using Tesseract OCR. Give it public URLs and get back the recognized text for every page, with a confidence score and word/character counts. Nothing is installed on your side, and no file is ever sent to a third-party OCR or vision API: recognition runs entirely inside this Actor's own container.
What you get (output fields)
For each page — a plain image counts as one page — one dataset row with:
url— the file URL you requested.fileType— detected type:png,jpeg,tiff,webp,bmp,gif, orpdf. Detected from the file's own bytes, never from the URL's extension.page— 1-based page/frame number within the document.pageCount— total pages/frames in the document, even ifmaxPagesPerDocumentlimited how many were actually processed.text— the OCR text recognized on this page.meanConfidence— Tesseract's mean word confidence for this page, 0-100. Null if no word was recognized on the page at all.wordCount/charCount— size of the recognized text.languages— language(s) Tesseract used for this page.width/height— pixel size of the image actually OCRed (the rendered PDF page, or the source image).wordBoxes— only when the Include word boxes input is on: every recognized word with its bounding box and confidence.error— null on success, or a message explaining why this page/document failed.
Who it's for
- Anyone with scanned paperwork — old letters, forms, receipts, faxed contracts — who needs the text out of them without retyping.
- Developers who need OCR in a pipeline (search indexing, data extraction, archiving) without standing up Tesseract themselves.
- Anyone who wants OCR without sending private documents to Google Cloud Vision, AWS Textract, or a similar third-party API — see "Where does OCR actually happen?" below.
Where does OCR actually happen?
Everything runs inside this Actor's own container: PDF pages are rasterized with pypdfium2 (the same rendering engine behind Chrome's built-in PDF viewer), and the resulting page images are recognized with Tesseract (the open-source OCR engine originally developed at HP and now maintained by Google), called through pytesseract. Your files are never uploaded to Google Cloud Vision, AWS Textract, or any other third-party OCR/vision service — they exist only for the lifetime of the run, inside this Actor.
This Actor deliberately does not use PyMuPDF for PDF rendering, even though it's a common choice: PyMuPDF is licensed AGPL, which would require this Actor's own source to be AGPL too (or a commercial license purchased from its maker). pypdfium2 (Apache/BSD) and Tesseract (Apache 2.0) avoid that entirely.
How to use
- List your file URLs — public PNG, JPEG, TIFF (including multi-page), WebP, BMP, GIF, or PDF links.
- Pick the language(s) the text is in (default: English). Pick more than one if a page mixes languages.
- Run the Actor. Each page becomes one row, with its recognized text and confidence.
- Optionally adjust PDF render DPI (sharper but slower for small print), Max pages per document, Page segmentation mode (for a single line/word instead of a full page), or turn on Include word boxes if you need per-word positions.
Input example (JSON)
{"urls": ["https://example.com/scanned-letter.jpg","https://example.com/scanned-report.pdf"],"languages": ["eng"],"maxPagesPerDocument": 20,"dpi": 300,"includeWordBoxes": false,"pageSegmentationMode": 3,"maxFileSizeMb": 50}
Output example (JSON)
{"url": "https://example.com/scanned-report.pdf","fileType": "pdf","page": 1,"pageCount": 2,"text": "TOP SECRET/SENSITIVE\n\nTALKING POINTS\n\n1. AS WE DISCUSSED BEFORE MY MIDDLE EAST TRIP, I PROPOSED TO\nPRESIDENT SADAT, PRIME MINISTER BEGIN AND CROWN PRINCE FAHD...","meanConfidence": 94.32,"wordCount": 187,"charCount": 1204,"languages": ["eng"],"width": 2550,"height": 3301,"wordBoxes": null,"error": null}
If a page or file fails
If a URL can't be downloaded, isn't one of the supported file types, or is an encrypted/corrupted document, that URL gets one error row (page and pageCount null, error explaining why) — the rest of your list still runs. If a document opens fine but one specific page inside it fails to render or to OCR, only that page becomes an error row; the other pages of the same document are unaffected. No row that has error set is ever charged.
Pricing
Pay only for pages where text was actually found — nothing charged for blank pages or for any row with an error. Pricing model: pay-per-event.
| Event | When it's charged | Price |
|---|---|---|
page-ocr | a page/frame produced non-empty OCR text | 0.002 USD ($2 per 1,000 pages) |
Limitations
- Printed text works well; handwriting does not. Tesseract is a printed-text OCR engine — cursive or handwritten pages will come back with low confidence and garbled text, not a clear error. This is a real limit of the engine, not a bug.
- Quality is below commercial OCR services (Google Lens, ABBYY, cloud vision APIs) on hard cases: low-resolution scans, unusual fonts, heavy skew, or dense multi-column layouts. Tesseract is free and runs locally; those trade-offs are the cost of that.
- No automatic upright-rotation for photos without EXIF data. A photo's EXIF orientation tag is corrected automatically, and a PDF page renders with its own declared rotation — but a scanned image with no EXIF tag that is genuinely sideways or upside-down is OCRed as-is and will return poor results. Rotate it before submitting if you know it's misoriented.
maxPagesPerDocumentandmaxFileSizeMbare hard cutoffs, not smart sampling: a 100-page PDF with the default 20-page cap only returns its first 20 pages (pageCountstill reports the true total).- Multi-page OCR is TIFF-only. A multi-page TIFF becomes one row per page, like a PDF. An animated GIF or WebP is always read as a single still frame (its first one) —
pageCountis 1 even if the file has more; this Actor treats a GIF/WebP as one image, not a document, so an animation is never charged once per frame. - Very large pages are capped before OCR to keep memory and run time predictable: a PDF page is rendered at a reduced DPI if needed so it never exceeds a 40-megapixel render; an image over roughly 100 megapixels is rejected outright as an error row rather than processed.
- Mixing many unrelated languages in one request can reduce accuracy — Tesseract tries to match text against all requested languages' dictionaries at once.
FAQ
Does this send my files to Google Cloud Vision, AWS Textract, or another OCR API?
No. OCR runs inside this Actor with the open-source Tesseract engine. See "Where does OCR actually happen?" above.
Can it read handwriting?
Not reliably. Tesseract is built and trained for printed/typed text; handwritten pages will get low confidence scores and often garbled output.
What's the difference between submitting an image vs. a PDF?
None from your side — list either kind of URL, mixed freely, in the same run. Internally, a PDF page is rendered to an image at the DPI you choose before OCR; a plain image is read at its native resolution.
Can I OCR a scanned PDF that has no text layer at all?
Yes — that's exactly the case this Actor is built for. Every PDF page is rasterized to an image and OCRed from scratch; this Actor never tries to reuse a PDF's existing embedded text layer (if it has one), so it works identically whether the PDF has one or not.
Why is a page free even though the run succeeded?
Pages with no recognizable text (a blank page, a mostly-graphical page) are not charged — you only pay for pages Tesseract actually extracted text from.
Does it support multi-page TIFF files?
Yes. Each page of a multi-page TIFF becomes its own dataset row, the same as a PDF's pages.
What about an animated GIF or WebP — does every frame get OCRed?
No, only the first frame. Multi-page OCR is offered for TIFF specifically; a GIF/WebP is treated as one image, so you're never charged once per frame of an animation.
Can I get the position of each word, not just the page text?
Yes — turn on Include word boxes in the input. Each row then also includes a wordBoxes array with every word's text, confidence, and pixel bounding box.
Can I use this through the Apify API or an MCP server?
Yes, like any Apify Actor — through the standard Apify API, or through the Apify MCP server if you use Claude, Cursor, or another MCP-enabled client.
Export
Results can be downloaded from the Apify dataset as JSON, CSV, or Excel, or accessed via the Apify API.