Go to example tasks
OCR a Scanned PDF With Page-Level Provenance
Created by
Automation Lab
Render and OCR a public English scanned PDF, returning ordered text, page-level extraction methods, counts, and OCR confidence.
Document Text Extractorautomation-lab/layout-aware-text-extractor
Source URL
Type
Title
Extracted text
+8 fieldsTextNumberBooleanListObject
Input
Public document URLs(required)
url:https://raw.githubusercontent.com/ocrmypdf/OCRmyPDF/main/tests/resources/aspect.pdf
Maximum documents:1
Maximum PDF OCR pages:5
Include source metadata:true
Output fields
Source URL
Type
Title
Extracted text
Method
Pages
PDF pages
Words
OCR confidence
Warnings
Error
Extracted at
Sign up on Apify01
Create your Apify account to access the Document Text Extractor.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
