Go to example tasks
Extract text from scanned PDFs with OCR
Read scanned PDFs and images (TIFF, PNG, JPG) with Tesseract OCR and get the text as Markdown. Supports English, Spanish, German, French, Portuguese, Italian and Dutch. Native text pages are read directly and are not charged as OCR.
PDF to Markdown Extractor & Document Parser: Tables, JSON, OCRfguiraud/document-to-markdown-tables
Document
Status
Type
Pages
+5 fieldsTextNumberBooleanListObject
Input
Document URLs
url:https://raw.githubusercontent.com/ocrmypdf/OCRmyPDF/main/tests/resources/ccitt.pdf
OCR mode:auto
OCR languages:eng
What to return:markdown+1
Output fields
Document
Status
Type
Pages
Tables
OCR pages
Characters
Title
Error
Sign up on Apify01
Create your Apify account to access the PDF to Markdown Extractor & Document Parser: Tables, JSON, OCR.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
