Go to example tasks
Find which PDFs are scans needing OCR
Audits a list of documents and keeps only the ones with no text layer, so you know exactly which files have to go through OCR before anything can read them. Documents that could not be reached at all are not charged.
PDF Extractor: Bulk PDF to Text, Tables & Markdown from a CSVnerolabs/dataset-pdf-extract
PDF URL
Status
Has text layer
Needs OCR
+10 fieldsTextNumberBooleanListObject
Input
PDF URLs:https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf+1
Which rows to keep:needs_ocr
Text output:none
Output fields
PDF URL
Status
Has text layer
Needs OCR
Pages
Words
Tables
Title
Author
Created
Size (KB)
Text
Tables
Detail
Sign up on Apify01
Create your Apify account to access the PDF Extractor: Bulk PDF to Text, Tables & Markdown from a CSV.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
