Mixed Document Text Extraction Workflow
Created by
Stas Persiianenko
Turn a PDF, text image, and web page into consistent typed dataset records for downstream search, RAG, spreadsheet, or ETL workflows.
Document Text Extractorautomation-lab/layout-aware-text-extractor
Source URL
Type
Title
Extracted text
+7 fieldsTextNumberBooleanListObject
Input
Public document URLs(required)
url:https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf+2
Maximum documents:3
Maximum source size (MB):20
Request timeout (seconds):30
Include source metadata:true
Output fields
Source URL
Type
Title
Extracted text
Method
Pages
Words
OCR confidence
Warnings
Error
Extracted at
Sign up on Apify01
Create your Apify account to access the Document Text Extractor.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
