Web Page to Structured Plain Text Example
Created by
Stas Persiianenko
Convert a real standards page into clean plain text while retaining semantic headings, paragraphs, and lists for analysis or indexing.
Document Text Extractorautomation-lab/layout-aware-text-extractor
Source URL
Type
Title
Extracted text
+7 fieldsTextNumberBooleanListObject
Input
Public document URLs(required)
url:https://www.w3.org/WAI/standards-guidelines/wcag/
Maximum documents:1
Include source metadata:true
Output fields
Source URL
Type
Title
Extracted text
Method
Pages
Words
OCR confidence
Warnings
Error
Extracted at
Sign up on Apify01
Create your Apify account to access the Document Text Extractor.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
