Go to example tasks
Convert a PDF to clean text and markdown
Downloads a PDF from a link and returns its text and markdown with real line breaks, per-page text, page count and metadata. Ready for search indexing or an LLM.
Document Text Extractor - PDF, DOCX & HTML to Text/Markdownclearfetch/document-text-extractor
Title
Type
Pages
Words
+2 fieldsTextNumberBooleanListObject
Input
Documents:https://arxiv.org/pdf/1706.03762
Maximum pages per document:15
Output fields
Title
Type
Pages
Words
Has text
Document
Sign up on Apify01
Create your Apify account to access the Document Text Extractor - PDF, DOCX & HTML to Text/Markdown.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
