Go to example tasks
PDF text and tables as Markdown for an LLM
Returns each document's text with its tables appended as proper Markdown tables, which is the friendliest shape to hand to a language model. Pair it with Dataset AI Enrich to turn the text into typed columns.
PDF Extractor: Bulk PDF to Text, Tables & Markdown from a CSVnerolabs/dataset-pdf-extract
PDF URL
Status
Has text layer
Needs OCR
+10 fieldsTextNumberBooleanListObject
Input
PDF URLs:https://arxiv.org/pdf/1706.03762
Text output:markdown
Include per-page text:true
Output fields
PDF URL
Status
Has text layer
Needs OCR
Pages
Words
Tables
Title
Author
Created
Size (KB)
Text
Tables
Detail
Sign up on Apify01
Create your Apify account to access the PDF Extractor: Bulk PDF to Text, Tables & Markdown from a CSV.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
