Go to example tasks
Convert a PDF to Markdown chunks for RAG
Created by
Argentin Vazdautan
Convert a PDF (here the 'Attention Is All You Need' paper) into clean Markdown split into chunks by heading, with page numbers and heading paths as metadata, ready to embed into a vector database for RAG. Replace the link with your own PDF, Word, PowerPoint or Excel files, or point it at a whole list of documents.
Document to Markdown for AI & RAGegra_van/document-to-markdown
File
Type
Chunk
Page from
+7 fieldsTextNumberBooleanListObject
Input
Document URLs
url:https://arxiv.org/pdf/1706.03762
Dataset URL field:url
Output format:markdown
Chunking:heading
Chunk size:800
Chunk size unit:tokens
Chunk overlap:100
Insert page markers:true
Include PowerPoint speaker notes:true
OCR scanned pages:false
OCR languages:eng
Max pages per document:0
Max documents:1000
Max file size (MB):100
Output fields
File
Type
Chunk
Page from
Page to
Section
Tokens
Text
Full Markdown
Status
Error
Sign up on Apify01
Create your Apify account to access the Document to Markdown for AI & RAG.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
