Go to example tasks![Wikipedia Scraper [$0.7💰/1k] | RAG & AI Training](https://images.apifyusercontent.com/gg37N0QNbUV-DSIzXLR92d25XpDHlnpLaA9yBSnlxrY/rs:fill:250:250/cb:1/aHR0cHM6Ly9hcGlmeS1pbWFnZS11cGxvYWRzLXByb2QuczMudXMtZWFzdC0xLmFtYXpvbmF3cy5jb20vZHRSN3FuVGJRck5pU0NtajctYWN0b3ItRmcyMlQ1Yncxa0N2MExKOVctNmVZNXYzM0RWeC1pbWFnZXNfJTI4OSUyOS5qZmlm.webp)
Wikipedia RAG & AI Training Dataset
Extract clean Wikipedia article text for RAG pipelines and LLM fine-tuning datasets. Returns chunked, structured content ready for vectorization.
Wikipedia Scraper [$0.7💰/1k] | RAG & AI Trainingahmed_jasarevic/wikipedia-scraper
Title
URL
Page Title
Language
+1 fieldTextNumberBooleanListObject
Input
Start URLs(required)
url:https://en.wikipedia.org/wiki/Large_language_model+2
Extract Infobox:true
Extract Section Headings:true
Extract Article Text:true
Max Text Length per Section:50000
Wikipedia Language:en
Output fields
Title
URL
Page Title
Language
Crawled At
Sign up on Apify01
Create your Apify account to access the Wikipedia Scraper [$0.7💰/1k] | RAG & AI Training.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
