Go to example tasks
Crawl a Docs Subtree into Markdown Chunks
Crawl an authorized public docs or help subtree into Markdown pages with nested chunks for RAG. Replace the sandbox URL with your /docs or /help section and declare the correct rights basis. The first run returns bounded PAGE rows; later runs create fresh snapshots for index refreshes.
Website Content Crawler: Markdown for AI and RAGautoma-flow/website-content-crawler
Final source URL
Title
Markdown
Content hash
+1 fieldTextNumberBooleanListObject
Input
Authorized sources(required)
Start URL(required):https://books.toscrape.com/
Crawl scope:subtree
Rights basis(required):permission
Rights reference:https://books.toscrape.com/
Attribution:Scraping practice sandbox (books.toscrape.com)
Maximum pages:10
Maximum link depth:2
Include chunks:true
Chunk size (characters):6000
Output fields
Final source URL
Title
Markdown
Content hash
Semantic fingerprint
Sign up on Apify01
Create your Apify account to access the Website Content Crawler: Markdown for AI and RAG.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
