Go to example tasks
Crawl One Authorized Page to Markdown
Extract one authorized public HTML page into clean Markdown with source-linked chunks. Replace the sandbox URL with a page you may process and declare the correct rights basis. Use the PAGE row in your RAG or docs pipeline; run again after the page changes.
Website Content Crawler: Markdown for AI and RAGautoma-flow/website-content-crawler
Final source URL
Title
Markdown
Content hash
+1 fieldTextNumberBooleanListObject
Input
Authorized sources(required)
Start URL(required):https://books.toscrape.com/
Crawl scope:page
Rights basis(required):permission
Rights reference:https://books.toscrape.com/
Attribution:Scraping practice sandbox (books.toscrape.com)
Include chunks:true
Chunk size (characters):6000
Output fields
Final source URL
Title
Markdown
Content hash
Semantic fingerprint
Sign up on Apify01
Create your Apify account to access the Website Content Crawler: Markdown for AI and RAG.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
