日本語ドキュメントをRAG用Markdownに変換
Created by
nezha
日本語のドキュメントサイトをクロールし、RAG、ベクトルDB、AI検索、社内ナレッジベースに使いやすいMarkdownとテキストを抽出します。
Japanese Website Content Crawler for RAGnezha/website-content-crawler-japan
Url
Title
Description
Content format
+6 fieldsTextNumberBooleanListObject
Input
日本語サイト、ドキュメント、ヘルプセンターのURL(required)
url:https://www.digital.go.jp/resources/introduction-to-web-accessibility-guidebook
抽出する最大ページ数:5
ページ探索方法:auto
リンクの深さ:1
対象範囲だけをクロール:true
メインコンテンツ形式:markdown
Output fields
Url
Title
Description
Content format
Word count
Language
Canonical url
Depth
Http status code
Crawled at
Sign up on Apify01
Create your Apify account to access the Japanese Website Content Crawler for RAG.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
