RAG Web Browser
Pricing
$3.00 / 1,000 page crawleds
RAG Web Browser
Crawl any website into clean, LLM-ready markdown for RAG pipelines and AI agents. Handles JS pages, dedupes content. Pay per page crawled.
Pricing
$3.00 / 1,000 page crawleds
Rating
0.0
(0)
Developer
Travel Monitor Lab
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
RAG Web Browser — crawl any website into clean, LLM-ready markdown
Crawl web pages and extract their content as clean markdown or plain text, optimized for RAG pipelines, LLM context windows and AI agents. Search Google and extract the top results, or pass direct URLs — handles JavaScript-rendered pages, boilerplate removal and content deduplication.
Use this tool when…
- You're building a RAG pipeline and need clean page content without the nav/ads/footer noise.
- Your AI agent needs to read web pages — via the Apify MCP server, Claude/GPT/Cursor/n8n can call it with zero glue code.
- You want to turn a Google search into extracted content in one step: query in, clean markdown out.
- You need multilingual content extraction (en, fr, de and more).
Features
- Google search mode: pass a query, get the top results fetched and extracted in one run.
- Direct URL mode: extract content from any list of page URLs.
- LLM-optimized output: markdown (best for embeddings and LLM context) or plain text.
- JavaScript rendering: handles modern JS-heavy sites that break naive fetchers.
- Boilerplate stripping and deduplication: content, not chrome.
- Language control for search results (en, fr, de…).
- AI-agent ready: README and input schema readable by LLMs through MCP.
Use cases
RAG knowledge base ingestion
Feed clean markdown from documentation sites, blogs and help centers straight into your vector store.
Research automation for agents
Let an agent answer "what's the latest on X" by searching and reading the top sources itself.
Competitor and market monitoring
Extract and diff content from competitor pages, changelogs and pricing pages.
Content summarization pipelines
Pull articles into your LLM for summaries, translations or structured extraction.
Input
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
query | string | one of the two | Google search query; top results are extracted | "apify mcp server" |
urls | array | one of the two | Direct page URLs to extract | ["https://docs.apify.com"] |
maxResults | integer | no | Max search results / pages to process | 5 |
outputFormat | string | no | markdown (best for LLM/RAG) or text | "markdown" |
lang | string | no | Language code for search results | "en" |
{ "query": "apify mcp server", "maxResults": 5, "outputFormat": "markdown", "lang": "en" }
Output
One dataset item per page: url, title, content (markdown or text), format, extractedAt.
Pricing (pay per event)
| Event | Price |
|---|---|
page-crawled | $0.003 per page |
1,000 pages = $3.00. No subscription. Platform margin: Apify takes 20%. Tiered Apify plans (Bronze/Silver/Gold) get progressively lower prices.
Agentic payments (x402 / USDC)
This actor supports x402 agentic payments: AI agents without an Apify account can call it and pay per call in USDC on Base — verified live end-to-end.
FAQ
Markdown or plain text? Markdown for LLM/RAG (keeps headings and links structure); plain text for simple pipelines.
Does it handle JavaScript sites? Yes — pages are rendered before extraction.
Does it work with Claude, Cursor, GPT or n8n? Yes, via the Apify MCP server — it is the tool agents use to read the web.
Is it legal? It extracts publicly available pages. Respect robots policies and applicable laws.
More tools from this toolbox
Built and maintained by Adaga solutions — full-stack & AI agents in production, Luxembourg.