RAG Web Browser avatar

RAG Web Browser

Pricing

$3.00 / 1,000 page crawleds

Go to Apify Store
RAG Web Browser

RAG Web Browser

Crawl any website into clean, LLM-ready markdown for RAG pipelines and AI agents. Handles JS pages, dedupes content. Pay per page crawled.

Pricing

$3.00 / 1,000 page crawleds

Rating

0.0

(0)

Developer

Travel Monitor Lab

Travel Monitor Lab

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Share

RAG Web Browser — crawl any website into clean, LLM-ready markdown

Crawl web pages and extract their content as clean markdown or plain text, optimized for RAG pipelines, LLM context windows and AI agents. Search Google and extract the top results, or pass direct URLs — handles JavaScript-rendered pages, boilerplate removal and content deduplication.

Use this tool when…

  • You're building a RAG pipeline and need clean page content without the nav/ads/footer noise.
  • Your AI agent needs to read web pages — via the Apify MCP server, Claude/GPT/Cursor/n8n can call it with zero glue code.
  • You want to turn a Google search into extracted content in one step: query in, clean markdown out.
  • You need multilingual content extraction (en, fr, de and more).

Features

  • Google search mode: pass a query, get the top results fetched and extracted in one run.
  • Direct URL mode: extract content from any list of page URLs.
  • LLM-optimized output: markdown (best for embeddings and LLM context) or plain text.
  • JavaScript rendering: handles modern JS-heavy sites that break naive fetchers.
  • Boilerplate stripping and deduplication: content, not chrome.
  • Language control for search results (en, fr, de…).
  • AI-agent ready: README and input schema readable by LLMs through MCP.

Use cases

RAG knowledge base ingestion

Feed clean markdown from documentation sites, blogs and help centers straight into your vector store.

Research automation for agents

Let an agent answer "what's the latest on X" by searching and reading the top sources itself.

Competitor and market monitoring

Extract and diff content from competitor pages, changelogs and pricing pages.

Content summarization pipelines

Pull articles into your LLM for summaries, translations or structured extraction.

Input

FieldTypeRequiredDescriptionExample
querystringone of the twoGoogle search query; top results are extracted"apify mcp server"
urlsarrayone of the twoDirect page URLs to extract["https://docs.apify.com"]
maxResultsintegernoMax search results / pages to process5
outputFormatstringnomarkdown (best for LLM/RAG) or text"markdown"
langstringnoLanguage code for search results"en"
{ "query": "apify mcp server", "maxResults": 5, "outputFormat": "markdown", "lang": "en" }

Output

One dataset item per page: url, title, content (markdown or text), format, extractedAt.

Pricing (pay per event)

EventPrice
page-crawled$0.003 per page

1,000 pages = $3.00. No subscription. Platform margin: Apify takes 20%. Tiered Apify plans (Bronze/Silver/Gold) get progressively lower prices.

Agentic payments (x402 / USDC)

This actor supports x402 agentic payments: AI agents without an Apify account can call it and pay per call in USDC on Base — verified live end-to-end.

FAQ

Markdown or plain text? Markdown for LLM/RAG (keeps headings and links structure); plain text for simple pipelines.

Does it handle JavaScript sites? Yes — pages are rendered before extraction.

Does it work with Claude, Cursor, GPT or n8n? Yes, via the Apify MCP server — it is the tool agents use to read the web.

Is it legal? It extracts publicly available pages. Respect robots policies and applicable laws.

More tools from this toolbox


Built and maintained by Adaga solutions — full-stack & AI agents in production, Luxembourg.