Website Content Crawler API: Web Pages to Markdown for RAG
Pricing
from $1.06 / 1,000 pages
Website Content Crawler API: Web Pages to Markdown for RAG
Website content crawler API: turn any web page or whole site into clean Markdown and text for RAG, LLMs and AI agents. Crawls links and sitemaps, renders JavaScript pages when needed. RAG Web Browser alternative at $1 per 1,000 pages.
Pricing
from $1.06 / 1,000 pages
Rating
0.0
(0)
Developer
Saulius AutomatesIT
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
Turn any web page, docs section or whole website into clean Markdown and text for RAG pipelines, LLM context, vector databases and AI agents. Give it URLs; it fetches them, strips navigation, footers, scripts and cookie banners, keeps the main content, and returns Markdown with headings, lists, links, code blocks and tables. Crawl links and sitemaps when you need a whole site. $1.25 per 1,000 pages, failed and empty pages free.
A cheaper alternative to RAG Web Browser ($2.28 per 1,000 fetched pages) and to compute-billed website crawlers: you pay a flat price per page with text, so a big crawl costs what you expect.
What you get per page
| Field | Example |
|---|---|
url, requestedUrl, statusCode | final URL after redirects, the URL you gave, 200 |
title, description | page title and meta description |
markdown | the page as clean Markdown: # Title, ## Sections, lists, links, fenced code, tables |
language, author, publishedAt | from the page's meta tags and JSON-LD |
canonicalUrl, image | canonical link and share image |
wordCount | words in the text |
mainContentOnly | true when only the main article or docs block was kept |
depth, referrerUrl, startUrl | where the crawler found the page |
text, html, links | optional: plain text, the cleaned HTML, every link on the page |
Input
- Start URLs: pages to read. A bare domain (
apify.com) works. - Crawl depth: 0 (default) reads only your URLs. 1 also reads the pages they link to, and so on. Links
are followed on the same site and inside the start URL's folder:
https://docs.apify.com/platform/actorscrawlshttps://docs.apify.com/platform/.... - Use sitemaps: add every page in the site's
sitemap.xml(and sitemaps listed inrobots.txt) inside the start folder. The fastest way to get a whole docs site. - Max pages per start URL (100) and max pages in total.
- Only URLs containing / Skip URLs containing: simple text filters such as
/blog/or?page=. - Main content only (default on): Mozilla Readability keeps the article or docs body on content pages; home pages and listings keep everything except navigation, footers, scripts and cookie banners.
- Remove elements: extra CSS selectors to drop, such as
.sidebar, #comments. - Include plain text / cleaned HTML / links: extra fields, same price.
{"startUrls": ["https://docs.apify.com/platform/"],"maxCrawlDepth": 2,"useSitemaps": true,"maxPagesPerStartUrl": 500}
Pricing
$1.25 per 1,000 pages with text ($0.00125 a page), plus Apify's $0.00005 Actor start. Apify Store discounts apply: Bronze $1.19, Silver $1.13, Gold and above $1.06 per 1,000. No charge for pages that fail, are not found, are blocked, are not HTML (PDF, images, files), have no text or only show text after JavaScript runs.
Examples: a 300 page docs site = $0.38. 100,000 blog posts for a vector database = $125. One page for an AI agent = $0.0013.
Speed: a run gets one CPU core per 4 GB of memory. The default 1 GB reads about 35 pages a minute; give big crawls 4 GB (about 4 times faster). The price per page is the same at any memory.
Use cases
- RAG and vector databases: load documentation, help centres, blogs and knowledge bases as Markdown chunks with titles and source URLs.
- AI agents and MCP: fetch a page as Markdown inside an agent loop; small, cheap, predictable.
- LLM training and evaluation data: clean text of whole sites with language and dates.
- Content monitoring: schedule a crawl of a competitor's blog or a docs section and diff the Markdown.
- SEO audits: titles, descriptions, canonical links, word counts and internal links per page.
Error rows (free)
error | Meaning |
|---|---|
NOT_FOUND | the page answered 404 or 410 |
HTTP_403, HTTP_429 and so on | the site refused the page |
BLOCKED | a bot check (Cloudflare challenge, captcha) |
NOT_HTML | a PDF, image, archive or other file |
NEEDS_JAVASCRIPT | the server sends an empty app shell; the text appears only after JavaScript runs |
EMPTY | the page has no text |
FAILED | the site could not be reached |
NO_DATA | the input had no usable URL |
API, MCP and integrations
Call it from Python, JavaScript, Make, Zapier, n8n, LangChain or LlamaIndex like any Apify Actor, or give
an AI agent this MCP server: https://mcp.apify.com/?tools=sauliusautomatesit/website-content-crawler-api.
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("sauliusautomatesit/website-content-crawler-api").call(run_input={"startUrls": ["https://docs.apify.com/platform/"], "maxCrawlDepth": 1, "maxPagesPerStartUrl": 50})for page in client.dataset(run["defaultDatasetId"]).iterate_items():if page.get("type") != "error":print(page["url"], page["wordCount"])
Notes
- Pages are read as the server sends them (plain HTTP with a Chrome fingerprint through Apify datacenter
proxies). Server rendered sites (docs, blogs, news, Wikipedia, most CMSs and Next.js sites) work; client
only app shells come back as free
NEEDS_JAVASCRIPTrows. - The crawler stays on the start URL's site and folder. Use several start URLs for several sections.
- Respect each site's terms and robots rules for your use; this Actor reads only public pages.
Related
- Contact Details Scraper API: emails, phones and social profiles from company websites.
- YouTube Transcript Scraper: video transcripts for the same RAG pipelines.