Website Content Crawler API: Web Pages to Markdown for RAG avatar

Website Content Crawler API: Web Pages to Markdown for RAG

Pricing

from $1.06 / 1,000 pages

Go to Apify Store
Website Content Crawler API: Web Pages to Markdown for RAG

Website Content Crawler API: Web Pages to Markdown for RAG

Website content crawler API: turn any web page or whole site into clean Markdown and text for RAG, LLMs and AI agents. Crawls links and sitemaps, renders JavaScript pages when needed. RAG Web Browser alternative at $1 per 1,000 pages.

Pricing

from $1.06 / 1,000 pages

Rating

0.0

(0)

Developer

Saulius AutomatesIT

Saulius AutomatesIT

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Turn any web page, docs section or whole website into clean Markdown and text for RAG pipelines, LLM context, vector databases and AI agents. Give it URLs; it fetches them, strips navigation, footers, scripts and cookie banners, keeps the main content, and returns Markdown with headings, lists, links, code blocks and tables. Crawl links and sitemaps when you need a whole site. $1.25 per 1,000 pages, failed and empty pages free.

A cheaper alternative to RAG Web Browser ($2.28 per 1,000 fetched pages) and to compute-billed website crawlers: you pay a flat price per page with text, so a big crawl costs what you expect.

What you get per page

FieldExample
url, requestedUrl, statusCodefinal URL after redirects, the URL you gave, 200
title, descriptionpage title and meta description
markdownthe page as clean Markdown: # Title, ## Sections, lists, links, fenced code, tables
language, author, publishedAtfrom the page's meta tags and JSON-LD
canonicalUrl, imagecanonical link and share image
wordCountwords in the text
mainContentOnlytrue when only the main article or docs block was kept
depth, referrerUrl, startUrlwhere the crawler found the page
text, html, linksoptional: plain text, the cleaned HTML, every link on the page

Input

  • Start URLs: pages to read. A bare domain (apify.com) works.
  • Crawl depth: 0 (default) reads only your URLs. 1 also reads the pages they link to, and so on. Links are followed on the same site and inside the start URL's folder: https://docs.apify.com/platform/actors crawls https://docs.apify.com/platform/....
  • Use sitemaps: add every page in the site's sitemap.xml (and sitemaps listed in robots.txt) inside the start folder. The fastest way to get a whole docs site.
  • Max pages per start URL (100) and max pages in total.
  • Only URLs containing / Skip URLs containing: simple text filters such as /blog/ or ?page=.
  • Main content only (default on): Mozilla Readability keeps the article or docs body on content pages; home pages and listings keep everything except navigation, footers, scripts and cookie banners.
  • Remove elements: extra CSS selectors to drop, such as .sidebar, #comments.
  • Include plain text / cleaned HTML / links: extra fields, same price.
{
"startUrls": ["https://docs.apify.com/platform/"],
"maxCrawlDepth": 2,
"useSitemaps": true,
"maxPagesPerStartUrl": 500
}

Pricing

$1.25 per 1,000 pages with text ($0.00125 a page), plus Apify's $0.00005 Actor start. Apify Store discounts apply: Bronze $1.19, Silver $1.13, Gold and above $1.06 per 1,000. No charge for pages that fail, are not found, are blocked, are not HTML (PDF, images, files), have no text or only show text after JavaScript runs.

Examples: a 300 page docs site = $0.38. 100,000 blog posts for a vector database = $125. One page for an AI agent = $0.0013.

Speed: a run gets one CPU core per 4 GB of memory. The default 1 GB reads about 35 pages a minute; give big crawls 4 GB (about 4 times faster). The price per page is the same at any memory.

Use cases

  • RAG and vector databases: load documentation, help centres, blogs and knowledge bases as Markdown chunks with titles and source URLs.
  • AI agents and MCP: fetch a page as Markdown inside an agent loop; small, cheap, predictable.
  • LLM training and evaluation data: clean text of whole sites with language and dates.
  • Content monitoring: schedule a crawl of a competitor's blog or a docs section and diff the Markdown.
  • SEO audits: titles, descriptions, canonical links, word counts and internal links per page.

Error rows (free)

errorMeaning
NOT_FOUNDthe page answered 404 or 410
HTTP_403, HTTP_429 and so onthe site refused the page
BLOCKEDa bot check (Cloudflare challenge, captcha)
NOT_HTMLa PDF, image, archive or other file
NEEDS_JAVASCRIPTthe server sends an empty app shell; the text appears only after JavaScript runs
EMPTYthe page has no text
FAILEDthe site could not be reached
NO_DATAthe input had no usable URL

API, MCP and integrations

Call it from Python, JavaScript, Make, Zapier, n8n, LangChain or LlamaIndex like any Apify Actor, or give an AI agent this MCP server: https://mcp.apify.com/?tools=sauliusautomatesit/website-content-crawler-api.

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("sauliusautomatesit/website-content-crawler-api").call(
run_input={"startUrls": ["https://docs.apify.com/platform/"], "maxCrawlDepth": 1, "maxPagesPerStartUrl": 50}
)
for page in client.dataset(run["defaultDatasetId"]).iterate_items():
if page.get("type") != "error":
print(page["url"], page["wordCount"])

Notes

  • Pages are read as the server sends them (plain HTTP with a Chrome fingerprint through Apify datacenter proxies). Server rendered sites (docs, blogs, news, Wikipedia, most CMSs and Next.js sites) work; client only app shells come back as free NEEDS_JAVASCRIPT rows.
  • The crawler stays on the start URL's site and folder. Use several start URLs for several sections.
  • Respect each site's terms and robots rules for your use; this Actor reads only public pages.