Website to Markdown — Content Crawler for LLM & RAG
Pricing
from $0.60 / 1,000 results
Website to Markdown — Content Crawler for LLM & RAG
Content Crawler turns any site into clean Markdown per page for LLMs, RAG pipelines and vector DBs — no headless browser, $1 per 1,000 pages.
Pricing
from $0.60 / 1,000 results
Rating
0.0
(0)
Developer
Murat Uzun
Maintained by CommunityActor stats
0
Bookmarked
5
Total users
2
Monthly active users
2 days ago
Last modified
Categories
Share
What is Website to Markdown Crawler?
Website to Markdown Crawler is an Apify Actor that turns any website into clean, LLM-ready Markdown — one row per page, with navigation, footers, sidebars and cookie banners stripped, headings and code blocks preserved, and every link made absolute. It runs on plain HTTP with Crawlee's CheerioCrawler — no headless browser — which is what makes it fast and roughly 10× cheaper than browser-based content crawlers. Feed it a start URL (or a whole /sitemap.xml) and it does a breadth-first crawl up to your page budget, ready to feed an LLM, a RAG pipeline, LangChain/LlamaIndex loader, or a vector DB (Pinecone, Qdrant, Weaviate, pgvector) in minutes.
Pricing: Website to Markdown costs $1.00 per 1,000 results (pay per result, no subscription; Apify's free plan credit covers small runs).
Why use Website to Markdown Crawler?
- Feed an LLM or RAG pipeline in minutes. Point it at your docs site, a competitor's help center, or your own product site and get clean Markdown ready for chunking and embedding — no HTML soup to clean up first.
- 10× cheaper, no browser.
apify/website-content-crawlerand similar Actors spin up a browser per page. This Actor fetches raw HTML over HTTP, so it's faster and priced at $1 per 1,000 pages. - Predictable cost.
maxPagesis the billed unit and the crawl stops at exactly that count — no surprise runs. - Sitemap-aware BFS crawl. Optionally seeds from
/sitemap.xml(including nested sitemap indexes) so you reach real content pages immediately instead of waiting for link discovery.
How to use Website to Markdown Crawler
- Paste one or more pages into Start URLs, e.g.
https://docs.apify.com/platformor your own docs/blog root. - Set Max pages — this is the billed unit ($1 per 1,000 pages) and the crawl stops exactly there.
- Optionally set Include path prefixes (e.g.
/docs) to stay inside one section of a larger site. - Choose Output format: Markdown, plain text, or both.
- Click Start, then download the dataset as JSON, CSV, Excel or HTML, or pull it straight into your RAG pipeline via the API.
Example input
{"startUrls": [{ "url": "https://docs.apify.com/platform" }],"maxPages": 50,"maxDepth": 3,"sameDomainOnly": true,"useSitemap": true,"outputFormat": "markdown"}
Example output
{"url": "https://crawlee.dev/docs","finalUrl": "https://crawlee.dev/js/docs/quick-start","statusCode": 200,"title": "Quick Start | Crawlee for JavaScript","description": "Crawlee helps you build reliable scrapers. Fast.","lang": "en","canonical": "https://crawlee.dev/js/docs/quick-start","markdown": "# Quick Start\n\nWith this short tutorial you can start scraping with Crawlee in a minute or two...\n\n## Choose your crawler\n\n### CheerioCrawler\n\nThis is a plain HTTP crawler...","text": null,"wordCount": 1484,"headings": ["Quick Start", "Choose your crawler", "CheerioCrawler", "PuppeteerCrawler"],"links": ["https://crawlee.dev/js/docs/introduction", "https://crawlee.dev/js/docs/guides"],"depth": 0,"contentHash": "3f9a1c2e...","crawledAt": "2026-09-12T18:21:07.000Z","error": null}
What data does Website to Markdown Crawler extract?
One row per crawled page:
| Field | Type | Description |
|---|---|---|
url / finalUrl | string | Requested URL and URL after redirects |
statusCode | integer | HTTP status code, null on a failed request |
title / description | string | Page <title> (or first <h1>) and meta description |
lang | string | <html lang> attribute |
canonical | string | <link rel=canonical>, resolved to an absolute URL |
markdown | string | Main content as Markdown — headings, lists, code blocks, tables, absolute links/images |
text | string | Main content as plain text |
wordCount | integer | Word count of the extracted content |
headings | array | H1-H3 text, in order, up to 50 |
links | array | Same-domain absolute links found on the page, up to 200 |
depth | integer | Link depth from the nearest start URL |
contentHash | string | SHA-1 of the plain text — diff two runs to detect changed pages |
crawledAt | string | ISO timestamp |
error | string | Always null: pages that fail are not charged and are listed in the run's ERRORS record instead |
chunks | array | With Add RAG chunks: { text, tokens, headingPath } per chunk, otherwise null |
changeStatus | string | With Detect changes: new, changed or unchanged, otherwise null |
markdown/text are populated according to the Output format input (markdown, text, or both); the unused one is null.
Input parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
startUrls | array | docs.apify.com/platform | Page(s) to start crawling from |
maxPages | integer | 50 | Billed unit; the crawl stops at exactly this count (1-5000) |
maxDepth | integer | 3 | Link depth from a start URL (0-10) |
sameDomainOnly | boolean | true | Only follow links on the same domain (www. counts as the same) |
includePathPrefixes | array | none | Only crawl paths starting with one of these, e.g. /docs |
excludePathPatterns | array | images/PDF/CSS/JS, /login, /signup, ?replytocom= | Regexes tested against the full URL; a match is skipped |
useSitemap | boolean | true | Also seed from /sitemap.xml (follows one nested sitemap-index level) |
outputFormat | string | markdown | markdown, text, or both |
removeSelectors | array | none | Extra CSS selectors to strip before extraction |
maxConcurrency | integer | 10 | Parallel requests (1-50) |
respectRobots | boolean | true | Skip URLs disallowed by robots.txt |
proxyConfiguration | object | none | Optional Apify Proxy configuration |
chunkForRag | boolean | false | Add a chunks array (heading-aware, exact o200k token counts) to every page |
chunkSize | integer | 500 | Max tokens per chunk (50-8000) |
changeDetection | boolean | false | Add changeStatus (new / changed / unchanged) versus the previous run with the same start URLs |
skipUnchangedContent | boolean | true | With change detection on, unchanged pages get empty markdown/text/chunks |
Pricing
Website to Markdown Crawler uses pay-per-event pricing: $1 per 1,000 pages ($0.001 per page), plus a negligible actor-start fee. There is no separate compute-unit or browser billing because no browser is used. Set Max pages (and, on the platform, Maximum cost per run) to cap spend — the Actor trims its own crawl to stay within budget and never bills more pages than it delivers.
Website to Markdown Crawler vs. browser-based content crawlers
Browser-based crawlers such as apify/website-content-crawler load every page in a real (often headless) browser before extracting content — that's necessary for JavaScript-rendered single-page apps, but it's slow and expensive for the huge majority of documentation sites, blogs and marketing pages that are plain server-rendered HTML. This Actor skips the browser entirely: a page that takes a browser-based crawler several seconds and a full render cycle is a sub-second HTTP fetch here, at a fraction of the price. The trade-off is explicit: sites that only render content client-side with JavaScript (pure SPAs) will come back with little or no text — see Limitations below.
RAG-ready chunks in the same run
Turn on Add RAG chunks and every page row gets a chunks array: heading-aware pieces of the Markdown, each with its heading path and an exact GPT (o200k) token count — embed them straight into Pinecone, Qdrant, pgvector or any vector DB. Code blocks stay whole and every chunk is at most Chunk size tokens. No extra charge: chunks are part of the page result.
Keep a knowledge base fresh: change detection
Turn on Detect changes since last run and schedule the Actor. Each page gets changeStatus — new, changed or unchanged — compared with the previous run that used the same start URLs (the state is kept in a key-value store in your own account). With Leave content empty for unchanged pages on, unchanged rows carry no Markdown, so you only re-embed what changed. Every crawled page still counts as one result.
Using Website to Markdown Crawler with AI agents, LangChain and LlamaIndex
Website to Markdown Crawler runs on pay-per-event pricing with limited permissions, so it's callable through the Apify MCP server directly from an AI agent — pass startUrls and maxPages and get back one Markdown row per page. The dataset also drops straight into a RAG pipeline: use Apify's LangChain or LlamaIndex ApifyDatasetLoader to turn the dataset into Document objects, then chunk and embed into Pinecone, Qdrant, Weaviate or pgvector. It also connects through n8n, Make and Zapier via Apify's standard integrations.
Limitations
- JavaScript-rendered single-page apps are the main limitation. This Actor fetches raw HTML over HTTP and never executes page JavaScript, so a site that renders its content entirely client-side (no content in the initial HTML) will come back with an empty or near-empty row. It does follow
<meta http-equiv="refresh">redirect stubs (one hop) and ordinary HTTP redirects, which covers a common class of "empty landing page" issue, but not full client-side rendering. If you hit this, look for a browser-based crawler instead. - Interactive widgets that swap content via JavaScript (tabs, accordions with lazy content) are captured in whatever state they render in the raw HTML — usually just the first/default tab.
- Short run timeouts end cleanly. If the run timeout is reached before
maxPages, the crawler stops shortly before the deadline and the run finishes as SUCCEEDED with every page crawled so far (you pay only for those). Raise the timeout or lowermaxPagesfor a full crawl. - The main-content heuristic (no
@mozilla/readability) is tuned for documentation and blog layouts (article,main,#content,.content); unusual layouts may needremoveSelectorsto clean up.
FAQ
Is this legal? The Actor reads only public HTML. You're responsible for crawling sites you're allowed to and respecting their terms of service.
Why is a row empty? Almost always a JavaScript-rendered page — see Limitations above. Check the statusCode field (failed pages are in the ERRORS record); a 200 with an empty markdown usually means client-side rendering.
How do I stay inside one section of a site? Set includePathPrefixes, e.g. ["/docs"], and/or turn off useSitemap if the sitemap covers more than you want.
How do I control cost? Set maxPages; the Actor never crawls or bills past it. On the platform, also set Maximum cost per run.
Can I detect changed pages between runs? Yes — compare contentHash (SHA-1 of the extracted text) across runs to see which pages changed.
Does this work with n8n, Make or Zapier? Yes, through Apify's standard integrations, and through LangChain/LlamaIndex document loaders for RAG pipelines.
Related Actors
Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:
Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.
Website & domain intelligence
- Email Extractor — Website Contact & Social Finder — e-mails, phones and social profiles per domain
- Tech Stack Detector — Wappalyzer & BuiltWith Alternative — CMS, e-commerce, analytics, pixels and payments per domain
- Domain DNS & Email Security Checker — SPF, DKIM, DMARC, MX provider, registrar and domain age
- Domain Security Audit (TLS, HTTP headers, redirects, robots) — TLS expiry, security headers, redirect chain, robots and llms.txt
- Subdomain Finder (Certificate Transparency) — every subdomain seen in CT logs, with a live DNS check
- Bulk Core Web Vitals & PageSpeed Audit — Lighthouse scores, LCP, CLS, INP and top fixes per URL
- On-Page SEO Audit — title, meta, headings, links, images and schema issues per page
- Sitemap URL Extractor & Change Monitor — every sitemap URL, or new and removed pages between runs
- Wayback Machine Snapshot & Page Change Tracker — how a page changed over time, or every archived snapshot
- Bulk Domain WHOIS & RDAP Lookup — registrar, dates, status and nameservers per domain
- Web Scraper — CSS Selector & Data Extractor — pull any CSS selector off any page, one row per URL
- Website Screenshot Generator — full-page or viewport PNG/JPEG screenshots of any URL
Content for AI, LLMs and RAG
- AI Web Search & Read: Google results as clean Markdown — a query turned into clean Markdown from the top search results
- llms.txt Generator (Website to llms.txt & llms-full.txt) — llms.txt and llms-full.txt for any website or docs site
- PDF to Markdown Converter (Text, Headings, Metadata) — PDF files to clean Markdown with headings, metadata and token counts
- RAG Text Chunker (Markdown to Chunks, Token Counts) — heading-aware RAG chunks with exact GPT token counts from any dataset
- Article & News Extractor (clean text, author, date, markdown) — clean article text, author, date and Markdown per URL
- Structured Data & JSON-LD Extractor (Schema.org, Open Graph) — Schema.org and Open Graph data from any page
- Google News Scraper (RSS search by keyword, topic, site) — news results by keyword, topic or site
- Press Release Monitor: PR Newswire, BusinessWire, GlobeNewswire — PR Newswire, Business Wire and GlobeNewswire releases
Search, video and social
- YouTube Shorts Scraper — Shorts from channels, hashtags and searches with view counts
- Pinterest Pins Scraper — latest pins of public Pinterest profiles and boards
- YouTube Transcript Scraper — captions and subtitles as text + timed segments, per video or channel
- Google Search Results Scraper — SERP API — organic SERP results per keyword and country
- YouTube Comments Scraper — Comments & Replies — comments and replies with likes, no API key
- YouTube Channel Latest Videos (RSS, no API key) — the latest 15 videos of any channel from RSS
- YouTube Channel Scraper (videos, shorts, live) — a channel's full video, shorts and stream list
- YouTube Search Results Scraper (videos, channels, no API key) — videos, channels and playlists per query
- YouTube Video Details Scraper (views, likes, description, tags) — views, likes, description, tags and chapters per video
- Apple Podcasts Lookup & Episodes Scraper — podcast metadata and episodes from iTunes and RSS
- Bluesky Post, Search & Profile Scraper — posts, profiles, followers and threads from the AT Protocol API
- X Tweet Scraper (Twitter Posts by URL, No Login) — X (Twitter) posts by URL with likes, replies, author, media and MP4 links
- TikTok Profile Scraper (Followers, Likes, Bio, Video Stats) — TikTok profiles with exact follower and like counts, plus video stats
- Telegram Channel Posts Scraper — posts, views and media flags from any public channel
- Substack Publication & Posts Scraper — archive, authors and paywall status per publication
- Google Play Reviews Scraper — reviews, ratings, replies and app versions per app
- App Store Reviews Scraper — iOS reviews and ratings per app and country
- Trustpilot Reviews Scraper (Ratings, Replies, No Login) — Trustpilot reviews with stars, full text, verification and company replies
- Google Trends Scraper — interest over time, by region, and related queries per keyword
- Google Ads Transparency Scraper — ads any advertiser runs on Google, with format and dates
- Keyword Suggestions Scraper (Google, YouTube, Amazon, Bing) — autocomplete keyword ideas from four search engines
- Bilibili Scraper (Videos, Search, Popular) — Chinese video platform: views, likes, coins, danmaku, uploader
- Mastodon Scraper (Hashtags, Accounts, Trending) — public fediverse posts by hashtag, account or trending
- Meetup Events Scraper (Search by Keyword & City) — upcoming events with RSVPs, fees, venues and groups
- Eventbrite Scraper (Events by Keyword & City) — events by keyword and city with venue, dates and organizer
Leads, jobs and company data
- Career Site Jobs API (Greenhouse, Lever, Ashby, Workday +1) — company domains in, their open jobs out, ATS detected automatically
- Workday Jobs Scraper — jobs with full descriptions from any Workday career site
- Google Maps Scraper — businesses with phone, website, address, rating and coordinates per search
- LinkedIn Jobs Scraper — job titles, companies, locations and full descriptions from LinkedIn job search
- Company 360: full company profile from a domain — one row per domain: contacts, tech, security, hiring and company facts
- Hiring Signals Scraper (Greenhouse, Lever, Ashby, Workable) — open jobs and hiring velocity from 10 public ATS boards
- Y Combinator Companies & Founders Scraper — YC startups by batch, industry and hiring status
- Wikidata Entity & Company Enrichment (facts, IDs, links) — HQ, founders, employees, revenue and social IDs per company
- Email Validator & Verifier — Bulk Email Check — syntax, MX, disposable, role and free-provider checks
- OpenStreetMap POI Extractor (Overpass API: shops, amenities) — shops and amenities by radius, bbox or area
- Stock, Crypto & FX Quotes — one row per symbol from Yahoo, Binance and ECB rates
- Remote Jobs Aggregator (RemoteOK, WWR, Hacker News) — one clean row per remote job, de-duplicated across feeds
- Greenhouse Jobs Scraper — jobs with descriptions from any Greenhouse job board
- Lever Jobs Scraper — jobs with descriptions from any Lever careers page
- Ashby Jobs Scraper — jobs, salaries and descriptions from any Ashby job board
- SmartRecruiters Jobs Scraper — jobs with descriptions from any SmartRecruiters company
- Seek Jobs Scraper (Australia & New Zealand) — Seek job ads with salary, work type and location
- Dice Jobs Scraper — US tech jobs from Dice with salary and remote flag
- AutoScout24 Scraper — European car listings with price, mileage and seller
- Rightmove Scraper — UK property for sale or rent with price and agent
- Wellfound Jobs Scraper (AngelList Startup Jobs) — startup jobs with salary and equity ranges, company size and stage
- Yandex Maps Scraper (Places, Ratings, Phones) — businesses in Russia, Türkiye and the CIS with phones, ratings, hours
- Craigslist Scraper (Listings, Prices, Locations) — listings in any area and category with price, date and coordinates
- JobStreet Scraper (Malaysia, Singapore, PH, ID + JobsDB) — JobStreet and JobsDB jobs in 6 Asian countries with parsed salaries
- InfoJobs Scraper (Spain Jobs, Salaries, Companies) — Spanish jobs with salary range, contract type and full description
- StepStone Scraper (Germany, Austria, Belgium Jobs) — StepStone jobs in Germany, Austria and Belgium by keyword and city
- Zillow Scraper (Homes for Sale, Rent & Sold, Zestimates) — Zillow homes for sale, rent or sold with Zestimates, beds, baths and photos
- Zillow Home Details Scraper (Description, Photos, History) — full Zillow property pages: description, photos, HOA, tax rate, agent
- Redfin Scraper (Homes for Sale, Prices, Details) — US homes for sale or sold from any Redfin search, with price and details
- Kleinanzeigen Scraper (Ads, Prices, Locations) — German classifieds with price, VB flag, ZIP, city and seller type
Developer, app and research data
- npm, PyPI & Crates.io Package Health Checker — releases, downloads, deprecation and a health score
- GitHub Repository Health & Activity Report — stars, commits, contributors and risk flags per repo
- VS Code Marketplace Extension Scraper (installs, ratings) — installs, ratings and versions per extension
- Chrome Web Store Extension Scraper (installs, ratings) — users, rating, version and developer per extension
- Google Play Scraper — apps, ratings, installs, developer contact and reviews
- App Store (iOS) App Metadata, Ratings & Top Charts Lookup — ratings, price, version and charts per app
- CrossRef DOI & Citation Metadata Lookup — papers, authors, journals and citation counts
- FDA Recalls & Adverse Events Monitor (openFDA) — food, drug and device recalls from openFDA
- iCal / ICS Calendar Feed to Events Extractor — any public calendar feed as event rows
- Shopify Store Products Scraper — catalog, prices, variants and stock per store
- Hacker News Search & Front Page Scraper — stories, comments and points by query or front page
- GitHub Trending Repositories Scraper — trending repos and developers by language and period
- Stack Overflow & Stack Exchange Q&A Scraper — questions, answers and scores by query, tag or site
- Bulk Image Downloader — download image URLs to storage with size, dimensions and a ZIP
- Google Flights Scraper (Prices, Airlines, Stops) — flight prices, airlines, times, stops and CO2 by route and date
- Google Hotels Scraper (Prices, Ratings, Reviews) — hotel prices per night, stars, rating and reviews by city and dates
- Booking.com Scraper (Hotel Prices, Ratings, Availability) — Booking.com hotels for any city and dates with prices, scores and deals
- Airbnb Scraper (Listings, Prices, Ratings, Coordinates) — Airbnb stays for any city and dates with prices, ratings and coordinates
- AliExpress Scraper (Search Products & Prices) — AliExpress search results with USD price, discount and rank
- Lazada Scraper (Products, Prices, Sold, Ratings) — Lazada products in 6 countries with price, rating, units sold and seller
- Trendyol Scraper (Turkey Products, Prices, Ratings) — Trendyol products with price, basket discount, rating and promotions
Support and feedback
Found a bug or want a feature? Open an issue on the Issues tab.