Site to Markdown Crawler: Website Content Crawler Alternative
Pricing
from $0.50 / 1,000 page crawleds
Site to Markdown Crawler: Website Content Crawler Alternative
Crawls a website and converts each page to clean Markdown for AI and RAG use. Plain HTTP by default (256 MB), spend cap on, priced per page.
Pricing
from $0.50 / 1,000 page crawleds
Rating
0.0
(0)
Developer
Jack Valmadre
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
19 hours ago
Last modified
Categories
Share
Site to Markdown Crawler
Crawls a website and converts each page to clean Markdown, ready for AI ingestion, RAG pipelines or offline archiving. Uses plain HTTP (no browser), so it runs in 512 MB by default and keeps costs predictable.
Default protections: a page cap (500 pages) and a spend cap (US$5.00) are on by default. An unattended run stops at whichever limit it hits first. Adjust either value in the input, or set to 0 to disable.
What you get
One dataset row per crawled page:
| Field | Description |
|---|---|
url | URL as requested |
finalUrl | URL after redirects |
httpStatus | HTTP status code |
title | Page title, if found |
markdown | Page body as Markdown: headings, lists, emphasis, tables, code blocks |
wordCount | Word count of the extracted body |
status | ok, no-content, blocked, disallowed-by-robots, or error |
error | Error message, when status is error |
fetchedAt | ISO timestamp of the fetch |
no-content is set when the page is a navigation or index page with too little prose to be useful.
How it crawls
- Reads
robots.txtand honoursCrawl-delayand per-agent disallow rules before fetching any page. - Seeds the URL queue from the site's
sitemap.xml(found viarobots.txtor/sitemap.xml), then follows same-domain<a href>links in BFS order. - Extracts body text and converts to Markdown using trafilatura (Apache-2.0). Navigation, header and footer boilerplate are stripped at the library level.
- Stops at the page cap or spend cap, whichever comes first.
When to use this
- Feeding a documentation site, blog or wiki into a RAG pipeline or AI assistant.
- Archiving a site as structured Markdown.
- Any job where you want clean text at a known cost rather than raw HTML.
Pricing
US$0.0005 per page crawled (US$0.50 per 1,000 pages), plus a one-time US$0.00005 start charge per run.
Only pages with successfully extracted Markdown content are charged (status: ok). Pages that are blocked, disallowed by robots.txt, have no extractable text, or are skipped due to the spend cap are not charged.
The default spend cap is US$5.00, which covers up to 10,000 pages per run. Adjust maxTotalChargeUsd to match your job.
Why this is cheaper than Apify's Website Content Crawler
Apify's Website Content Crawler defaults to 8 GB of memory and bills by compute time. On a plain-HTTP crawl, that is roughly US$0.20 per 1,000 pages at its minimum; with a browser it is US$0.50–5 per 1,000 pages (Apify's own published range). An unmonitored run can accumulate large charges.
This actor runs at 512 MB, uses plain HTTP by default, and stops at the page and spend caps. The underlying platform compute cost on our runs was US$0.09–0.14 per 1,000 pages; the US$0.50 per 1,000 pages list price covers overhead and keeps your charges predictable.
| Site | Pages | Run time | Platform cost |
|---|---|---|---|
| flask.palletsprojects.com | 76 ok / 78 total | 252 s | US$0.0070 (US$0.090/1k) |
| www.sphinx-doc.org | 202 ok / 202 total | 754 s | US$0.0222 (US$0.110/1k) |
| docs.djangoproject.com | 204 ok / 204 total | 968 s | US$0.0275 (US$0.135/1k) |
These platform costs are from real Apify runs (actor HrXtDuKlhELYgA9UT, builds 0.1.3–0.1.7) at 512 MB. Your actual charge is the list price (US$0.0005/page), not the underlying compute cost. Your cost will vary with page count, page size and server speed.
Input
| Field | Default | Description |
|---|---|---|
startUrls | (required) | One or more URLs to start crawling from |
maxPages | 500 | Stop after this many pages |
maxTotalChargeUsd | 5.0 | Stop before the run charge exceeds this (pay-per-event billing only) |
sameDomainOnly | true | Only follow links on the same domain |
useSitemap | true | Seed the URL queue from the site's sitemap.xml |
timeoutSecs | 20 | Per-page request timeout in seconds |
maxConcurrency | 2 | Parallel requests (default 2 is safe at 512 MB; increase with memory: up to 4 at 1 GB) |
Use with AI agents
Pass startUrls as a list of {"url": "..."} objects, plain URL strings, or a single URL string. All discovered pages are returned in the default dataset as structured rows. The markdown field is ready to insert directly into a prompt or embed in a vector store. Filter by status: "ok" to drop navigation pages and errors before chunking.
Robots.txt and terms of service
This actor reads every site's robots.txt before crawling and honours Disallow rules for User-agent: * and User-agent: MadrascoSiteMarkdownCrawler (RFC 9309). Sites that block our user agent or return a captcha or 403 are recorded as blocked; they are not retried or bypassed.
You are responsible for ensuring your use of this actor complies with the target site's terms of service.
Limitations
- JavaScript-rendered content: the actor uses plain HTTP, not a browser. Pages that require JavaScript to render their main content will return incomplete or empty Markdown. For those sites, expect
no-contentresults. - Slow servers: the default 20-second per-page timeout may cause errors on slow hosts. Raise
timeoutSecsif you see a high error rate. - Large sites: the default page cap is 500. Set
maxPagesandmaxTotalChargeUsdto match the job; a test run on a sample of URLs is a good first step for unfamiliar sites. - Sparse pages: index and search pages with little prose are returned with
status: no-contentrather than discarded, so you can see what the crawler found.
Publisher
Built by Madrasco. Support: use the Issues tab on this actor's page.