Website to Markdown Crawler | Pre-Chunked for RAG avatar

Website to Markdown Crawler | Pre-Chunked for RAG

Pricing

from $1.00 / 1,000 page crawleds

Go to Apify Store
Website to Markdown Crawler | Pre-Chunked for RAG

Website to Markdown Crawler | Pre-Chunked for RAG

Crawl any website into clean, pre-chunked Markdown with per-chunk token counts for RAG pipelines, vector DBs (Pinecone, Qdrant) and LLM context. MCP native for Claude & ChatGPT. SPA support via Playwright. Pay only for pages that crawl. A Firecrawl alternative.

Pricing

from $1.00 / 1,000 page crawleds

Rating

0.0

(0)

Developer

The Mine Works

The Mine Works

Maintained by Community

Actor stats

0

Bookmarked

7

Total users

3

Monthly active users

10 minutes ago

Last modified

Share

29 docs pages to 303 RAG chunks in 26 s

From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors.

Point it at a website and get every page back as clean Markdown, already cut into chunks along its headings, each chunk with its heading path and a token estimate. That is the shape a RAG pipeline, a vector database (Pinecone, Qdrant, Weaviate) or an LLM context window wants, without writing HTML cleanup or chunking code. Plain HTTP by default, a real browser when a site needs JavaScript. No login, no API key.

Why choose this actor?

  • A docs site turned into chunks in under half a minute. In a recorded run on 2 October 2026, Apify's documentation (from https://docs.apify.com/platform, 30 pages requested) came back as 29 pages and 303 chunks in 26 seconds, about 30,000 tokens in all, with no chunk over the 500-token limit we set.
  • Chunks that keep their place in the document. Every chunk carries heading_path, the headings above it from the top level down (two levels deep on Apify's schedules page, for example), and a section split into several chunks repeats its heading at the top of each, so retrieval keeps its context.
  • Pay only for pages that worked. A page that times out, fails to render, is blocked or has no readable text comes back as a free failed row with the reason. The Gold price is $1.00 per 1,000 pages.

Run it on Apify

Part of The Mine Works Developer and AI tools family: GitHub Skill Finder, GitHub Repo Scraper, GitHub Trending Scraper.

Try it in one minute

Paste this into the JSON tab of the input page and press Start. It crawls 10 pages of a documentation site in well under a minute.

{
"startUrls": [{ "url": "https://docs.apify.com/platform" }],
"maxPages": 10,
"outputFormat": "both"
}

The only input you must give is startUrls: one or more pages to start from, each as { "url": "https://..." } (in the Console you can paste URLs one per line). The crawler follows links from those pages within the same domain until it reaches maxPages. Use excludeUrlPatterns to keep it out of sections you do not want, and outputFormat to choose chunks, full-page Markdown or both.

Apify's free plan includes $5 of credit every month, which covers about 2,400 pages at this actor's Free plan price ($0.002 a page plus the $0.005 start fee, in runs of 100 pages).

Copy to your AI assistant

Paste this block into ChatGPT, Claude, Cursor or any assistant that can write code, and it can run the actor for you.

themineworks/rag-crawler on Apify. Crawls a website from start URLs within the same domain and returns one row per page with url, canonical, title, description, language, word_count, token_count, crawled_at, status, plus chunks (array of {heading_path, text, token_count, chunk_index}) and/or markdown. Call ApifyClient("TOKEN").actor("themineworks/rag-crawler").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Required: startUrls ([{"url": "https://..."}]). Optional: maxPages (1 to 500, default 10), renderJs (default false; true uses a headless browser), waitStrategy ("networkidle" default, "domcontentloaded" or "load", browser mode only), outputFormat ("chunks" default, "full" or "both"), maxTokensPerChunk (50 to 4000, default 800), stripBoilerplate (default true), excludeUrlPatterns (regex strings), customCss (CSS selector for the main content). Rows with status "failed" are not billed; rows with _type "summary" or "info" are run reports. Full spec: GET https://api.apify.com/v2/acts/themineworks~rag-crawler/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations

Key features

  • Up to 500 pages per run, following links within the start URL's domain. In our test, 30 page requests produced 29 pages in 26 seconds on plain HTTP.
  • Heading-aware chunking. Markdown is split at every heading level (# to ######); a section longer than maxTokensPerChunk is packed paragraph by paragraph into several chunks, each starting with the section's heading line.
  • Three output shapes. chunks (default) for RAG, full for one Markdown document per page, or both.
  • Boilerplate removal. Mozilla Readability picks the main article, and navigation, headers, footers, sidebars, forms, scripts and common ad or cookie elements are dropped. A customCss selector, when it matches, overrides all of that and keeps exactly the element you name.
  • JavaScript sites. renderJs: true switches to a headless browser (Playwright) with a choice of when to read the page (waitStrategy).
  • Token estimates for budgeting. Every page and every chunk carries token_count, an estimate of words times 1.35, so you can size embedding costs and context windows before indexing.

How to use it

Basic: one site, default settings

{
"startUrls": [{ "url": "https://docs.apify.com/platform" }],
"maxPages": 50
}

You get up to 50 pages as heading-split chunks of at most about 800 tokens each.

Tight chunks for embeddings, plus the full page

{
"startUrls": [{ "url": "https://docs.apify.com/platform" }],
"maxPages": 30,
"outputFormat": "both",
"maxTokensPerChunk": 500
}

This is the exact input of our 2 October test: 29 pages, 303 chunks, no chunk over 500 estimated tokens, in 26 seconds. Index chunks; keep markdown for display or re-chunking later.

Keep the crawl inside one section

{
"startUrls": [{ "url": "https://docs.example.com/guides/" }],
"maxPages": 200,
"excludeUrlPatterns": ["/blog/", "/changelog/", "/api-reference/", "\\.(pdf|zip|png|jpg|jpeg|gif|svg)$"]
}

The crawler follows any link on the same domain, including other subdomains of it (in our test it reached console.apify.com from docs.apify.com). List the sections you do not want as regular expressions; matching links are never queued and never charged.

A JavaScript-only site

{
"startUrls": [{ "url": "https://app.example.com/help" }],
"maxPages": 25,
"renderJs": true,
"waitStrategy": "networkidle"
}

Use this only when the plain mode returns no_content failures or empty pages, since a browser is slower. domcontentloaded is quicker than networkidle when the content is in the first HTML.

Nightly refresh of a help center for a support bot

{
"startUrls": [{ "url": "https://help.example.com/" }],
"maxPages": 300,
"outputFormat": "chunks",
"customCss": "article"
}

Save this as a task and schedule it nightly (for example 0 2 * * *), then upsert chunks into your vector store keyed on url plus chunk_index. customCss keeps only the matched element when it exists on a page, which is the cleanest option for a site with one consistent article template.

Input parameters

ParameterTypeDefaultWhat it does
startUrlsarrayrequired (form prefill: https://example.com)Pages to start from, as { "url": "..." } objects.
maxPagesinteger (1 to 500)10Most pages to request in the run, including failed ones.
renderJsbooleanfalsefalse fetches HTML directly (fast); true renders each page in a headless browser.
waitStrategystringnetworkidleBrowser mode only: networkidle, domcontentloaded or load. If the chosen event does not come within 30 seconds, the actor falls back to domcontentloaded.
outputFormatstringchunkschunks adds the chunks array, full adds markdown, both adds both.
maxTokensPerChunkinteger (50 to 4,000)800Target size of a chunk, in estimated tokens. A single paragraph longer than this is never split, so the odd chunk can be larger.
stripBoilerplatebooleantrueUse Mozilla Readability to keep the main content. Off, the whole page body is converted (navigation and footers are still dropped).
excludeUrlPatternsarray of strings[] (form prefill: file extensions, /tag/, /author/)Regular expressions, case-insensitive. Matching links are skipped.
customCssstringempty (form prefill: article.main-content)A CSS selector for the main content. When it matches, only that element is converted; when it does not, the actor falls back to Readability.

"Form prefill" values fill the Console form for you; an API call that leaves a field out gets the default shown. The Console prefill for customCss is only an example; it is ignored on pages where nothing matches.

There is no proxy setting: pages are fetched directly from Apify's servers, which suits public documentation and content sites. The default timeout is 300 seconds; for runs of a few hundred pages, or any run with renderJs, give the run more time in Run options.

What data do you get?

One row per page.

The page: url, canonical (from the page's canonical link, or the URL itself when there is none), title, description (the meta description), language (the lang attribute of the page's <html> tag), crawled_at, status (success).

Size: word_count and token_count for the whole cleaned page. token_count is an estimate (words times 1.35, rounded up), not the count of any one model's tokenizer; expect it to be within about 10% for English prose and further off for code.

Content: chunks (with outputFormat chunks or both), an array of { heading_path, text, token_count, chunk_index }, and markdown (with full or both), the whole cleaned page. heading_path is the list of headings above the chunk, empty for text before the first heading; chunk_index counts from 0 within the page. Headings are copied as the page writes them, so documentation sites that put anchor links in their headings keep those (for example Apify Console[](#apify-console)).

description and language are read from the raw HTML and need quoted attributes: on Wikipedia language came back as en, but on the minified Apify docs both were empty strings on all 29 pages.

Failed pages come back as { url, status: "failed", reason, message, charged: false }, where reason is no_content, timeout, render_error or blocked. They are never billed.

Each run ends with a _type: "summary" row: pages_requested, pages_crawled, pages_failed, total_tokens, charged_for and the run's own cost-guard figures; a _type: "info" row follows when pages were delivered. Neither is billed.

Stable fields for automations

These fields were present in every successful page row we sampled (32 pages across two runs on 2 October 2026):

FieldWhat it holds
urlThe page address that was crawled
canonicalCanonical URL, or url when the page declares none
titlePage title
descriptionMeta description, or an empty string
language<html lang> value, or an empty string
word_countWords in the cleaned Markdown
token_countEstimated tokens in the cleaned Markdown
crawled_atISO timestamp of the crawl
statussuccess on page rows (failed rows are described above)
chunksPresent when outputFormat is chunks or both
markdownPresent when outputFormat is full or both

Inside chunks, every item has heading_path, text, token_count and chunk_index. We will not rename these fields. New fields may be added over time; existing ones keep their names.

Output examples

Real rows from our 2 October 2026 runs, with text trimmed.

A documentation page with nested headings (outputFormat: "both", maxTokensPerChunk: 500; 8 chunks, the first three shown):

{
"url": "https://docs.apify.com/actors/running/schedules",
"canonical": "https://docs.apify.com/actors/running/schedules",
"title": "Actor and task schedules | Platform | Apify Documentation",
"description": "",
"language": "",
"word_count": 1278,
"token_count": 1726,
"crawled_at": "2026-10-02T08:25:21.241Z",
"status": "success",
"chunks": [
{
"heading_path": [],
"text": "Schedules allow you to run your Actors and tasks at specific times. You schedule the run frequency using [cron expressions](#cron-expressions)…",
"token_count": 185,
"chunk_index": 0
},
{
"heading_path": ["Set up a new schedule[](#set-up-a-new-schedule)"],
"text": "## Set up a new schedule[](#set-up-a-new-schedule)\n\nBefore setting up a new schedule, you should have the [Actor](https://docs.apify.com/actors) or [task](https…",
"token_count": 126,
"chunk_index": 1
},
{
"heading_path": ["Set up a new schedule[](#set-up-a-new-schedule)", "Apify Console[](#apify-console)"],
"text": "### Apify Console[](#apify-console)\n\nIn [Apify Console](https://console.apify.com/schedules), click on the **Schedules** in the navigation menu…",
"token_count": 424,
"chunk_index": 2
}
],
"markdown": "Schedules allow you to run your Actors and tasks at specific times. You schedule the run frequency using [cron expressions](#cron-expressions).\n\nTimezone & Daylight Savings Time…"
}

A Wikipedia article (outputFormat: "chunks", default 800-token limit; 14 chunks, two shown):

{
"url": "https://en.wikipedia.org/wiki/Web_scraping",
"canonical": "https://en.wikipedia.org/wiki/Web_scraping",
"title": "Web scraping - Wikipedia",
"description": "",
"language": "en",
"word_count": 4523,
"token_count": 6107,
"crawled_at": "2026-10-02T08:26:10.977Z",
"status": "success",
"chunks": [
{
"heading_path": ["Human copy-and-paste"],
"text": "### Human copy-and-paste\n\n\\[[edit](https://en.wikipedia.org/w/index.php?title=Web_scraping&action=edit&section=3 …",
"token_count": 134,
"chunk_index": 2
},
{
"heading_path": ["Methods to prevent web scraping"],
"text": "## Methods to prevent web scraping\n\n- [Academic journal publishing reform](https://en.wikipedia.org/wiki/Academic_journa…",
"token_count": 168,
"chunk_index": 11
}
]
}

The title here is Wikipedia's own page title, which contains its own hyphen. Note also that Wikipedia's "[edit]" links survive the cleanup; strip them in your pipeline if they matter to you.

The smallest possible run, from our daily platform check on 1 October (example.com, one page, 22 seconds), with its summary row:

{
"url": "https://example.com",
"canonical": "https://example.com",
"title": "Example Domain",
"description": "",
"language": "",
"word_count": 27,
"token_count": 37,
"crawled_at": "2026-10-01T09:00:49.448Z",
"status": "success",
"chunks": [
{
"heading_path": [],
"text": "This domain is for use in documentation examples without needing permission…",
"token_count": 37,
"chunk_index": 0
}
]
}
{
"_type": "summary",
"pages_requested": 1,
"pages_crawled": 1,
"pages_failed": 0,
"total_tokens": 37,
"charged_for": 1
}

Pricing

Pay per event: you pay for each page crawled and delivered to your dataset, plus a small start fee per run.

EventFreeBronzeSilverGold and above
page-crawled, per page$0.002$0.0016$0.00125$0.001
page-crawled, per 1,000 pages$2.00$1.60$1.25$1.00
apify-actor-start, per run$0.005 per GB of run memory, minimum one eventsamesamesame

The start fee, exactly. Apify's apify-actor-start event is charged once when a run starts, at $0.005 for each GB of memory the run uses, with a minimum of one event. This actor runs on 1 GB by default, so a default run pays $0.005. If you raise memory for browser mode (for example to 4 GB), the start fee rises to $0.02.

Worked examples on the Gold tier: 30 pages cost $0.03 plus $0.005. 500 pages, the most in one run, cost $0.50 plus $0.005. On the Free tier, 500 pages cost $1.00 plus $0.005. The price is per page, not per chunk: the 303 chunks of our test cost the same as 29 pages.

Never charged: pages that time out, fail to render, are blocked or have no readable text (the failed rows), links skipped by excludeUrlPatterns, retries, and the summary and info rows. A run where every page fails pays only the start fee, and the crawl stops after 5 failed pages in a row.

There is no scheduled price change for this actor; these prices have applied since 14 September 2026. The Pricing tab on this page always shows the rate for your own plan; if it and this table ever differ, the Pricing tab is right.

Run it on a schedule

Turn on monitorMode and each run delivers only the pages that are new or have changed since an earlier run with the same input, so a daily run costs you only for what moved.

  1. Enter your input, switch on Monitor mode and click Save as a task.
  2. In Apify Console open Schedules, click Add schedule and pick Daily (or any time and timezone you like).
  3. Under Actors or tasks to run, add the task you saved and save the schedule.
{
"startUrls": [
{
"url": "https://docs.apify.com/platform"
}
],
"maxPages": 20,
"outputFormat": "chunks",
"monitorMode": true
}

The first run delivers every page it finds. After that, a page is delivered again only when one of its fields changed, and unchanged pages are skipped and never charged. The summary row at the end shows new_this_run, changed_this_run and skipped_duplicates. Changing the input starts a fresh history; changing only the result limit does not. A changed page comes back whole, with all of its chunks, so you can replace that page's chunks in your vector store. Unchanged pages are still crawled, so changes deeper in the site are found, but they are not delivered and cost you nothing. A site where little changed can stop before maxPages to keep request costs in line; the summary row then shows profit_guard_tripped.

FAQ

What does it crawl? Any public website. It starts at your startUrls, follows links on the same domain (including its other subdomains) and stops at maxPages. It does not log in, and it reads only what a visitor can see.

How accurate are the token counts? They are estimates: the word count times 1.35, rounded up. That is close for English prose and rougher for code or other languages. Use them for budgeting and chunk sizing; if you need exact counts for a specific model, run that model's tokenizer on text.

How does the chunking work? The cleaned Markdown is split at every heading. A section that fits within maxTokensPerChunk becomes one chunk; a longer one is packed paragraph by paragraph into several chunks that each start with the section's heading. A paragraph is never cut in half, so one very long paragraph or list can produce a chunk over the limit (one of 14 chunks on the Wikipedia page in our test, at 1,201 estimated tokens against 800).

Does it work on JavaScript-heavy sites? Yes, with renderJs: true, which loads each page in a headless browser before converting it. Browser mode is slower and uses more compute, so try the default plain mode first; our tests on 2 October used plain mode.

Which proxy does it use? None. Pages are requested directly from Apify's servers, which works for public documentation and content sites. Sites that block datacenter traffic will show up as blocked or timeout rows, which are free.

How fresh is the data? Every run fetches the pages live. Nothing is cached between runs.

How do I keep the crawl focused? Start from the section you want, set maxPages, and list unwanted paths in excludeUrlPatterns as regular expressions (for example /blog/, /tag/ or \.(pdf|zip)$; inside JSON the backslash is written twice). For a site with one consistent article template, customCss gives the cleanest text.

Can I run it on a schedule? Yes. Save your input as a task, then in Apify Console open Schedules, create a schedule (for example nightly, 0 2 * * *) and pick the task. Each run crawls the pages again; key your vector store on url and chunk_index and replace a page's chunks when it is re-crawled.

How do I export the data? From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API. JSON keeps chunks intact; in CSV they are flattened into numbered columns. Skip rows with _type or status: "failed" if you want pages only.

Can I use it from Claude, ChatGPT or another AI assistant?

  • Connector URL: https://mcp.apify.com/?tools=themineworks/rag-crawler.
  • Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
  • ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
  • Cursor or VS Code: add it as an HTTP MCP server with that URL.
  • Claude Code: claude mcp add -t http rag-crawler "https://mcp.apify.com/?tools=themineworks/rag-crawler".

Is it legal to crawl websites this way? The actor reads only public pages and does not get past logins or paywalls. Website content is usually copyrighted, and many sites set terms or robots rules for automated access, so crawl sites you own, sites whose terms allow it, or content you are licensed to use, and respect data protection laws such as GDPR and CCPA for any personal data on the pages. You are responsible for how you use the data. This is general information, not legal advice.

Integrations

  • Vector databases: load chunks into Pinecone, Qdrant, Weaviate or pgvector from a webhook or a short script on the dataset.
  • Make, Zapier and n8n: use the Apify app or node to run a crawl and pass the chunks to an embedding step.
  • Google Sheets: export a run to a sheet to review titles, word counts and failures.
  • Webhooks: have Apify call your URL when a run succeeds, then fetch the dataset.
  • API and client libraries: start runs and read datasets from Python, JavaScript or any HTTP client. See the "Copy to your AI assistant" block above for the exact call.
  • MCP clients: Claude, ChatGPT, Cursor, VS Code and other MCP clients can call the actor through https://mcp.apify.com.

More from The Mine Works

Developer and AI tools

Social media and video

Leads and business directories

Marketing, SEO and reviews

LinkedIn

Real estate

Science, health and government data

Jobs and hiring

E-commerce and marketplaces

Company and business data

Food and local services

More tools

Support

Found a bug or need a field we do not return yet? Open an issue on the Issues tab of this page and we will reply there. To ask for a new source, email dmineworks@gmail.com. A guide for this actor lives at themineworks.com, with a comparison in Firecrawl vs RAG Crawler.

Website to Markdown Crawler turns a website into clean, heading-aware Markdown chunks for RAG, billed only for the pages it actually crawls.