Website to Markdown – URL & Site Crawler for LLM/RAG
Pricing
from $0.40 / 1,000 page converted to markdowns
Website to Markdown – URL & Site Crawler for LLM/RAG
Convert any URL, sitemap or whole website to clean Markdown for LLMs, RAG and AI agents. Removes nav, footers and cookie banners, keeps headings, tables and code, adds RAG-ready chunks with token counts, llms.txt generation and an only-changed-pages mode. No browser: fast and cheap.
Pricing
from $0.40 / 1,000 page converted to markdowns
Rating
0.0
(0)
Developer
Cemal Atakli
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 days ago
Last modified
Categories
Share
Convert any URL, sitemap or whole website to clean Markdown for LLMs, RAG pipelines, vector databases and AI agents. The Actor removes navigation, headers, footers, sidebars and cookie banners. It keeps headings, lists, tables, code blocks and absolute links, and splits every page into RAG-ready chunks with token counts and heading paths. It can also write an llms.txt / llms-full.txt for a site and output only new or changed pages on scheduled re-indexing runs.
There is no headless browser, so it is fast and costs $0.40 per 1,000 pages.
What it does
- Three input modes:
- A list of URLs.
- A same-domain crawl from a start URL, with max depth and include/exclude globs.
- Every URL in a sitemap: give a
sitemap.xml, or just the site and the sitemap is found via robots.txt. Sitemap indexes and.xml.gzwork.
- Main-content extraction:
- Scores
<main>,<article>,role=mainand common content containers, with a text-density fallback. - Strips scripts, styles, forms, nav, footer, aside and cookie or consent banners.
- Scores
- High-quality HTML → Markdown:
- Headings, nested lists, GFM tables, fenced code blocks with the language, blockquotes and bold/italic.
- Absolute links. Images are optional.
- RAG chunks:
chunks[]with{index, headingPath, text, tokens}. Chunks split on headings and paragraphs, and code blocks are never cut. The size is configurable (default 800 tokens). - llms.txt generator (llmstxt.org format): pages are grouped by path into sections, as
- [title](url): description.llms-full.txtholds all the Markdown. - Only-changed mode for scheduled runs. ETag, Last-Modified and SHA-256 content hash are remembered between runs, so you pay only for pages that are new or changed.
- JS-only pages are detected (
needsBrowser: true) and not charged. - Polite by default:
- Respects robots.txt and Crawl-delay.
- 2 requests per host.
- Retries with backoff on 429/5xx.
- Honest User-Agent with a contact address.
Use cases
- Feed documentation sites into a RAG / vector database (Pinecone, Qdrant, Weaviate, pgvector) using the ready-made chunks.
- Give an AI agent a web-fetch tool that returns clean Markdown instead of raw HTML.
- Generate
llms.txtfor your own site or a competitor's docs. - Re-index a knowledge base nightly with
onlyChanged, so you pay only for pages that changed. - Build fine-tuning or evaluation datasets from blogs, docs and help centres.
Input example
{"startUrls": [{ "url": "https://docs.apify.com/sitemap_base.xml" }],"crawlMode": "sitemap","maxPages": 200,"includeGlobs": ["https://docs.apify.com/academy/**"],"chunkSize": 800,"generateLlmsTxt": true,"onlyChanged": false}
The default input converts 2 pages in a few seconds for less than $0.001.
| Field | Default | Notes |
|---|---|---|
startUrls | 2 sample pages | URLs, a site, or a sitemap URL |
crawlMode | single | single / sameDomain / sitemap |
maxPages | 20 | Hard cap on pages fetched |
maxDepth | 2 | Same-domain crawl only |
includeGlobs / excludeGlobs | – | e.g. https://example.com/docs/**, **/tag/** |
chunkSize | 800 | Tokens per chunk, 0 = off |
contentMode | main | main = article/docs body only, full = whole page |
keepLinksInMarkdown / includeImages | true / false | Token control |
generateLlmsTxt | false | One llms.txt + llms-full.txt per site |
onlyChanged | false | Output only new or changed pages since the last run |
respectRobots | true |
Output example
One dataset row per page. The Overview, Markdown and Metadata views are in the Output tab.
{"url": "https://docs.apify.com/academy/scraping-basics-javascript","finalUrl": "https://docs.apify.com/academy/scraping-basics-javascript","status": "ok","httpStatus": 200,"title": "Web scraping basics for JavaScript devs | Academy | Apify Documentation","description": "Learn how to use JavaScript to extract information from websites ...","lang": "en","canonical": "https://docs.apify.com/academy/scraping-basics-javascript","wordCount": 979,"tokenCount": 1580,"markdown": "# Web scraping basics for JavaScript devs\n\n**Learn how to use JavaScript ...**\n\n## What we'll do\n\n- Inspect pages using browser DevTools.\n...","chunks": [{ "index": 0, "headingPath": "Web scraping basics for JavaScript devs", "text": "# Web scraping basics ...", "tokens": 283 },{ "index": 1, "headingPath": "Web scraping basics for JavaScript devs > Requirements", "text": "## Requirements ...", "tokens": 350 }],"links": ["https://docs.apify.com/get-started", "..."],"contentHash": "2573bbdf63d0cd97...","depth": 0,"needsBrowser": false,"fetchedAt": "2026-10-01T09:00:12+00:00","error": null}
statusis one of:ok: charged.needsBrowser: JS-rendered page, not charged.failed: HTTP error or timeout, not charged.skipped: PDF or other non-HTML, or blocked by robots.txt. Not charged.
- With
onlyChanged, each row also haschangeStatus(new/changed). - The key-value store holds
llms.txt,llms-full.txt(with several sites:llms-<host>.txt) and anOUTPUTrun summary.
Pricing
Pay per event. You pay only for pages that are converted successfully.
| Event | Price |
|---|---|
| Page converted to Markdown (chunks, links and metadata included) | $0.0004 ($0.40 / 1,000) |
| llms.txt + llms-full.txt generated, per site | $0.005 |
| Actor start | $0.00005 |
Compared with other Store Actors (public Store prices, 2026-10-01):
| Actor | Price per 1,000 pages |
|---|---|
| Website to Markdown (this Actor) | $0.40 |
| apify/web-fetch | $1.50 |
| 6sigmag/fast-website-content-crawler | $3.00 |
| parseforge | $25.00 |
Set Maximum cost per run on the run options to cap spending. The Actor stops cleanly when the limit is reached, and the money for llms.txt is held back in advance.
FAQ
Does it render JavaScript?
No. It uses plain HTTP, which is why it is fast and cheap. Pages that only render with JavaScript are flagged needsBrowser: true and are not charged. Most docs sites, blogs, news sites and help centres are server-rendered and work well.
How are tokens counted? As characters / 4. This is a fast estimate that is close to the OpenAI and Anthropic tokenizers for English text.
What happens with PDFs, images and other files?
Links to them appear in links[], but they are not downloaded or converted. A PDF given as a start URL is returned as skipped.
How does only-changed mode work?
State is kept in a named key-value store (wtm-state-…) derived from your start URLs, or set stateKey. Re-runs send If-None-Match / If-Modified-Since and compare content hashes. Unchanged pages are neither output nor charged.
Can I crawl only part of a site?
Yes. Use includeGlobs, e.g. https://example.com/docs/**, together with maxDepth and maxPages.
Is it legal? The Actor fetches only the public URLs you supply, respects robots.txt by default and identifies itself. You are responsible for having the rights to use the content you convert, for example under the site's terms and copyright.
Use with AI agents / Apify MCP
The Actor works as a web-fetch tool for LLM agents.
- Apify MCP server: add
gazidev/website-to-markdownto your MCP client (Claude Desktop, Cursor, VS Code) throughhttps://mcp.apify.com?actors=gazidev/website-to-markdown. The agent can then call it with{"startUrls":[{"url":"..."}]}and get Markdown back. - API:
POST https://api.apify.com/v2/acts/gazidev~website-to-markdown/run-sync-get-dataset-items?token=...with the input JSON. It returns the rows directly, which suits LangChain / LlamaIndex loaders. - RAG pipelines: embed
chunks[].textand storeurlandheadingPathas metadata for citations.