Website to Markdown – URL & Site Crawler for LLM/RAG avatar

Website to Markdown – URL & Site Crawler for LLM/RAG

Pricing

from $0.40 / 1,000 page converted to markdowns

Go to Apify Store
Website to Markdown – URL & Site Crawler for LLM/RAG

Website to Markdown – URL & Site Crawler for LLM/RAG

Convert any URL, sitemap or whole website to clean Markdown for LLMs, RAG and AI agents. Removes nav, footers and cookie banners, keeps headings, tables and code, adds RAG-ready chunks with token counts, llms.txt generation and an only-changed-pages mode. No browser: fast and cheap.

Pricing

from $0.40 / 1,000 page converted to markdowns

Rating

0.0

(0)

Developer

Cemal Atakli

Cemal Atakli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 days ago

Last modified

Categories

Share

Convert any URL, sitemap or whole website to clean Markdown for LLMs, RAG pipelines, vector databases and AI agents. The Actor removes navigation, headers, footers, sidebars and cookie banners. It keeps headings, lists, tables, code blocks and absolute links, and splits every page into RAG-ready chunks with token counts and heading paths. It can also write an llms.txt / llms-full.txt for a site and output only new or changed pages on scheduled re-indexing runs.

There is no headless browser, so it is fast and costs $0.40 per 1,000 pages.

What it does

  • Three input modes:
    • A list of URLs.
    • A same-domain crawl from a start URL, with max depth and include/exclude globs.
    • Every URL in a sitemap: give a sitemap.xml, or just the site and the sitemap is found via robots.txt. Sitemap indexes and .xml.gz work.
  • Main-content extraction:
    • Scores <main>, <article>, role=main and common content containers, with a text-density fallback.
    • Strips scripts, styles, forms, nav, footer, aside and cookie or consent banners.
  • High-quality HTML → Markdown:
    • Headings, nested lists, GFM tables, fenced code blocks with the language, blockquotes and bold/italic.
    • Absolute links. Images are optional.
  • RAG chunks: chunks[] with {index, headingPath, text, tokens}. Chunks split on headings and paragraphs, and code blocks are never cut. The size is configurable (default 800 tokens).
  • llms.txt generator (llmstxt.org format): pages are grouped by path into sections, as - [title](url): description. llms-full.txt holds all the Markdown.
  • Only-changed mode for scheduled runs. ETag, Last-Modified and SHA-256 content hash are remembered between runs, so you pay only for pages that are new or changed.
  • JS-only pages are detected (needsBrowser: true) and not charged.
  • Polite by default:
    • Respects robots.txt and Crawl-delay.
    • 2 requests per host.
    • Retries with backoff on 429/5xx.
    • Honest User-Agent with a contact address.

Use cases

  • Feed documentation sites into a RAG / vector database (Pinecone, Qdrant, Weaviate, pgvector) using the ready-made chunks.
  • Give an AI agent a web-fetch tool that returns clean Markdown instead of raw HTML.
  • Generate llms.txt for your own site or a competitor's docs.
  • Re-index a knowledge base nightly with onlyChanged, so you pay only for pages that changed.
  • Build fine-tuning or evaluation datasets from blogs, docs and help centres.

Input example

{
"startUrls": [{ "url": "https://docs.apify.com/sitemap_base.xml" }],
"crawlMode": "sitemap",
"maxPages": 200,
"includeGlobs": ["https://docs.apify.com/academy/**"],
"chunkSize": 800,
"generateLlmsTxt": true,
"onlyChanged": false
}

The default input converts 2 pages in a few seconds for less than $0.001.

FieldDefaultNotes
startUrls2 sample pagesURLs, a site, or a sitemap URL
crawlModesinglesingle / sameDomain / sitemap
maxPages20Hard cap on pages fetched
maxDepth2Same-domain crawl only
includeGlobs / excludeGlobs–e.g. https://example.com/docs/**, **/tag/**
chunkSize800Tokens per chunk, 0 = off
contentModemainmain = article/docs body only, full = whole page
keepLinksInMarkdown / includeImagestrue / falseToken control
generateLlmsTxtfalseOne llms.txt + llms-full.txt per site
onlyChangedfalseOutput only new or changed pages since the last run
respectRobotstrue

Output example

One dataset row per page. The Overview, Markdown and Metadata views are in the Output tab.

{
"url": "https://docs.apify.com/academy/scraping-basics-javascript",
"finalUrl": "https://docs.apify.com/academy/scraping-basics-javascript",
"status": "ok",
"httpStatus": 200,
"title": "Web scraping basics for JavaScript devs | Academy | Apify Documentation",
"description": "Learn how to use JavaScript to extract information from websites ...",
"lang": "en",
"canonical": "https://docs.apify.com/academy/scraping-basics-javascript",
"wordCount": 979,
"tokenCount": 1580,
"markdown": "# Web scraping basics for JavaScript devs\n\n**Learn how to use JavaScript ...**\n\n## What we'll do\n\n- Inspect pages using browser DevTools.\n...",
"chunks": [
{ "index": 0, "headingPath": "Web scraping basics for JavaScript devs", "text": "# Web scraping basics ...", "tokens": 283 },
{ "index": 1, "headingPath": "Web scraping basics for JavaScript devs > Requirements", "text": "## Requirements ...", "tokens": 350 }
],
"links": ["https://docs.apify.com/get-started", "..."],
"contentHash": "2573bbdf63d0cd97...",
"depth": 0,
"needsBrowser": false,
"fetchedAt": "2026-10-01T09:00:12+00:00",
"error": null
}
  • status is one of:
    • ok: charged.
    • needsBrowser: JS-rendered page, not charged.
    • failed: HTTP error or timeout, not charged.
    • skipped: PDF or other non-HTML, or blocked by robots.txt. Not charged.
  • With onlyChanged, each row also has changeStatus (new / changed).
  • The key-value store holds llms.txt, llms-full.txt (with several sites: llms-<host>.txt) and an OUTPUT run summary.

Pricing

Pay per event. You pay only for pages that are converted successfully.

EventPrice
Page converted to Markdown (chunks, links and metadata included)$0.0004 ($0.40 / 1,000)
llms.txt + llms-full.txt generated, per site$0.005
Actor start$0.00005

Compared with other Store Actors (public Store prices, 2026-10-01):

ActorPrice per 1,000 pages
Website to Markdown (this Actor)$0.40
apify/web-fetch$1.50
6sigmag/fast-website-content-crawler$3.00
parseforge$25.00

Set Maximum cost per run on the run options to cap spending. The Actor stops cleanly when the limit is reached, and the money for llms.txt is held back in advance.

FAQ

Does it render JavaScript? No. It uses plain HTTP, which is why it is fast and cheap. Pages that only render with JavaScript are flagged needsBrowser: true and are not charged. Most docs sites, blogs, news sites and help centres are server-rendered and work well.

How are tokens counted? As characters / 4. This is a fast estimate that is close to the OpenAI and Anthropic tokenizers for English text.

What happens with PDFs, images and other files? Links to them appear in links[], but they are not downloaded or converted. A PDF given as a start URL is returned as skipped.

How does only-changed mode work? State is kept in a named key-value store (wtm-state-…) derived from your start URLs, or set stateKey. Re-runs send If-None-Match / If-Modified-Since and compare content hashes. Unchanged pages are neither output nor charged.

Can I crawl only part of a site? Yes. Use includeGlobs, e.g. https://example.com/docs/**, together with maxDepth and maxPages.

Is it legal? The Actor fetches only the public URLs you supply, respects robots.txt by default and identifies itself. You are responsible for having the rights to use the content you convert, for example under the site's terms and copyright.

Use with AI agents / Apify MCP

The Actor works as a web-fetch tool for LLM agents.

  • Apify MCP server: add gazidev/website-to-markdown to your MCP client (Claude Desktop, Cursor, VS Code) through https://mcp.apify.com?actors=gazidev/website-to-markdown. The agent can then call it with {"startUrls":[{"url":"..."}]} and get Markdown back.
  • API: POST https://api.apify.com/v2/acts/gazidev~website-to-markdown/run-sync-get-dataset-items?token=... with the input JSON. It returns the rows directly, which suits LangChain / LlamaIndex loaders.
  • RAG pipelines: embed chunks[].text and store url and headingPath as metadata for citations.