Website to Markdown for AI: Clean Content, RAG Chunks
Pricing
$1.00 / 1,000 page converteds
Website to Markdown for AI: Clean Content, RAG Chunks
Turn web pages or whole websites into clean Markdown for LLMs and RAG: main content only (no menus, ads or share buttons), title, author, date, word and token counts, optional chunks. Crawl a site or give a list. Fixed price per page.
Pricing
$1.00 / 1,000 page converteds
Rating
0.0
(0)
Developer
SwiftKit
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 hours ago
Last modified
Categories
Share
Feed web pages to ChatGPT, Claude, your RAG pipeline or your AI agent without the junk. Give it a list of pages or a site to crawl, and get back:
- Clean Markdown of the main content: headings, lists, tables, code blocks and links kept; menus, footers, cookie banners, ads and share buttons removed
- Metadata: title, description, author, publish date, language, site name, image, canonical URL
- Word count and approximate token count per page, so you can plan context windows and costs
- Optional RAG chunks: pieces of about N tokens, split on paragraph boundaries with overlap
- Fixed price per page. No compute-time surprises.
Main content is found with Mozilla Readability, the engine behind Firefox Reader View. Links and images are made absolute, so the Markdown works anywhere.
Who it's for
- RAG and chatbots: index docs, help centers and blogs as clean chunks.
- AI agents: read any page as Markdown with one call.
- Content teams and researchers: archive or analyze articles in a readable format.
- Training and evaluation data: clean text with word counts and dates.
Input
| Option | Default | What it does |
|---|---|---|
| URLs | – | Pages to convert, or start pages to crawl |
| Crawl the site | off | Follow links on the same site |
| Max pages | 50 | Stop after this many converted pages |
| Max link depth | 3 | How far from the start pages to crawl |
| Only / skip URLs containing | – | Keep the crawl to e.g. /docs/ |
| Content | Main content | Or the whole page minus menus and footer |
| Chunk size, overlap | 0, 50 | Tokens per RAG chunk (0 = off) |
| Include plain text, links | off | Extra fields |
| Respect robots.txt | on | Skip pages the site asks bots not to visit |
Output
Real result, shortened:
{"url": "https://github.blog/engineering/the-cost-of-saying-yes-has-changed/","status": "ok","title": "The cost of saying yes has changed","author": "Dalia Abuadas","published": "2026-07-17T16:46:47+00:00","language": "en-US","markdown": "The cost of writing code dropped; the cost of owning it didn’t. …","wordCount": 1174,"approxTokens": 1877,"chunks": [{ "index": 0, "text": "…", "approxTokens": "…" }]}
| status | meaning |
|---|---|
ok | Converted. |
unreachable | The page didn't load. Not charged. |
not_html | Not a web page (PDF, image…). Not charged. |
blocked_by_robots_txt | The site asks bots not to visit. Not charged. |
blocked_by_bot_protection | The site showed a bot check. We don't try to get around it. Not charged. |
Pricing
You pay per page converted, the same for every page size. Failed and blocked pages are free. See the Pricing tab.
Limits, honestly
- No browser is used. That keeps it fast and cheap, but pages that build their content with JavaScript after loading (some single-page apps) may come back thin, and crawling them can find few links. Most blogs, docs, news and company sites work well.
- Token counts are approximate (about 4 characters per token). Exact counts depend on your model.
- Pages behind a login, a paywall or a bot check aren't converted.
- You're responsible for how you use the content; respect each site's terms and copyright.
More tools from SwiftKit
- Sitemap Extractor & Broken Link Checker: every URL from a site’s sitemaps, plus 404 and redirect checks
- RSS Feed Reader: news, blogs and podcasts from any feed or website, only new items
- Website Screenshot Tool: bulk full-page and mobile screenshots, PNG/JPEG/PDF
- PDF to Text & Markdown: clean text from PDF links for AI, priced per page
Questions?
Open an issue on the Issues tab.