llms.txt Generator – Make Your Website AI-Ready
Pricing
from $1.00 / 1,000 pages
llms.txt Generator – Make Your Website AI-Ready
Generate a spec-compliant llms.txt (llmstxt.org) and optional llms-full.txt for any website. Crawls the sitemap or internal links, groups pages into sections, adds an Optional section and clean Markdown full text. Ready to upload.
Pricing
from $1.00 / 1,000 pages
Rating
0.0
(0)
Developer
Deepak Ganesh
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share

Generate a spec-compliant llms.txt file, plus an optional llms-full.txt, for any website. Upload the files to your site root so that ChatGPT, Claude, Perplexity, Gemini, Cursor and other AI assistants can understand your site.
- 📄 Follows the llmstxt.org spec:
# Title,> summaryblockquote,## Sectionheadings with- [Title](url): descriptionlists, and an## Optionalsection for secondary pages. Brackets in titles are escaped, and every file is checked against the spec before it is saved. - 🗂️ Smart sections: pages are grouped by URL path (
/docs→ Docs,/blog→ Blog, large areas are split further, e.g. JS / API), or put into one flat list. - 🧭 Covers the whole site: pages come from
sitemap.xml(including sitemaps listed in robots.txt and sitemap indexes) or from following internal links. Pages are picked evenly from each part of the site, so a smallmaxPagesstill gives a full overview. Versioned doc copies and alternate-language duplicates are pushed to the back. - 📚 llms-full.txt: the clean Markdown content of every page in one file. Navigation, headers, footers, sidebars, scripts, forms and cookie banners are removed.
- ✨ Clean titles and descriptions: uses og:title, then h1, then
<title>, with " | Brand" suffixes removed. Descriptions come from the meta description or the first real paragraph (max 200 characters). Generic titles or descriptions repeated on many pages are replaced with page-specific text. - 🤝 Polite crawling: plain HTTP (no browser), at most 5 parallel requests, follows robots.txt, skips
noindexpages, never logs in. - 💸 You only pay for pages that end up in the file. Websites that fail to load are free.
Use cases
- Website owners and SEO / GEO teams: publish
/llms.txtand/llms-full.txtso AI search and assistants describe your product correctly. - Docs teams: give coding assistants (Cursor, Copilot, Claude Code) an index of your documentation, or the whole docs as one Markdown file.
- Agencies: create llms.txt files for many client sites in one run.
- AI and RAG developers: turn any site into clean Markdown context for LLMs.
Input
| Field | Type | Default | Description |
|---|---|---|---|
startUrl | string | – | Website to process, e.g. https://example.com or example.com. A URL with a path (e.g. https://example.com/docs) limits the file to that part of the site. |
startUrls | string list | – | Optional: several websites. One file set is generated per website. |
maxPages | integer | 50 | Max pages per website (1–500). Each included page is one billed Page event. |
includeFullText | boolean | false | Also generate llms-full.txt with the Markdown content of every page. |
sectionStrategy | path / flat | path | Group links by URL path, or put them in a single list. |
useSitemap | boolean | true | Find pages via sitemap.xml. Falls back to following internal links when there is no sitemap. |
excludePatterns | string list | login, signup, account, cart, checkout, tag, category, feed, search, privacy, terms, cookie, legal, author and pagination pages | URL path globs (**/login*, /blog/tag/**). Matching pages go to ## Optional and are crawled last. |
dropExcludedPages | boolean | false | Skip matching pages completely instead of listing them under ## Optional. |
proxyConfiguration | object | no proxy | Use Apify Proxy only if a site blocks you. |
Example:
{"startUrl": "https://llmstxt.org","maxPages": 50,"includeFullText": true}
Output
Files. Each website gets these files in the run's key-value store:
llms-<domain>.txt: the llms.txt file (text/plain; charset=utf-8). Rename it tollms.txtand upload it to your site root.llms-full-<domain>.txt: the full-text version, ifincludeFullTextis enabled.
The Output tab links to the dataset and to the list of files. Each dataset item also contains direct download URLs.
Dataset. One item per website. This example is trimmed from a real run:
{"site": "https://llmstxt.org","success": true,"error": null,"siteTitle": "llms-txt","summary": "A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.","pagesCrawled": 8,"sectionCount": 1,"sections": [{"name": "Pages","linkCount": 8,"links": [{ "title": "The /llms.txt file, v2", "url": "https://llmstxt.org/", "description": "A proposal to standardise on using an /llms.txt file to provide information to help agents use a website." },{ "title": "Python source", "url": "https://llmstxt.org/core.html", "description": "Source code for llms_txt Python module, containing helpers to create and use llms.txt files" }]}],"llmsTxt": "# llms-txt\n\n> A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.\n\n## Pages\n\n- [The /llms.txt file, v2](https://llmstxt.org/): A proposal ...\n- [Python source](https://llmstxt.org/core.html): Source code for llms_txt Python module, ...\n","llmsTxtKey": "llms-llmstxt.org.txt","llmsTxtUrl": "https://api.apify.com/v2/key-value-stores/MuEO6MqBjJWvCYjhX/records/llms-llmstxt.org.txt?signature=…","llmsFullTxtKey": "llms-full-llmstxt.org.txt","llmsFullTxtUrl": "https://api.apify.com/v2/key-value-stores/MuEO6MqBjJWvCYjhX/records/llms-full-llmstxt.org.txt?signature=…","warnings": [],"generatedAt": "2026-10-07T16:36:22.683Z"}
A website that cannot be loaded returns success: false with an error, and you are not charged for it:
{ "site": "https://this-domain-does-not-exist-xyz123.com", "success": false, "error": "RequestError: getaddrinfo ENOTFOUND this-domain-does-not-exist-xyz123.com", "pagesCrawled": 0, "llmsTxt": null }
The generated llms.txt for a larger site (crawlee.dev, trimmed) looks like this:
# Crawlee> Crawlee helps you build and maintain your crawlers. It's open source, but built by developers who scrape millions of pages every day for a living.## Pages- [Build reliable crawlers. Fast.](https://crawlee.dev/): Crawlee helps you build and maintain your crawlers. ...## Blog- [Crawlee v3.18: Type-safe routers](https://crawlee.dev/blog/crawlee-v3-18): Crawlee v3.18 brings type-safe router labels, ...## Python / Docs- [Quick start](https://crawlee.dev/python/docs/quick-start): This short tutorial will help you start scraping with Crawlee in just a minute or two. ...## JS- [Introduction](https://crawlee.dev/js/docs/introduction): Your first steps into the world of scraping with Crawlee
The warnings array explains anything that was left out, for example: maxPages reached, pages blocked by robots.txt, noindex pages, alternate-language URLs, pages that failed to load, or a missing meta description.
Dataset views: Overview, llms.txt content and Links (one row per link).
Pricing
Pay per event: you pay only for pages included in a generated file.
| Event | Price | When it is charged |
|---|---|---|
| Page | $0.001 ($1 per 1,000 pages) | Each page included in a generated llms.txt |
| Actor start | $0.00005 | Once per run (per GB of memory) |
- A 50-page site costs about $0.05.
- Failed websites are free. So are invalid URLs, pages that return errors, and duplicate,
noindexor robots-blocked pages. llms-full.txtcosts nothing extra.- If you set a maximum cost per run, the Actor stops when it is reached and still saves the files built so far.
FAQ
Where do I put the file? Rename llms-<domain>.txt to llms.txt and upload it to your website root (https://example.com/llms.txt). Do the same with llms-full.txt. On WordPress, Webflow, Shopify and similar, use a file manager, a redirect, or a plugin that serves static files.
Is the output spec-compliant? Yes. There is exactly one H1, an optional blockquote summary, only H2 section headings, and link lists in the form - [name](url): notes. Secondary pages go under ## Optional. Every file is checked by a validator before it is saved, and any problem is listed in warnings.
Why are some pages missing? maxPages limits how many pages are included. Pages are chosen evenly from each part of the site. Raise maxPages (up to 500), or start from a path like https://example.com/docs to focus on one part.
Does it work on JavaScript-heavy sites? It reads the server-rendered HTML, which works for most sites, including docs frameworks, WordPress, Webflow and Next.js. Pure client-side apps that render nothing without JavaScript will give sparse descriptions.
Does it respect robots.txt? Yes. URLs disallowed for all user agents (*) are skipped, and so are pages marked noindex.
Can I process several websites? Yes. Add them to startUrls. Each website gets its own files and dataset item.
This Actor is not affiliated with llmstxt.org, Answer.AI, OpenAI, Anthropic, Google or Perplexity.
Changelog
- 0.1 (2026-10): First release. llms.txt and llms-full.txt generation, sitemap and link discovery, path-based sections, Optional section, spec validation, pay per page.