llms.txt Generator – Make Your Website AI-Ready avatar

llms.txt Generator – Make Your Website AI-Ready

Pricing

from $1.00 / 1,000 pages

Go to Apify Store
llms.txt Generator – Make Your Website AI-Ready

llms.txt Generator – Make Your Website AI-Ready

Generate a spec-compliant llms.txt (llmstxt.org) and optional llms-full.txt for any website. Crawls the sitemap or internal links, groups pages into sections, adds an Optional section and clean Markdown full text. Ready to upload.

Pricing

from $1.00 / 1,000 pages

Rating

0.0

(0)

Developer

Deepak Ganesh

Deepak Ganesh

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

llms.txt Generator – Make Your Website AI-Ready

Generate a spec-compliant llms.txt file, plus an optional llms-full.txt, for any website. Upload the files to your site root so that ChatGPT, Claude, Perplexity, Gemini, Cursor and other AI assistants can understand your site.

  • 📄 Follows the llmstxt.org spec: # Title, > summary blockquote, ## Section headings with - [Title](url): description lists, and an ## Optional section for secondary pages. Brackets in titles are escaped, and every file is checked against the spec before it is saved.
  • 🗂️ Smart sections: pages are grouped by URL path (/docs → Docs, /blog → Blog, large areas are split further, e.g. JS / API), or put into one flat list.
  • 🧭 Covers the whole site: pages come from sitemap.xml (including sitemaps listed in robots.txt and sitemap indexes) or from following internal links. Pages are picked evenly from each part of the site, so a small maxPages still gives a full overview. Versioned doc copies and alternate-language duplicates are pushed to the back.
  • 📚 llms-full.txt: the clean Markdown content of every page in one file. Navigation, headers, footers, sidebars, scripts, forms and cookie banners are removed.
  • ✨ Clean titles and descriptions: uses og:title, then h1, then <title>, with " | Brand" suffixes removed. Descriptions come from the meta description or the first real paragraph (max 200 characters). Generic titles or descriptions repeated on many pages are replaced with page-specific text.
  • 🤝 Polite crawling: plain HTTP (no browser), at most 5 parallel requests, follows robots.txt, skips noindex pages, never logs in.
  • 💸 You only pay for pages that end up in the file. Websites that fail to load are free.

Use cases

  • Website owners and SEO / GEO teams: publish /llms.txt and /llms-full.txt so AI search and assistants describe your product correctly.
  • Docs teams: give coding assistants (Cursor, Copilot, Claude Code) an index of your documentation, or the whole docs as one Markdown file.
  • Agencies: create llms.txt files for many client sites in one run.
  • AI and RAG developers: turn any site into clean Markdown context for LLMs.

Input

FieldTypeDefaultDescription
startUrlstring–Website to process, e.g. https://example.com or example.com. A URL with a path (e.g. https://example.com/docs) limits the file to that part of the site.
startUrlsstring list–Optional: several websites. One file set is generated per website.
maxPagesinteger50Max pages per website (1–500). Each included page is one billed Page event.
includeFullTextbooleanfalseAlso generate llms-full.txt with the Markdown content of every page.
sectionStrategypath / flatpathGroup links by URL path, or put them in a single list.
useSitemapbooleantrueFind pages via sitemap.xml. Falls back to following internal links when there is no sitemap.
excludePatternsstring listlogin, signup, account, cart, checkout, tag, category, feed, search, privacy, terms, cookie, legal, author and pagination pagesURL path globs (**/login*, /blog/tag/**). Matching pages go to ## Optional and are crawled last.
dropExcludedPagesbooleanfalseSkip matching pages completely instead of listing them under ## Optional.
proxyConfigurationobjectno proxyUse Apify Proxy only if a site blocks you.

Example:

{
"startUrl": "https://llmstxt.org",
"maxPages": 50,
"includeFullText": true
}

Output

Files. Each website gets these files in the run's key-value store:

  • llms-<domain>.txt: the llms.txt file (text/plain; charset=utf-8). Rename it to llms.txt and upload it to your site root.
  • llms-full-<domain>.txt: the full-text version, if includeFullText is enabled.

The Output tab links to the dataset and to the list of files. Each dataset item also contains direct download URLs.

Dataset. One item per website. This example is trimmed from a real run:

{
"site": "https://llmstxt.org",
"success": true,
"error": null,
"siteTitle": "llms-txt",
"summary": "A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.",
"pagesCrawled": 8,
"sectionCount": 1,
"sections": [
{
"name": "Pages",
"linkCount": 8,
"links": [
{ "title": "The /llms.txt file, v2", "url": "https://llmstxt.org/", "description": "A proposal to standardise on using an /llms.txt file to provide information to help agents use a website." },
{ "title": "Python source", "url": "https://llmstxt.org/core.html", "description": "Source code for llms_txt Python module, containing helpers to create and use llms.txt files" }
]
}
],
"llmsTxt": "# llms-txt\n\n> A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.\n\n## Pages\n\n- [The /llms.txt file, v2](https://llmstxt.org/): A proposal ...\n- [Python source](https://llmstxt.org/core.html): Source code for llms_txt Python module, ...\n",
"llmsTxtKey": "llms-llmstxt.org.txt",
"llmsTxtUrl": "https://api.apify.com/v2/key-value-stores/MuEO6MqBjJWvCYjhX/records/llms-llmstxt.org.txt?signature=…",
"llmsFullTxtKey": "llms-full-llmstxt.org.txt",
"llmsFullTxtUrl": "https://api.apify.com/v2/key-value-stores/MuEO6MqBjJWvCYjhX/records/llms-full-llmstxt.org.txt?signature=…",
"warnings": [],
"generatedAt": "2026-10-07T16:36:22.683Z"
}

A website that cannot be loaded returns success: false with an error, and you are not charged for it:

{ "site": "https://this-domain-does-not-exist-xyz123.com", "success": false, "error": "RequestError: getaddrinfo ENOTFOUND this-domain-does-not-exist-xyz123.com", "pagesCrawled": 0, "llmsTxt": null }

The generated llms.txt for a larger site (crawlee.dev, trimmed) looks like this:

# Crawlee
> Crawlee helps you build and maintain your crawlers. It's open source, but built by developers who scrape millions of pages every day for a living.
## Pages
- [Build reliable crawlers. Fast.](https://crawlee.dev/): Crawlee helps you build and maintain your crawlers. ...
## Blog
- [Crawlee v3.18: Type-safe routers](https://crawlee.dev/blog/crawlee-v3-18): Crawlee v3.18 brings type-safe router labels, ...
## Python / Docs
- [Quick start](https://crawlee.dev/python/docs/quick-start): This short tutorial will help you start scraping with Crawlee in just a minute or two. ...
## JS
- [Introduction](https://crawlee.dev/js/docs/introduction): Your first steps into the world of scraping with Crawlee

The warnings array explains anything that was left out, for example: maxPages reached, pages blocked by robots.txt, noindex pages, alternate-language URLs, pages that failed to load, or a missing meta description.

Dataset views: Overview, llms.txt content and Links (one row per link).

Pricing

Pay per event: you pay only for pages included in a generated file.

EventPriceWhen it is charged
Page$0.001 ($1 per 1,000 pages)Each page included in a generated llms.txt
Actor start$0.00005Once per run (per GB of memory)
  • A 50-page site costs about $0.05.
  • Failed websites are free. So are invalid URLs, pages that return errors, and duplicate, noindex or robots-blocked pages.
  • llms-full.txt costs nothing extra.
  • If you set a maximum cost per run, the Actor stops when it is reached and still saves the files built so far.

FAQ

Where do I put the file? Rename llms-<domain>.txt to llms.txt and upload it to your website root (https://example.com/llms.txt). Do the same with llms-full.txt. On WordPress, Webflow, Shopify and similar, use a file manager, a redirect, or a plugin that serves static files.

Is the output spec-compliant? Yes. There is exactly one H1, an optional blockquote summary, only H2 section headings, and link lists in the form - [name](url): notes. Secondary pages go under ## Optional. Every file is checked by a validator before it is saved, and any problem is listed in warnings.

Why are some pages missing? maxPages limits how many pages are included. Pages are chosen evenly from each part of the site. Raise maxPages (up to 500), or start from a path like https://example.com/docs to focus on one part.

Does it work on JavaScript-heavy sites? It reads the server-rendered HTML, which works for most sites, including docs frameworks, WordPress, Webflow and Next.js. Pure client-side apps that render nothing without JavaScript will give sparse descriptions.

Does it respect robots.txt? Yes. URLs disallowed for all user agents (*) are skipped, and so are pages marked noindex.

Can I process several websites? Yes. Add them to startUrls. Each website gets its own files and dataset item.

This Actor is not affiliated with llmstxt.org, Answer.AI, OpenAI, Anthropic, Google or Perplexity.

Changelog

  • 0.1 (2026-10): First release. llms.txt and llms-full.txt generation, sitemap and link discovery, path-based sections, Optional section, spec validation, pay per page.