llms.txt Generator - Make Your Site LLM-Readable
Pricing
Pay per event
llms.txt Generator - Make Your Site LLM-Readable
Generate a ready-to-upload llms.txt for any website. Crawl from a start URL or read an XML sitemap, group pages into sections and build the file from each page's own title and meta description. One dataset row per page so you can review before publishing.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Cuantic Data
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Generate a ready-to-upload llms.txt for your website. Point the actor at a start page or an XML sitemap, and get the finished file plus one dataset row per page, so you can review every entry before you publish.
What is llms.txt?
llms.txt is a proposed standard: a Markdown file served at the root of a site (https://your-site.com/llms.txt) that gives large language models and AI assistants a short, curated map of the site. It starts with the site name as an H1 heading, an optional one-line summary in a blockquote, and then sections of links, each with a short description. AI tools and agents can read it to find the right pages quickly instead of parsing every HTML page.
What it does
- Crawls your site from a start URL (same-origin links only, up to a link depth you choose), or reads the pages listed in an XML sitemap (a sitemap index is followed one level deep).
- Fetches each page with a plain HTTP request and reads its own metadata: the title (
og:title, otherwise<title>, otherwise the first<h1>) and the description (<meta name="description">, otherwiseog:description). - Groups pages into sections by the first path segment:
/docs/...goes to Docs,/blog/...to Blog,/api-reference/...to Api Reference. Root pages (/) go to Pages, and so does any page whose first path segment no other listed page shares (for example a lone/pricingor/en), so the file has no one-page sections. - Leaves out a page's description when it is identical to the site summary, so a meta description repeated across the whole site is not repeated on every line.
- Assembles the
llms.txtfile per the llmstxt.org format and stores it in the key-value store recordllms.txt. - Writes one dataset row per page (included or not, with the reason), so you can see exactly what went into the file and why a page was left out.
- Respects
robots.txtby default, removes duplicate URLs, skips binary and asset files (PDF, images, CSS, JS, fonts, video, XML, ZIP) and isolates failures: a page that cannot be fetched gets an error row and the run continues.
Who it is for
- Documentation and developer portals that want AI assistants to point users to the right guide or API page.
- Product and SaaS sites that want to be described accurately by AI search and chat tools.
- Agencies and SEO teams preparing
llms.txtfiles for several client sites without writing them by hand.
Input
| Field | Type | Default | Description |
|---|---|---|---|
startUrl | string | - | Page where the crawl starts. Only same-origin links are followed. |
sitemapUrl | string | - | XML sitemap or sitemap index. When set, pages come from its <loc> entries and no links are followed. |
maxPages | integer | 100 | Pages processed per run, 1 to 500 (fetched pages plus pages skipped by robots.txt). |
maxDepth | integer | 3 | Link depth from the start URL, 0 to 6. 0 fetches the start page only. Crawl mode only. |
includePatterns | array of strings | - | When set, only URLs whose path (with query string) contains one of these substrings are kept. |
excludePatterns | array of strings | - | URLs whose path (with query string) contains any of these substrings are skipped. Applied after the include patterns. |
siteName | string | og:site_name of the start page, else its hostname | H1 heading of the file. |
siteSummary | string | meta description of the start page | Blockquote summary. Omitted when there is none. |
respectRobotsTxt | boolean | true | Read robots.txt once per site and skip disallowed pages. |
maxConcurrency | integer | 4 | Pages fetched in parallel, 1 to 10. |
At least one of startUrl or sitemapUrl is required. If both are given, the sitemap is used, a warning is logged, and the start URL only provides the site name and summary.
In crawl mode the start page is always fetched (its links and metadata are needed), even when it does not match the include patterns; in that case it is simply not listed in the file. In sitemap mode, if the site's home page is not in the sitemap, it is fetched once for the site name and summary only (no dataset row, no charge).
Example input:
{"startUrl": "https://docs.example-app.com/","maxPages": 200,"maxDepth": 3,"excludePatterns": ["/changelog/", "?page="],"respectRobotsTxt": true,"maxConcurrency": 4}
Output
The llms.txt file
Stored in the default key-value store under the key llms.txt (Markdown). Download it and upload it to the root of your site. Example:
# Example App> Example App is a hosted API for sending transactional email.## Pages- [Example App - Transactional email API](https://docs.example-app.com/)- [Pricing](https://docs.example-app.com/pricing): Free tier, then pay per email sent.## Guides- [Quickstart](https://docs.example-app.com/guides/quickstart): Send your first email in five minutes.- [Domains and DNS](https://docs.example-app.com/guides/domains): Verify a sending domain with SPF, DKIM and DMARC.- [Webhooks](https://docs.example-app.com/guides/webhooks): Receive delivery, bounce and complaint events.## Api Reference- [Send an email](https://docs.example-app.com/api-reference/send): POST /v1/emails parameters and responses.- [Errors](https://docs.example-app.com/api-reference/errors)
Sections are ordered with Pages first, then by number of pages (largest first). Inside a section, pages keep the order in which they were discovered. A page without a description, or whose description is identical to the site summary, gets a link line without the : description part.
Dataset (one row per page)
{"url": "https://docs.example-app.com/guides/quickstart","finalUrl": "https://docs.example-app.com/guides/quickstart","httpStatus": 200,"title": "Quickstart","description": "Send your first email in five minutes.","section": "Guides","included": true,"reason": null,"error": null}
finalUrlis the address after redirects; that is the URL written to the file.sectionanddescriptionare the page's own values. In the file, a page whose section would hold only that page is listed under Pages, and a description identical to the site summary is omitted.includedtells whether the page is listed inllms.txt. When it isfalse,reasonsays why:no title,not an HTML page,disallowed by robots.txt,filtered by include/exclude patterns,redirected to another site,duplicate of <url>(two URLs that redirect to the same page) orfetch failed.errorholds the error message when the page could not be fetched (for exampleHTTP 404 for ...), otherwisenull.
Rows are written as each page finishes, so their order follows completion, not discovery.
Summary
The key-value store record OUTPUT:
{"pagesCrawled": 137,"pagesIncluded": 128,"pagesSkipped": 9,"pagesFailed": 2,"chargeLimitReached": false,"sections": [{ "name": "Pages", "pages": 4 },{ "name": "Guides", "pages": 71 },{ "name": "Api Reference", "pages": 53 }],"llmsTxtChars": 14210}
pagesCrawled counts pages requested over HTTP; pagesSkipped counts dataset rows not listed in the file (errors included).
Combine with robots-llms-txt-monitor
Generate your llms.txt here, then monitor it with robots-llms-txt-monitor to make sure it stays live and valid after you publish it (and that your robots.txt does not accidentally block the crawlers you care about).
Limitations
- Plain HTML fetch only. Pages are fetched without running JavaScript. Content that only appears after client-side rendering (single-page apps) can be missed, including links, titles and descriptions set by scripts.
- No AI generation or summarization. Titles and descriptions come verbatim from each page's own meta tags. Weak or missing meta tags give weak or missing descriptions; improve the meta tags or edit the file.
- Sections follow URL structure. The section is the first path segment, so on a site where every page sits under a locale prefix (
/en/...) most pages land in one section named after it. Use include/exclude patterns or edit the file to adjust. - Include/exclude patterns apply to the discovered URL, before redirects.
- robots.txt is respected by default. Disallowed pages are not fetched and get a
disallowed by robots.txtrow. You can disable this withrespectRobotsTxt: false, which is only appropriate for sites you own or are allowed to crawl. - Max 500 pages per run. Extra URLs are dropped with a warning.
- Binary and asset URLs are skipped by file extension, and responses that are not HTML are listed as
not an HTML page. Response bodies over 5 MB are refused. - Plain XML sitemaps only. Gzipped (
.xml.gz) sitemaps are not supported; a sitemap index is followed one level deep (up to 50 child sitemaps). - Review the generated file before publishing it. It is a starting point built from your pages' metadata: check the titles, remove pages that should not be recommended to AI tools and add context where it helps.
Pricing
Pay-per-event: $0.002 USD per page fetched and parsed successfully ($2.00 per 1,000 pages). Pages that fail, are not HTML or are skipped because of robots.txt are not charged. Apify platform usage is included in this price; there is no separate platform-usage charge.
An Actor Start event costs $0.00005 USD. One start event is charged per GB of Actor memory, with a minimum of one event per run.
If the maximum charge you set for a run is reached, the actor stops starting new pages, finishes the ones in progress and still writes the llms.txt from the pages it has.
Support
Cuantic Data - cuanticwindows@gmail.com
Terms of use: see Terms of use below.
Terms of use
Provided by Cuantic Data (cuanticwindows@gmail.com).
1. What the actor does
The actor fetches public pages of the website you point it to, either by following links from a start URL or by reading the XML sitemap you provide. It makes plain HTTP requests (no browser, no JavaScript), reads each page's title and description from its own HTML metadata, and stores the resulting rows and the generated llms.txt file in your Apify storage.
2. Your responsibility
- You are responsible for having the right to crawl the site you submit. Only submit sites you own, manage, or are otherwise allowed to crawl.
- robots.txt is respected by default. If you disable that option, you confirm that you are allowed to fetch the disallowed pages.
- Each run generates real traffic on the target site; choose the number of pages and the concurrency accordingly.
- You are responsible for reviewing the generated
llms.txtbefore publishing it, and for what you publish on your site. - You must comply with the terms of service of the crawled site and with the Apify Terms of Service.
3. About the results
- Titles and descriptions are copied from the pages' own metadata; the actor does not write, summarize or verify content. Pages that rely on JavaScript to render can be missed or listed with incomplete metadata.
- llms.txt is a proposed convention (https://llmstxt.org). Publishing the file does not guarantee that any AI system will read it or use it in a particular way.
- Results are provided "as is", for informational purposes, without warranty of accuracy, completeness or fitness for a particular purpose.
4. Data
The actor stores only what it produces (the dataset rows, the llms.txt file and the summary record) in your own Apify storage. It does not keep copies of the crawled pages or results elsewhere.
5. Liability
To the extent permitted by law, Cuantic Data is not liable for any damage arising from the use of the actor or its results, including the content of a published llms.txt file or any effect of the crawl traffic on the crawled site.
6. Changes
These terms may be updated together with the actor. The version published with the actor is the one that applies.