# llms.txt Generator - Make Your Site LLM-Readable (`cuantic_data/llms-txt-generator`) Actor

Generate a ready-to-upload llms.txt for any website. Crawl from a start URL or read an XML sitemap, group pages into sections and build the file from each page's own title and meta description. One dataset row per page so you can review before publishing.

- **URL**: https://apify.com/cuantic\_data/llms-txt-generator.md
- **Developed by:** [Cuantic Data](https://apify.com/cuantic_data) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## llms.txt Generator - Make Your Site LLM-Readable

Generate a ready-to-upload `llms.txt` for your website. Point the actor at a start page or an XML sitemap, and get the finished file plus one dataset row per page, so you can review every entry before you publish.

### What is llms.txt?

[llms.txt](https://llmstxt.org) is a proposed standard: a Markdown file served at the root of a site (`https://your-site.com/llms.txt`) that gives large language models and AI assistants a short, curated map of the site. It starts with the site name as an H1 heading, an optional one-line summary in a blockquote, and then sections of links, each with a short description. AI tools and agents can read it to find the right pages quickly instead of parsing every HTML page.

### What it does

- Crawls your site from a start URL (same-origin links only, up to a link depth you choose), **or** reads the pages listed in an XML sitemap (a sitemap index is followed one level deep).
- Fetches each page with a plain HTTP request and reads its own metadata: the title (`og:title`, otherwise `<title>`, otherwise the first `<h1>`) and the description (`<meta name="description">`, otherwise `og:description`).
- Groups pages into sections by the first path segment: `/docs/...` goes to **Docs**, `/blog/...` to **Blog**, `/api-reference/...` to **Api Reference**. Root pages (`/`) go to **Pages**, and so does any page whose first path segment no other listed page shares (for example a lone `/pricing` or `/en`), so the file has no one-page sections.
- Leaves out a page's description when it is identical to the site summary, so a meta description repeated across the whole site is not repeated on every line.
- Assembles the `llms.txt` file per the llmstxt.org format and stores it in the key-value store record `llms.txt`.
- Writes one dataset row per page (included or not, with the reason), so you can see exactly what went into the file and why a page was left out.
- Respects `robots.txt` by default, removes duplicate URLs, skips binary and asset files (PDF, images, CSS, JS, fonts, video, XML, ZIP) and isolates failures: a page that cannot be fetched gets an error row and the run continues.

### Who it is for

- **Documentation and developer portals** that want AI assistants to point users to the right guide or API page.
- **Product and SaaS sites** that want to be described accurately by AI search and chat tools.
- **Agencies and SEO teams** preparing `llms.txt` files for several client sites without writing them by hand.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `startUrl` | string | - | Page where the crawl starts. Only same-origin links are followed. |
| `sitemapUrl` | string | - | XML sitemap or sitemap index. When set, pages come from its `<loc>` entries and no links are followed. |
| `maxPages` | integer | `100` | Pages processed per run, 1 to 500 (fetched pages plus pages skipped by robots.txt). |
| `maxDepth` | integer | `3` | Link depth from the start URL, 0 to 6. `0` fetches the start page only. Crawl mode only. |
| `includePatterns` | array of strings | - | When set, only URLs whose path (with query string) contains one of these substrings are kept. |
| `excludePatterns` | array of strings | - | URLs whose path (with query string) contains any of these substrings are skipped. Applied after the include patterns. |
| `siteName` | string | og:site\_name of the start page, else its hostname | H1 heading of the file. |
| `siteSummary` | string | meta description of the start page | Blockquote summary. Omitted when there is none. |
| `respectRobotsTxt` | boolean | `true` | Read `robots.txt` once per site and skip disallowed pages. |
| `maxConcurrency` | integer | `4` | Pages fetched in parallel, 1 to 10. |

At least one of `startUrl` or `sitemapUrl` is required. If both are given, the sitemap is used, a warning is logged, and the start URL only provides the site name and summary.

In crawl mode the start page is always fetched (its links and metadata are needed), even when it does not match the include patterns; in that case it is simply not listed in the file. In sitemap mode, if the site's home page is not in the sitemap, it is fetched once for the site name and summary only (no dataset row, no charge).

Example input:

```json
{
    "startUrl": "https://docs.example-app.com/",
    "maxPages": 200,
    "maxDepth": 3,
    "excludePatterns": ["/changelog/", "?page="],
    "respectRobotsTxt": true,
    "maxConcurrency": 4
}
```

### Output

#### The llms.txt file

Stored in the default key-value store under the key `llms.txt` (Markdown). Download it and upload it to the root of your site. Example:

```markdown
## Example App

> Example App is a hosted API for sending transactional email.

### Pages

- [Example App - Transactional email API](https://docs.example-app.com/)
- [Pricing](https://docs.example-app.com/pricing): Free tier, then pay per email sent.

### Guides

- [Quickstart](https://docs.example-app.com/guides/quickstart): Send your first email in five minutes.
- [Domains and DNS](https://docs.example-app.com/guides/domains): Verify a sending domain with SPF, DKIM and DMARC.
- [Webhooks](https://docs.example-app.com/guides/webhooks): Receive delivery, bounce and complaint events.

### Api Reference

- [Send an email](https://docs.example-app.com/api-reference/send): POST /v1/emails parameters and responses.
- [Errors](https://docs.example-app.com/api-reference/errors)
```

Sections are ordered with **Pages** first, then by number of pages (largest first). Inside a section, pages keep the order in which they were discovered. A page without a description, or whose description is identical to the site summary, gets a link line without the `: description` part.

#### Dataset (one row per page)

```json
{
    "url": "https://docs.example-app.com/guides/quickstart",
    "finalUrl": "https://docs.example-app.com/guides/quickstart",
    "httpStatus": 200,
    "title": "Quickstart",
    "description": "Send your first email in five minutes.",
    "section": "Guides",
    "included": true,
    "reason": null,
    "error": null
}
```

- `finalUrl` is the address after redirects; that is the URL written to the file.
- `section` and `description` are the page's own values. In the file, a page whose section would hold only that page is listed under **Pages**, and a description identical to the site summary is omitted.
- `included` tells whether the page is listed in `llms.txt`. When it is `false`, `reason` says why: `no title`, `not an HTML page`, `disallowed by robots.txt`, `filtered by include/exclude patterns`, `redirected to another site`, `duplicate of <url>` (two URLs that redirect to the same page) or `fetch failed`.
- `error` holds the error message when the page could not be fetched (for example `HTTP 404 for ...`), otherwise `null`.

Rows are written as each page finishes, so their order follows completion, not discovery.

#### Summary

The key-value store record `OUTPUT`:

```json
{
    "pagesCrawled": 137,
    "pagesIncluded": 128,
    "pagesSkipped": 9,
    "pagesFailed": 2,
    "chargeLimitReached": false,
    "sections": [
        { "name": "Pages", "pages": 4 },
        { "name": "Guides", "pages": 71 },
        { "name": "Api Reference", "pages": 53 }
    ],
    "llmsTxtChars": 14210
}
```

`pagesCrawled` counts pages requested over HTTP; `pagesSkipped` counts dataset rows not listed in the file (errors included).

### Combine with robots-llms-txt-monitor

Generate your `llms.txt` here, then monitor it with **robots-llms-txt-monitor** to make sure it stays live and valid after you publish it (and that your `robots.txt` does not accidentally block the crawlers you care about).

### Limitations

- **Plain HTML fetch only.** Pages are fetched without running JavaScript. Content that only appears after client-side rendering (single-page apps) can be missed, including links, titles and descriptions set by scripts.
- **No AI generation or summarization.** Titles and descriptions come verbatim from each page's own meta tags. Weak or missing meta tags give weak or missing descriptions; improve the meta tags or edit the file.
- **Sections follow URL structure.** The section is the first path segment, so on a site where every page sits under a locale prefix (`/en/...`) most pages land in one section named after it. Use include/exclude patterns or edit the file to adjust.
- **Include/exclude patterns apply to the discovered URL**, before redirects.
- **robots.txt is respected by default.** Disallowed pages are not fetched and get a `disallowed by robots.txt` row. You can disable this with `respectRobotsTxt: false`, which is only appropriate for sites you own or are allowed to crawl.
- **Max 500 pages per run.** Extra URLs are dropped with a warning.
- **Binary and asset URLs are skipped** by file extension, and responses that are not HTML are listed as `not an HTML page`. Response bodies over 5 MB are refused.
- **Plain XML sitemaps only.** Gzipped (`.xml.gz`) sitemaps are not supported; a sitemap index is followed one level deep (up to 50 child sitemaps).
- **Review the generated file before publishing it.** It is a starting point built from your pages' metadata: check the titles, remove pages that should not be recommended to AI tools and add context where it helps.

### Pricing

Pay-per-event: $0.002 USD per page fetched and parsed successfully ($2.00 per 1,000 pages). Pages that fail, are not HTML or are skipped because of robots.txt are not charged. Apify platform usage is included in this price; there is no separate platform-usage charge.

An Actor Start event costs $0.00005 USD. One start event is charged per GB of Actor memory, with a minimum of one event per run.

If the maximum charge you set for a run is reached, the actor stops starting new pages, finishes the ones in progress and still writes the `llms.txt` from the pages it has.

### Support

Cuantic Data - cuanticwindows@gmail.com

Terms of use: see [Terms of use](#terms-of-use) below.

### Terms of use

Provided by Cuantic Data (cuanticwindows@gmail.com).

#### 1. What the actor does

The actor fetches public pages of the website you point it to, either by following links from a start URL or by reading the XML sitemap you provide. It makes plain HTTP requests (no browser, no JavaScript), reads each page's title and description from its own HTML metadata, and stores the resulting rows and the generated `llms.txt` file in your Apify storage.

#### 2. Your responsibility

- You are responsible for having the right to crawl the site you submit. Only submit sites you own, manage, or are otherwise allowed to crawl.
- robots.txt is respected by default. If you disable that option, you confirm that you are allowed to fetch the disallowed pages.
- Each run generates real traffic on the target site; choose the number of pages and the concurrency accordingly.
- You are responsible for reviewing the generated `llms.txt` before publishing it, and for what you publish on your site.
- You must comply with the terms of service of the crawled site and with the Apify Terms of Service.

#### 3. About the results

- Titles and descriptions are copied from the pages' own metadata; the actor does not write, summarize or verify content. Pages that rely on JavaScript to render can be missed or listed with incomplete metadata.
- llms.txt is a proposed convention (https://llmstxt.org). Publishing the file does not guarantee that any AI system will read it or use it in a particular way.
- Results are provided "as is", for informational purposes, without warranty of accuracy, completeness or fitness for a particular purpose.

#### 4. Data

The actor stores only what it produces (the dataset rows, the `llms.txt` file and the summary record) in your own Apify storage. It does not keep copies of the crawled pages or results elsewhere.

#### 5. Liability

To the extent permitted by law, Cuantic Data is not liable for any damage arising from the use of the actor or its results, including the content of a published `llms.txt` file or any effect of the crawl traffic on the crawled site.

#### 6. Changes

These terms may be updated together with the actor. The version published with the actor is the one that applies.

# Actor input Schema

## `startUrl` (type: `string`):

Page where the crawl starts, usually the home page or the docs root. Only links on the same site (same origin) are followed. Provide this or a sitemap URL.

## `sitemapUrl` (type: `string`):

An XML sitemap (or sitemap index, followed one level deep). When set, pages come from its <loc> entries instead of following links, and the start URL is not crawled.

## `maxPages` (type: `integer`):

Maximum number of pages processed in one run (fetched pages plus pages skipped because robots.txt disallows them). Extra URLs are dropped with a warning.

## `maxDepth` (type: `integer`):

How many clicks away from the start URL the crawl may go. 0 fetches the start page only. Ignored in sitemap mode.

## `includePatterns` (type: `array`):

Optional. When set, only URLs whose path (including the query string) contains at least one of these substrings are kept, for example /docs/ or /blog/.

## `excludePatterns` (type: `array`):

Optional. URLs whose path (including the query string) contains any of these substrings are skipped, for example /tag/ or ?page=. Applied after the include patterns.

## `siteName` (type: `string`):

Optional. Title used for the llms.txt H1 heading. Default: og:site\_name of the start page, otherwise its hostname.

## `siteSummary` (type: `string`):

Optional. Short summary placed in the llms.txt blockquote. Default: the meta description of the start page; the blockquote is omitted when there is none.

## `respectRobotsTxt` (type: `boolean`):

When enabled, robots.txt is read once per site and disallowed pages are not fetched (they get a skipped row). Disable only for sites you own.

## `maxConcurrency` (type: `integer`):

Pages fetched in parallel. Plain HTTP requests are light; keep this low to be polite to the target site.

## Actor input object example

```json
{
  "startUrl": "https://example.com/",
  "maxPages": 100,
  "maxDepth": 3,
  "respectRobotsTxt": true,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `llmsTxt` (type: `string`):

The generated llms.txt, ready to upload.

## `pages` (type: `string`):

One row per page: title, description, section and whether it is included.

## `summary` (type: `string`):

Counts and sections of the generated file.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrl": "https://example.com/"
};

// Run the Actor and wait for it to finish
const run = await client.actor("cuantic_data/llms-txt-generator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrl": "https://example.com/" }

# Run the Actor and wait for it to finish
run = client.actor("cuantic_data/llms-txt-generator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrl": "https://example.com/"
}' |
apify call cuantic_data/llms-txt-generator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cuantic_data/llms-txt-generator"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ghjp7TNDjuhptslrB/builds/LR8mldjX2dTGEJ08G/openapi.json
