# llms.txt Generator – Make Your Website AI-Ready (`webintel/llms-txt-generator`) Actor

Generate a spec-compliant llms.txt (llmstxt.org) and optional llms-full.txt for any website. Crawls the sitemap or internal links, groups pages into sections, adds an Optional section and clean Markdown full text. Ready to upload.

- **URL**: https://apify.com/webintel/llms-txt-generator.md
- **Developed by:** [Deepak Ganesh](https://apify.com/webintel) (community)
- **Categories:** AI, SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## llms.txt Generator – Make Your Website AI-Ready

![llms.txt Generator – Make Your Website AI-Ready](https://api.apify.com/v2/key-value-stores/21SBwDwdNtlnapIdO/records/llms-txt-generator.png?v=963848cb)

Generate a **spec-compliant [`llms.txt`](https://llmstxt.org)** file, plus an optional **`llms-full.txt`**, for any website. Upload the files to your site root so that ChatGPT, Claude, Perplexity, Gemini, Cursor and other AI assistants can understand your site.

- 📄 **Follows the llmstxt.org spec**: `# Title`, `> summary` blockquote, `## Section` headings with `- [Title](url): description` lists, and an `## Optional` section for secondary pages. Brackets in titles are escaped, and every file is checked against the spec before it is saved.
- 🗂️ **Smart sections**: pages are grouped by URL path (`/docs` → **Docs**, `/blog` → **Blog**, large areas are split further, e.g. **JS / API**), or put into one flat list.
- 🧭 **Covers the whole site**: pages come from `sitemap.xml` (including sitemaps listed in robots.txt and sitemap indexes) or from following internal links. Pages are picked evenly from each part of the site, so a small `maxPages` still gives a full overview. Versioned doc copies and alternate-language duplicates are pushed to the back.
- 📚 **llms-full.txt**: the clean Markdown content of every page in one file. Navigation, headers, footers, sidebars, scripts, forms and cookie banners are removed.
- ✨ **Clean titles and descriptions**: uses og:title, then h1, then `<title>`, with " | Brand" suffixes removed. Descriptions come from the meta description or the first real paragraph (max 200 characters). Generic titles or descriptions repeated on many pages are replaced with page-specific text.
- 🤝 **Polite crawling**: plain HTTP (no browser), at most 5 parallel requests, follows robots.txt, skips `noindex` pages, never logs in.
- 💸 **You only pay for pages that end up in the file.** Websites that fail to load are free.

### Use cases

- **Website owners and SEO / GEO teams**: publish `/llms.txt` and `/llms-full.txt` so AI search and assistants describe your product correctly.
- **Docs teams**: give coding assistants (Cursor, Copilot, Claude Code) an index of your documentation, or the whole docs as one Markdown file.
- **Agencies**: create llms.txt files for many client sites in one run.
- **AI and RAG developers**: turn any site into clean Markdown context for LLMs.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `startUrl` | string | – | Website to process, e.g. `https://example.com` or `example.com`. A URL with a path (e.g. `https://example.com/docs`) limits the file to that part of the site. |
| `startUrls` | string list | – | Optional: several websites. One file set is generated per website. |
| `maxPages` | integer | `50` | Max pages per website (1–500). Each included page is one billed **Page** event. |
| `includeFullText` | boolean | `false` | Also generate `llms-full.txt` with the Markdown content of every page. |
| `sectionStrategy` | `path` / `flat` | `path` | Group links by URL path, or put them in a single list. |
| `useSitemap` | boolean | `true` | Find pages via sitemap.xml. Falls back to following internal links when there is no sitemap. |
| `excludePatterns` | string list | login, signup, account, cart, checkout, tag, category, feed, search, privacy, terms, cookie, legal, author and pagination pages | URL path globs (`**/login*`, `/blog/tag/**`). Matching pages go to `## Optional` and are crawled last. |
| `dropExcludedPages` | boolean | `false` | Skip matching pages completely instead of listing them under `## Optional`. |
| `proxyConfiguration` | object | no proxy | Use Apify Proxy only if a site blocks you. |

Example:

```json
{
    "startUrl": "https://llmstxt.org",
    "maxPages": 50,
    "includeFullText": true
}
```

### Output

**Files.** Each website gets these files in the run's key-value store:

- `llms-<domain>.txt`: the llms.txt file (`text/plain; charset=utf-8`). Rename it to `llms.txt` and upload it to your site root.
- `llms-full-<domain>.txt`: the full-text version, if `includeFullText` is enabled.

The **Output** tab links to the dataset and to the list of files. Each dataset item also contains direct download URLs.

**Dataset.** One item per website. This example is trimmed from a real run:

```json
{
    "site": "https://llmstxt.org",
    "success": true,
    "error": null,
    "siteTitle": "llms-txt",
    "summary": "A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.",
    "pagesCrawled": 8,
    "sectionCount": 1,
    "sections": [
        {
            "name": "Pages",
            "linkCount": 8,
            "links": [
                { "title": "The /llms.txt file, v2", "url": "https://llmstxt.org/", "description": "A proposal to standardise on using an /llms.txt file to provide information to help agents use a website." },
                { "title": "Python source", "url": "https://llmstxt.org/core.html", "description": "Source code for llms_txt Python module, containing helpers to create and use llms.txt files" }
            ]
        }
    ],
    "llmsTxt": "# llms-txt\n\n> A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.\n\n## Pages\n\n- [The /llms.txt file, v2](https://llmstxt.org/): A proposal ...\n- [Python source](https://llmstxt.org/core.html): Source code for llms_txt Python module, ...\n",
    "llmsTxtKey": "llms-llmstxt.org.txt",
    "llmsTxtUrl": "https://api.apify.com/v2/key-value-stores/MuEO6MqBjJWvCYjhX/records/llms-llmstxt.org.txt?signature=…",
    "llmsFullTxtKey": "llms-full-llmstxt.org.txt",
    "llmsFullTxtUrl": "https://api.apify.com/v2/key-value-stores/MuEO6MqBjJWvCYjhX/records/llms-full-llmstxt.org.txt?signature=…",
    "warnings": [],
    "generatedAt": "2026-10-07T16:36:22.683Z"
}
```

A website that cannot be loaded returns `success: false` with an `error`, and you are not charged for it:

```json
{ "site": "https://this-domain-does-not-exist-xyz123.com", "success": false, "error": "RequestError: getaddrinfo ENOTFOUND this-domain-does-not-exist-xyz123.com", "pagesCrawled": 0, "llmsTxt": null }
```

The generated `llms.txt` for a larger site (crawlee.dev, trimmed) looks like this:

```markdown
## Crawlee

> Crawlee helps you build and maintain your crawlers. It's open source, but built by developers who scrape millions of pages every day for a living.

### Pages

- [Build reliable crawlers. Fast.](https://crawlee.dev/): Crawlee helps you build and maintain your crawlers. ...

### Blog

- [Crawlee v3.18: Type-safe routers](https://crawlee.dev/blog/crawlee-v3-18): Crawlee v3.18 brings type-safe router labels, ...

### Python / Docs

- [Quick start](https://crawlee.dev/python/docs/quick-start): This short tutorial will help you start scraping with Crawlee in just a minute or two. ...

### JS

- [Introduction](https://crawlee.dev/js/docs/introduction): Your first steps into the world of scraping with Crawlee
```

The `warnings` array explains anything that was left out, for example: maxPages reached, pages blocked by robots.txt, `noindex` pages, alternate-language URLs, pages that failed to load, or a missing meta description.

Dataset views: **Overview**, **llms.txt content** and **Links** (one row per link).

### Pricing

Pay per event: you pay only for pages included in a generated file.

| Event | Price | When it is charged |
|---|---|---|
| Page | $0.001 ($1 per 1,000 pages) | Each page included in a generated llms.txt |
| Actor start | $0.00005 | Once per run (per GB of memory) |

- A 50-page site costs about **$0.05**.
- **Failed websites are free.** So are invalid URLs, pages that return errors, and duplicate, `noindex` or robots-blocked pages.
- `llms-full.txt` costs nothing extra.
- If you set a maximum cost per run, the Actor stops when it is reached and still saves the files built so far.

### FAQ

**Where do I put the file?** Rename `llms-<domain>.txt` to `llms.txt` and upload it to your website root (`https://example.com/llms.txt`). Do the same with `llms-full.txt`. On WordPress, Webflow, Shopify and similar, use a file manager, a redirect, or a plugin that serves static files.

**Is the output spec-compliant?** Yes. There is exactly one H1, an optional blockquote summary, only H2 section headings, and link lists in the form `- [name](url): notes`. Secondary pages go under `## Optional`. Every file is checked by a validator before it is saved, and any problem is listed in `warnings`.

**Why are some pages missing?** `maxPages` limits how many pages are included. Pages are chosen evenly from each part of the site. Raise `maxPages` (up to 500), or start from a path like `https://example.com/docs` to focus on one part.

**Does it work on JavaScript-heavy sites?** It reads the server-rendered HTML, which works for most sites, including docs frameworks, WordPress, Webflow and Next.js. Pure client-side apps that render nothing without JavaScript will give sparse descriptions.

**Does it respect robots.txt?** Yes. URLs disallowed for all user agents (`*`) are skipped, and so are pages marked `noindex`.

**Can I process several websites?** Yes. Add them to `startUrls`. Each website gets its own files and dataset item.

This Actor is not affiliated with llmstxt.org, Answer.AI, OpenAI, Anthropic, Google or Perplexity.

### Changelog

- **0.1** (2026-10): First release. llms.txt and llms-full.txt generation, sitemap and link discovery, path-based sections, Optional section, spec validation, pay per page.

# Actor input Schema

## `startUrl` (type: `string`):

The website to generate llms.txt for, e.g. "https://example.com" or just "example.com". A URL with a path (e.g. "https://example.com/docs") limits the file to pages under that path.

## `startUrls` (type: `array`):

Optional: several websites at once, one per line. One llms.txt (and llms-full.txt) is generated per website. Duplicates are removed.

## `maxPages` (type: `integer`):

Maximum number of pages to include in each llms.txt. Each included page is billed as one "Page" event.

## `includeFullText` (type: `boolean`):

Also create llms-full.txt: the clean Markdown content of every included page in one file (navigation, headers, footers, scripts and forms removed).

## `sectionStrategy` (type: `string`):

"path" groups links into ## sections by the first URL path segment (/docs → Docs, /blog → Blog). "flat" puts all links in one section.

## `useSitemap` (type: `boolean`):

Find pages via sitemap.xml (including sitemaps listed in robots.txt). If disabled or no sitemap exists, pages are discovered by following internal links from the homepage.

## `excludePatterns` (type: `array`):

URL path globs (e.g. "**/login\*", "/blog/tag/**") for secondary pages. Matching pages are moved to the "## Optional" section (and crawled last) instead of the main sections. Leave empty to use the defaults: login, signup, account, cart, checkout, tag, category, feed, search, privacy, terms, cookie, legal, author and pagination pages.

## `dropExcludedPages` (type: `boolean`):

If enabled, pages matching the patterns above are skipped completely (not crawled, not billed) instead of being listed under "## Optional".

## `proxyConfiguration` (type: `object`):

Optional. Most websites work without a proxy. Use Apify Proxy only if the site blocks you.

## Actor input object example

```json
{
  "startUrl": "https://llmstxt.org",
  "maxPages": 50,
  "includeFullText": false,
  "sectionStrategy": "path",
  "useSitemap": true,
  "dropExcludedPages": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One item per website: site, success, siteTitle, summary, pagesCrawled, sections\[] {name, links\[]}, llmsTxt, llmsTxtUrl, llmsFullTxtUrl, warnings\[], error.

## `links` (type: `string`):

One row per link in the generated llms.txt files (section, title, URL, description).

## `files` (type: `string`):

List of the generated files in the key-value store (llms-<domain>.txt and llms-full-<domain>.txt). Download a file at .../records/<key>.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrl": "https://llmstxt.org",
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("webintel/llms-txt-generator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrl": "https://llmstxt.org",
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("webintel/llms-txt-generator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrl": "https://llmstxt.org",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call webintel/llms-txt-generator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,webintel/llms-txt-generator"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4Qpjgjwjxb1ksmoSw/builds/DqZTLg16OWAgdCILh/openapi.json
