# llms.txt Generator (Website to llms.txt & llms-full.txt) (`webdatatools/llms-txt-generator`) Actor

Turn any website or docs site into an llms.txt (sectioned links with one-line descriptions) and llms-full.txt (every page as clean Markdown) for AI assistants, RAG and AI search. Plain HTTP, sitemap-aware, robots.txt-friendly.

- **URL**: https://apify.com/webdatatools/llms-txt-generator.md
- **Developed by:** [Murat Uzun](https://apify.com/webdatatools) (community)
- **Categories:** AI, Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.60 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What is llms.txt Generator?

**llms.txt Generator** turns any website or documentation site into the two files AI assistants and AI search engines read: **`llms.txt`** (the site's name, a one-line summary and a sectioned list of its pages with one-line descriptions, in the [llmstxt.org](https://llmstxt.org) format) and **`llms-full.txt`** (every page as clean Markdown in one file). Give it a start URL, and it reads the site's sitemap and links with plain HTTP, strips navigation and boilerplate, and writes both files plus one dataset row per page with its title, description, section and token count. **$1 per 1,000 pages.**

Use it to publish an llms.txt for your own site or product docs, to feed a competitor's or a library's documentation into ChatGPT, Claude or Cursor, to build a RAG knowledge base from a docs site, or to check how an AI will see your content.

### What does llms.txt Generator produce?

**Files** (in the run's key-value store, linked from the Output tab):

| File | Content |
|---|---|
| `llms.txt` | `# Site name`, `> summary`, then `## Section` headings with `- [Page title](url): description` lines; blog, changelog, legal and similar sections go under `## Optional` |
| `llms-full.txt` | The same header, then each page's full Markdown under its title and source URL |
| `SUMMARY` | Per site: pages, sections, token counts of both files, download links, and whether the site already publishes its own `/llms.txt` |

With several websites in one run, each file is prefixed with its domain (`docs.example.com-llms.txt`).

**Dataset** (one row per page, the billed unit):

| Field | Description |
|---|---|
| `site`, `url` | Website and page |
| `title`, `description` | Clean page title (site-name suffix removed) and one-line description |
| `section` | llms.txt section the page is listed under |
| `wordCount`, `tokens` | Size of the page content (tokens ≈ characters / 4) |
| `markdown` | The page as Markdown (turn off **Markdown in the dataset** to leave it out) |

### How to use llms.txt Generator

1. Enter one or more **Start URLs**, e.g. `https://crawlee.dev/js/docs/quick-start` or your home page.
2. Set **Max pages per website** (default 100).
3. Run, then open **Output** and download `llms.txt` and `llms-full.txt`.
4. To publish it on your site, upload `llms.txt` to your web root so it is served at `https://yoursite.com/llms.txt`.

A start URL inside a folder keeps the crawl in that folder: `https://example.com/docs/intro` reads only `/docs` pages, a home-page start URL reads the whole site. Set **Include path prefixes** to choose the folders yourself.

### Example input

```json
{
    "startUrls": [{ "url": "https://crawlee.dev/js/docs/quick-start" }],
    "maxPages": 200
}
```

### Example output (`llms.txt`)

```markdown
## Crawlee for JavaScript

> With this short tutorial you can start scraping with Crawlee in a minute or two. To learn more, read the Introduction.

### Docs

- [Quick Start](https://crawlee.dev/js/docs/quick-start): With this short tutorial you can start scraping with Crawlee in a minute or two. To learn more, read the Introduction.
- [Deployment guides](https://crawlee.dev/js/docs/3.10/deployment): Here you can find guides on how to deploy your crawlers to various cloud providers.
- [Add data to dataset](https://crawlee.dev/js/docs/3.10/examples/add-data-to-dataset): This example saves data to the default dataset. If the dataset doesn't exist, it will be created.
```

### How much does it cost?

Pay per result: **$0.001 per page read ($1 per 1,000)**, with volume discounts on paid Apify plans. Pages that fail to load are listed in the `ERRORS` record and cost nothing; the llms.txt files themselves are free. Set **Maximum cost per run** to cap spend.

### Limits

- Plain HTTP only: pages that render their text with JavaScript after loading come out short. Sites that serve their text in the HTML work (tested on the Docusaurus-based Crawlee and Apify docs).
- Descriptions come from each page's meta description, or its first paragraph when the meta description is missing or the same site-wide text.
- Tag, author, archive, search, login and cart pages are skipped because they add noise to llms.txt.
- It respects `robots.txt` by default.

### Using llms.txt Generator with AI agents

The Actor is pay-per-event with limited permissions, so AI agents can call it through Apify's MCP server (`mcp.apify.com`), for example with `{"startUrls": [{"url": "https://docs.example.com"}], "maxPages": 50}`, then read `llms-full.txt` from the run's key-value store. It also accepts `url`, `urls`, `website`, `websites`, and `inputDatasetId` + `inputField` to take URLs from another run.

### FAQ

**What is llms.txt?** A proposed standard ([llmstxt.org](https://llmstxt.org)) for a Markdown file at `/llms.txt` that tells language models what a site contains and where its key pages are, the way `robots.txt` and `sitemap.xml` do for crawlers.

**The site already has an llms.txt. Why generate one?** `SUMMARY` tells you when it does. A generated one is still useful to compare coverage, or when you need `llms-full.txt` and the site does not publish it.

### Related Actors

Part of the **webdatatools** web-intelligence suite — every Actor is pay-per-event, reads public data
without a login, and returns one clean row per entity:

Browse the whole suite at [webdatatools](https://paulet4a-commits.github.io/webdatatools/), or call ten of
these Actors straight from Claude, Cursor or Cline with the
[webdatatools MCP server](https://github.com/paulet4a-commits/webdatatools-mcp-server).

**Website & domain intelligence**

- [Email Extractor — Website Contact & Social Finder](https://apify.com/webdatatools/contact-extractor) — e-mails, phones and social profiles per domain
- [Tech Stack Detector — Wappalyzer & BuiltWith Alternative](https://apify.com/webdatatools/tech-stack-detector) — CMS, e-commerce, analytics, pixels and payments per domain
- [Domain DNS & Email Security Checker](https://apify.com/webdatatools/dns-email-security-checker) — SPF, DKIM, DMARC, MX provider, registrar and domain age
- [Domain Security Audit (TLS, HTTP headers, redirects, robots)](https://apify.com/webdatatools/domain-security-audit) — TLS expiry, security headers, redirect chain, robots and llms.txt
- [Subdomain Finder (Certificate Transparency)](https://apify.com/webdatatools/subdomain-finder) — every subdomain seen in CT logs, with a live DNS check
- [Bulk Core Web Vitals & PageSpeed Audit](https://apify.com/webdatatools/core-web-vitals-audit) — Lighthouse scores, LCP, CLS, INP and top fixes per URL
- [On-Page SEO Audit](https://apify.com/webdatatools/seo-page-audit) — title, meta, headings, links, images and schema issues per page
- [Sitemap URL Extractor & Change Monitor](https://apify.com/webdatatools/sitemap-extractor) — every sitemap URL, or new and removed pages between runs
- [Wayback Machine Snapshot & Page Change Tracker](https://apify.com/webdatatools/wayback-page-diff) — how a page changed over time, or every archived snapshot
- [Bulk Domain WHOIS & RDAP Lookup](https://apify.com/webdatatools/domain-whois-rdap) — registrar, dates, status and nameservers per domain
- [Web Scraper — CSS Selector & Data Extractor](https://apify.com/webdatatools/css-selector-extractor) — pull any CSS selector off any page, one row per URL
- [Website Screenshot Generator](https://apify.com/webdatatools/website-screenshot) — full-page or viewport PNG/JPEG screenshots of any URL

**Content for AI, LLMs and RAG**

- [AI Web Search & Read: Google results as clean Markdown](https://apify.com/webdatatools/ai-web-search) — a query turned into clean Markdown from the top search results
- [Website to Markdown — Content Crawler for LLM & RAG](https://apify.com/webdatatools/website-to-markdown) — any site as clean Markdown per page, no browser
- [PDF to Markdown Converter (Text, Headings, Metadata)](https://apify.com/webdatatools/pdf-to-markdown) — PDF files to clean Markdown with headings, metadata and token counts
- [RAG Text Chunker (Markdown to Chunks, Token Counts)](https://apify.com/webdatatools/rag-text-chunker) — heading-aware RAG chunks with exact GPT token counts from any dataset
- [Article & News Extractor (clean text, author, date, markdown)](https://apify.com/webdatatools/article-extractor) — clean article text, author, date and Markdown per URL
- [Structured Data & JSON-LD Extractor (Schema.org, Open Graph)](https://apify.com/webdatatools/structured-data-extractor) — Schema.org and Open Graph data from any page
- [Google News Scraper (RSS search by keyword, topic, site)](https://apify.com/webdatatools/google-news-scraper) — news results by keyword, topic or site
- [Press Release Monitor: PR Newswire, BusinessWire, GlobeNewswire](https://apify.com/webdatatools/press-release-monitor) — PR Newswire, Business Wire and GlobeNewswire releases

**Search, video and social**

- [YouTube Shorts Scraper](https://apify.com/webdatatools/youtube-shorts-scraper) — Shorts from channels, hashtags and searches with view counts
- [Pinterest Pins Scraper](https://apify.com/webdatatools/pinterest-pins-scraper) — latest pins of public Pinterest profiles and boards
- [YouTube Transcript Scraper](https://apify.com/webdatatools/youtube-transcript-scraper) — captions and subtitles as text + timed segments, per video or channel
- [Google Search Results Scraper — SERP API](https://apify.com/webdatatools/google-search-scraper) — organic SERP results per keyword and country
- [YouTube Comments Scraper — Comments & Replies](https://apify.com/webdatatools/youtube-comments-scraper) — comments and replies with likes, no API key
- [YouTube Channel Latest Videos (RSS, no API key)](https://apify.com/webdatatools/youtube-channel-videos) — the latest 15 videos of any channel from RSS
- [YouTube Channel Scraper (videos, shorts, live)](https://apify.com/webdatatools/youtube-channel-scraper) — a channel's full video, shorts and stream list
- [YouTube Search Results Scraper (videos, channels, no API key)](https://apify.com/webdatatools/youtube-search-scraper) — videos, channels and playlists per query
- [YouTube Video Details Scraper (views, likes, description, tags)](https://apify.com/webdatatools/youtube-video-details) — views, likes, description, tags and chapters per video
- [Apple Podcasts Lookup & Episodes Scraper](https://apify.com/webdatatools/podcast-lookup) — podcast metadata and episodes from iTunes and RSS
- [Bluesky Post, Search & Profile Scraper](https://apify.com/webdatatools/bluesky-scraper) — posts, profiles, followers and threads from the AT Protocol API
- [X Tweet Scraper (Twitter Posts by URL, No Login)](https://apify.com/webdatatools/x-tweet-scraper) — X (Twitter) posts by URL with likes, replies, author, media and MP4 links
- [Telegram Channel Posts Scraper](https://apify.com/webdatatools/telegram-channel-scraper) — posts, views and media flags from any public channel
- [Substack Publication & Posts Scraper](https://apify.com/webdatatools/substack-scraper) — archive, authors and paywall status per publication
- [Google Play Reviews Scraper](https://apify.com/webdatatools/google-play-reviews-scraper) — reviews, ratings, replies and app versions per app
- [App Store Reviews Scraper](https://apify.com/webdatatools/app-store-reviews-scraper) — iOS reviews and ratings per app and country
- [Trustpilot Reviews Scraper (Ratings, Replies, No Login)](https://apify.com/webdatatools/trustpilot-reviews-scraper) — Trustpilot reviews with stars, full text, verification and company replies
- [Google Trends Scraper](https://apify.com/webdatatools/google-trends-scraper) — interest over time, by region, and related queries per keyword
- [Google Ads Transparency Scraper](https://apify.com/webdatatools/google-ads-transparency-scraper) — ads any advertiser runs on Google, with format and dates
- [Keyword Suggestions Scraper (Google, YouTube, Amazon, Bing)](https://apify.com/webdatatools/keyword-suggestions-scraper) — autocomplete keyword ideas from four search engines
- [Bilibili Scraper (Videos, Search, Popular)](https://apify.com/webdatatools/bilibili-scraper) — Chinese video platform: views, likes, coins, danmaku, uploader
- [Mastodon Scraper (Hashtags, Accounts, Trending)](https://apify.com/webdatatools/mastodon-scraper) — public fediverse posts by hashtag, account or trending
- [Meetup Events Scraper (Search by Keyword & City)](https://apify.com/webdatatools/meetup-events-scraper) — upcoming events with RSVPs, fees, venues and groups
- [Eventbrite Scraper (Events by Keyword & City)](https://apify.com/webdatatools/eventbrite-scraper) — events by keyword and city with venue, dates and organizer

**Leads, jobs and company data**

- [Career Site Jobs API (Greenhouse, Lever, Ashby, Workday +1)](https://apify.com/webdatatools/career-site-jobs-api) — company domains in, their open jobs out, ATS detected automatically
- [Workday Jobs Scraper](https://apify.com/webdatatools/workday-jobs-scraper) — jobs with full descriptions from any Workday career site
- [Google Maps Scraper](https://apify.com/webdatatools/google-maps-scraper) — businesses with phone, website, address, rating and coordinates per search
- [LinkedIn Jobs Scraper](https://apify.com/webdatatools/linkedin-jobs-scraper) — job titles, companies, locations and full descriptions from LinkedIn job search
- [Company 360: full company profile from a domain](https://apify.com/webdatatools/company-360) — one row per domain: contacts, tech, security, hiring and company facts
- [Hiring Signals Scraper (Greenhouse, Lever, Ashby, Workable)](https://apify.com/webdatatools/hiring-signals) — open jobs and hiring velocity from 10 public ATS boards
- [Y Combinator Companies & Founders Scraper](https://apify.com/webdatatools/yc-companies-scraper) — YC startups by batch, industry and hiring status
- [Wikidata Entity & Company Enrichment (facts, IDs, links)](https://apify.com/webdatatools/wikidata-entity-enrichment) — HQ, founders, employees, revenue and social IDs per company
- [Email Validator & Verifier — Bulk Email Check](https://apify.com/webdatatools/email-validator) — syntax, MX, disposable, role and free-provider checks
- [OpenStreetMap POI Extractor (Overpass API: shops, amenities)](https://apify.com/webdatatools/overpass-poi-extractor) — shops and amenities by radius, bbox or area
- [Stock, Crypto & FX Quotes](https://apify.com/webdatatools/market-quotes) — one row per symbol from Yahoo, Binance and ECB rates
- [Remote Jobs Aggregator (RemoteOK, WWR, Hacker News)](https://apify.com/webdatatools/remote-jobs-aggregator) — one clean row per remote job, de-duplicated across feeds
- [Greenhouse Jobs Scraper](https://apify.com/webdatatools/greenhouse-jobs-scraper) — jobs with descriptions from any Greenhouse job board
- [Lever Jobs Scraper](https://apify.com/webdatatools/lever-jobs-scraper) — jobs with descriptions from any Lever careers page
- [Ashby Jobs Scraper](https://apify.com/webdatatools/ashby-jobs-scraper) — jobs, salaries and descriptions from any Ashby job board
- [SmartRecruiters Jobs Scraper](https://apify.com/webdatatools/smartrecruiters-jobs-scraper) — jobs with descriptions from any SmartRecruiters company
- [Seek Jobs Scraper (Australia & New Zealand)](https://apify.com/webdatatools/seek-jobs-scraper) — Seek job ads with salary, work type and location
- [Dice Jobs Scraper](https://apify.com/webdatatools/dice-jobs-scraper) — US tech jobs from Dice with salary and remote flag
- [AutoScout24 Scraper](https://apify.com/webdatatools/autoscout24-scraper) — European car listings with price, mileage and seller
- [Rightmove Scraper](https://apify.com/webdatatools/rightmove-scraper) — UK property for sale or rent with price and agent
- [Wellfound Jobs Scraper (AngelList Startup Jobs)](https://apify.com/webdatatools/wellfound-jobs-scraper) — startup jobs with salary and equity ranges, company size and stage
- [Yandex Maps Scraper (Places, Ratings, Phones)](https://apify.com/webdatatools/yandex-maps-scraper) — businesses in Russia, Türkiye and the CIS with phones, ratings, hours
- [Craigslist Scraper (Listings, Prices, Locations)](https://apify.com/webdatatools/craigslist-scraper) — listings in any area and category with price, date and coordinates
- [JobStreet Scraper (Malaysia, Singapore, PH, ID + JobsDB)](https://apify.com/webdatatools/jobstreet-scraper) — JobStreet and JobsDB jobs in 6 Asian countries with parsed salaries
- [InfoJobs Scraper (Spain Jobs, Salaries, Companies)](https://apify.com/webdatatools/infojobs-scraper) — Spanish jobs with salary range, contract type and full description
- [Redfin Scraper (Homes for Sale, Prices, Details)](https://apify.com/webdatatools/redfin-scraper) — US homes for sale or sold from any Redfin search, with price and details
- [Kleinanzeigen Scraper (Ads, Prices, Locations)](https://apify.com/webdatatools/kleinanzeigen-scraper) — German classifieds with price, VB flag, ZIP, city and seller type

**Developer, app and research data**

- [npm, PyPI & Crates.io Package Health Checker](https://apify.com/webdatatools/package-health-checker) — releases, downloads, deprecation and a health score
- [GitHub Repository Health & Activity Report](https://apify.com/webdatatools/github-repo-health) — stars, commits, contributors and risk flags per repo
- [VS Code Marketplace Extension Scraper (installs, ratings)](https://apify.com/webdatatools/vscode-marketplace-extensions) — installs, ratings and versions per extension
- [Chrome Web Store Extension Scraper (installs, ratings)](https://apify.com/webdatatools/chrome-web-store-extensions) — users, rating, version and developer per extension
- [Google Play Scraper](https://apify.com/webdatatools/google-play-scraper) — apps, ratings, installs, developer contact and reviews
- [App Store (iOS) App Metadata, Ratings & Top Charts Lookup](https://apify.com/webdatatools/app-store-lookup) — ratings, price, version and charts per app
- [CrossRef DOI & Citation Metadata Lookup](https://apify.com/webdatatools/crossref-doi-lookup) — papers, authors, journals and citation counts
- [FDA Recalls & Adverse Events Monitor (openFDA)](https://apify.com/webdatatools/openfda-recall-monitor) — food, drug and device recalls from openFDA
- [iCal / ICS Calendar Feed to Events Extractor](https://apify.com/webdatatools/ical-calendar-extractor) — any public calendar feed as event rows
- [Shopify Store Products Scraper](https://apify.com/webdatatools/shopify-products-scraper) — catalog, prices, variants and stock per store
- [Hacker News Search & Front Page Scraper](https://apify.com/webdatatools/hacker-news-scraper) — stories, comments and points by query or front page
- [GitHub Trending Repositories Scraper](https://apify.com/webdatatools/github-trending-scraper) — trending repos and developers by language and period
- [Stack Overflow & Stack Exchange Q\&A Scraper](https://apify.com/webdatatools/stackexchange-scraper) — questions, answers and scores by query, tag or site
- [Bulk Image Downloader](https://apify.com/webdatatools/bulk-image-downloader) — download image URLs to storage with size, dimensions and a ZIP
- [Google Flights Scraper (Prices, Airlines, Stops)](https://apify.com/webdatatools/google-flights-scraper) — flight prices, airlines, times, stops and CO2 by route and date
- [Google Hotels Scraper (Prices, Ratings, Reviews)](https://apify.com/webdatatools/google-hotels-scraper) — hotel prices per night, stars, rating and reviews by city and dates
- [Booking.com Scraper (Hotel Prices, Ratings, Availability)](https://apify.com/webdatatools/booking-scraper) — Booking.com hotels for any city and dates with prices, scores and deals
- [AliExpress Scraper (Search Products & Prices)](https://apify.com/webdatatools/aliexpress-scraper) — AliExpress search results with USD price, discount and rank
- [Lazada Scraper (Products, Prices, Sold, Ratings)](https://apify.com/webdatatools/lazada-scraper) — Lazada products in 6 countries with price, rating, units sold and seller
- [Trendyol Scraper (Turkey Products, Prices, Ratings)](https://apify.com/webdatatools/trendyol-scraper) — Trendyol products with price, basket discount, rating and promotions

# Actor input Schema

## `startUrls` (type: `array`):

Enter the website(s) to turn into llms.txt, e.g. https://docs.apify.com. Each different domain gets its own llms.txt and llms-full.txt.

## `maxPages` (type: `integer`):

Enter the maximum number of pages to read per website, e.g. 100. This is the billed unit (one dataset row per page).

## `maxDepth` (type: `integer`):

Enter how many links deep to follow from a start URL, e.g. 3. Set 0 to crawl only the start URL(s).

## `includePathPrefixes` (type: `array`):

Optional. Only read pages whose path starts with one of these prefixes, e.g. /docs. Leave empty to stay inside the folder of each start URL (https://example.com/docs/intro → /docs), or the whole site for a home-page start URL.

## `excludePathPatterns` (type: `array`):

Optional. Regular expressions tested against the full URL; a match is skipped, e.g. .(png|jpe?g)$ to skip images. Defaults cover binary files and login/signup pages.

## `useSitemap` (type: `boolean`):

Turn this on to also read /sitemap.xml on each start URL's domain and add its URLs (filtered by the settings above) to the crawl queue, up to Max pages.

## `includeMarkdown` (type: `boolean`):

Turn this off to keep only title, description and token count per page in the dataset (llms-full.txt still has every page's Markdown).

## `removeSelectors` (type: `array`):

Optional. Extra CSS selectors to strip before extracting content, e.g. .cookie-banner or #newsletter-signup, on top of the built-in nav/header/footer/aside removal.

## `maxConcurrency` (type: `integer`):

Enter the maximum number of pages fetched in parallel, e.g. 10. Lower it if the target site rate-limits you.

## `respectRobots` (type: `boolean`):

Turn this on to skip URLs disallowed by the site's robots.txt file (recommended and on by default).

## `proxyConfiguration` (type: `object`):

Optional. Select a proxy configuration, e.g. Apify Proxy with the datacenter group, if the target site blocks requests from shared IPs.

## `inputDatasetId` (type: `string`):

Take the URLs from another Actor run's dataset, e.g. "aBcD1234". Combined with the list above.

## `inputField` (type: `string`):

Column holding the value, e.g. "url" (the default) or "website". Array columns are flattened.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://crawlee.dev/js/docs/quick-start"
    }
  ],
  "maxPages": 50,
  "maxDepth": 3,
  "includePathPrefixes": [],
  "excludePathPatterns": [
    "\\.(png|jpe?g|gif|svg|webp|pdf|zip|mp4|css|js)$",
    "/login",
    "/signup",
    "\\?replytocom="
  ],
  "useSitemap": true,
  "includeMarkdown": true,
  "removeSelectors": [],
  "maxConcurrency": 10,
  "respectRobots": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `llmsTxt` (type: `string`):

The llms.txt file (single-site runs; multi-site runs prefix each file with the domain, see Files).

## `llmsFullTxt` (type: `string`):

Every page as Markdown in one file.

## `summary` (type: `string`):

Per-site summary: pages, sections, token counts, file links.

## `files` (type: `string`):

All generated files.

## `pages` (type: `string`):

All pages read in this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://crawlee.dev/js/docs/quick-start"
        }
    ],
    "maxPages": 50,
    "maxDepth": 3,
    "includePathPrefixes": [],
    "excludePathPatterns": [
        "\\.(png|jpe?g|gif|svg|webp|pdf|zip|mp4|css|js)$",
        "/login",
        "/signup",
        "\\?replytocom="
    ],
    "removeSelectors": [],
    "maxConcurrency": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("webdatatools/llms-txt-generator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://crawlee.dev/js/docs/quick-start" }],
    "maxPages": 50,
    "maxDepth": 3,
    "includePathPrefixes": [],
    "excludePathPatterns": [
        "\\.(png|jpe?g|gif|svg|webp|pdf|zip|mp4|css|js)$",
        "/login",
        "/signup",
        "\\?replytocom=",
    ],
    "removeSelectors": [],
    "maxConcurrency": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("webdatatools/llms-txt-generator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://crawlee.dev/js/docs/quick-start"
    }
  ],
  "maxPages": 50,
  "maxDepth": 3,
  "includePathPrefixes": [],
  "excludePathPatterns": [
    "\\\\.(png|jpe?g|gif|svg|webp|pdf|zip|mp4|css|js)$",
    "/login",
    "/signup",
    "\\\\?replytocom="
  ],
  "removeSelectors": [],
  "maxConcurrency": 10
}' |
apify call webdatatools/llms-txt-generator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,webdatatools/llms-txt-generator"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/elRun5MdeavvM2AbA/builds/XNVzCLuBjpdoyr8EP/openapi.json
