llms.txt Generator (Website to llms.txt & llms-full.txt) avatar

llms.txt Generator (Website to llms.txt & llms-full.txt)

Pricing

from $0.60 / 1,000 pages

Go to Apify Store
llms.txt Generator (Website to llms.txt & llms-full.txt)

llms.txt Generator (Website to llms.txt & llms-full.txt)

Turn any website or docs site into an llms.txt (sectioned links with one-line descriptions) and llms-full.txt (every page as clean Markdown) for AI assistants, RAG and AI search. Plain HTTP, sitemap-aware, robots.txt-friendly.

Pricing

from $0.60 / 1,000 pages

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What is llms.txt Generator?

llms.txt Generator turns any website or documentation site into the two files AI assistants and AI search engines read: llms.txt (the site's name, a one-line summary and a sectioned list of its pages with one-line descriptions, in the llmstxt.org format) and llms-full.txt (every page as clean Markdown in one file). Give it a start URL, and it reads the site's sitemap and links with plain HTTP, strips navigation and boilerplate, and writes both files plus one dataset row per page with its title, description, section and token count. $1 per 1,000 pages.

Use it to publish an llms.txt for your own site or product docs, to feed a competitor's or a library's documentation into ChatGPT, Claude or Cursor, to build a RAG knowledge base from a docs site, or to check how an AI will see your content.

What does llms.txt Generator produce?

Files (in the run's key-value store, linked from the Output tab):

FileContent
llms.txt# Site name, > summary, then ## Section headings with - [Page title](url): description lines; blog, changelog, legal and similar sections go under ## Optional
llms-full.txtThe same header, then each page's full Markdown under its title and source URL
SUMMARYPer site: pages, sections, token counts of both files, download links, and whether the site already publishes its own /llms.txt

With several websites in one run, each file is prefixed with its domain (docs.example.com-llms.txt).

Dataset (one row per page, the billed unit):

FieldDescription
site, urlWebsite and page
title, descriptionClean page title (site-name suffix removed) and one-line description
sectionllms.txt section the page is listed under
wordCount, tokensSize of the page content (tokens ≈ characters / 4)
markdownThe page as Markdown (turn off Markdown in the dataset to leave it out)

How to use llms.txt Generator

  1. Enter one or more Start URLs, e.g. https://crawlee.dev/js/docs/quick-start or your home page.
  2. Set Max pages per website (default 100).
  3. Run, then open Output and download llms.txt and llms-full.txt.
  4. To publish it on your site, upload llms.txt to your web root so it is served at https://yoursite.com/llms.txt.

A start URL inside a folder keeps the crawl in that folder: https://example.com/docs/intro reads only /docs pages, a home-page start URL reads the whole site. Set Include path prefixes to choose the folders yourself.

Example input

{
"startUrls": [{ "url": "https://crawlee.dev/js/docs/quick-start" }],
"maxPages": 200
}

Example output (llms.txt)

# Crawlee for JavaScript
> With this short tutorial you can start scraping with Crawlee in a minute or two. To learn more, read the Introduction.
## Docs
- [Quick Start](https://crawlee.dev/js/docs/quick-start): With this short tutorial you can start scraping with Crawlee in a minute or two. To learn more, read the Introduction.
- [Deployment guides](https://crawlee.dev/js/docs/3.10/deployment): Here you can find guides on how to deploy your crawlers to various cloud providers.
- [Add data to dataset](https://crawlee.dev/js/docs/3.10/examples/add-data-to-dataset): This example saves data to the default dataset. If the dataset doesn't exist, it will be created.

How much does it cost?

Pay per result: $0.001 per page read ($1 per 1,000), with volume discounts on paid Apify plans. Pages that fail to load are listed in the ERRORS record and cost nothing; the llms.txt files themselves are free. Set Maximum cost per run to cap spend.

Limits

  • Plain HTTP only: pages that render their text with JavaScript after loading come out short. Sites that serve their text in the HTML work (tested on the Docusaurus-based Crawlee and Apify docs).
  • Descriptions come from each page's meta description, or its first paragraph when the meta description is missing or the same site-wide text.
  • Tag, author, archive, search, login and cart pages are skipped because they add noise to llms.txt.
  • It respects robots.txt by default.

Using llms.txt Generator with AI agents

The Actor is pay-per-event with limited permissions, so AI agents can call it through Apify's MCP server (mcp.apify.com), for example with {"startUrls": [{"url": "https://docs.example.com"}], "maxPages": 50}, then read llms-full.txt from the run's key-value store. It also accepts url, urls, website, websites, and inputDatasetId + inputField to take URLs from another run.

FAQ

What is llms.txt? A proposed standard (llmstxt.org) for a Markdown file at /llms.txt that tells language models what a site contains and where its key pages are, the way robots.txt and sitemap.xml do for crawlers.

The site already has an llms.txt. Why generate one? SUMMARY tells you when it does. A generated one is still useful to compare coverage, or when you need llms-full.txt and the site does not publish it.

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data