AI Crawler Access — robots.txt, llms.txt and sitemap per domain avatar

AI Crawler Access — robots.txt, llms.txt and sitemap per domain

Pricing

$1.00 / 1,000 row returneds

Go to Apify Store
AI Crawler Access — robots.txt, llms.txt and sitemap per domain

AI Crawler Access — robots.txt, llms.txt and sitemap per domain

For every domain: whether AI crawlers are allowed, whether llms.txt exists, and whether the sitemap is reachable. No browser needed. Blocked or unreachable files are reported as unknown, never as absent.

Pricing

$1.00 / 1,000 row returneds

Rating

0.0

(0)

Developer

The Artifact Machine

The Artifact Machine

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

For every domain: whether AI crawlers are allowed, whether llms.txt exists, and whether the sitemap is reachable. No browser needed. Blocked or unreachable files are reported as unknown, never as absent.

See which sites let AI crawlers in, and which quietly block them.

What you get

One row per domain, with these columns:

  • hasRobotsTxt
  • hasLlmsTxt
  • hasSitemap
  • blocksGptBot
  • blocksClaudeBot
  • blocksCCBot
  • robotsRetrieval

Real output

Measured on https://www.hubspot.com/:

{
"hasRobotsTxt": true,
"hasLlmsTxt": true,
"hasSitemap": true,
"blocksGptBot": false,
"blocksClaudeBot": false,
"blocksCCBot": false,
"robotsRetrieval": "present"
}

Input

Give it a list of domains:

{ "domains": ["example.com", "another.com"] }

Or chain it after any actor that produces a list — Google Maps Scraper, a CRM export, a prospect list — by passing that dataset as items and naming the column that holds the website:

{ "items": [{ "title": "Acme", "website": "acme.com" }], "field": "website" }

Domains it could not read

A domain that blocks the scan, times out, or errors comes back with measured: false, a reason, and every data column null — never a zero and never a false.

That matters at bulk. One fabricated false in a sheet of ten thousand rows is worse than a missing row, because nobody audits the row that looks plausible. Filter on measured and you know exactly what you are looking at.

Roughly one site in eight refuses an automated reader, and this actor uses plain HTTP rather than a browser, so its share is higher.

Notes

  • Reads the homepage only, one page per domain.
  • No browser — plain HTTP requests, so it runs fast and cheap over large lists.