Site QA Indexability AI Crawler Report avatar

Site QA Indexability AI Crawler Report

Pricing

from $30.00 / 1,000 ai crawler policy checkeds

Go to Apify Store
Site QA Indexability AI Crawler Report

Site QA Indexability AI Crawler Report

Try Input: 1 URL + 1 page, report/export OFF ≈ USD 0.03/row. From USD 0.03. Paste a URL and Start.

Pricing

from $30.00 / 1,000 ai crawler policy checkeds

Rating

0.0

(0)

Developer

naoki anzai

naoki anzai

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 days ago

Last modified

Categories

Share

Site QA Indexability AI Crawler Report Scraper

Paid first-run JSON ≈ USD 0.03/row (1 URL + 1 page, report/export OFF). Keep reports, exports, and movement options bounded for the first run. Open Try for free → Input and run this JSON.

{
"urls": [
"https://example.com/?siteQaCanary=indexability-p"
],
"maxPages": 1,
"checkRobotsTxt": true,
"checkLlmsTxt": true,
"authorizedUseConfirmed": true,
"aiCrawlerUserAgents": [
"GPTBot",
"Google-Extended"
],
"emitPageRows": false,
"generateReport": false,
"emitExport": false,
"emitUnchanged": false,
"dryRun": false,
"initialRunMode": "baseline_only",
"snapshotKey": "indexability-p"
}

Paste a URL and Start.

Site owners, SEO agencies, and content teams use this actor to audit public indexability and AI crawler access signals. Provide public URLs and optional AI crawler user-agent names. The actor returns source-linked policy observations, indexability issues, reports, and export rows.

The primary deliverable is the site-level indexability report or export. Policy observations and raw page rows are supporting evidence, not the main sales path.

Store Quickstart

Run with dryRun=false and public URLs that you own or are allowed to audit.

{
"urls": ["https://example.com/?siteQaCanary=indexability-ai-crawler-v1"],
"aiCrawlerUserAgents": ["GPTBot", "Google-Extended", "PerplexityBot", "ClaudeBot"],
"checkRobotsTxt": true,
"checkLlmsTxt": true,
"authorizedUseConfirmed": true,
"generateReport": true,
"emitExport": true,
"emitUnchanged": false,
"dryRun": false
}

Use with Apify MCP

Exact actor: taroyamada/site-qa-indexability-ai-crawler-report-scraper

Sample input JSON:

{
"urls": [
"https://example.com/?siteQaCanary=indexability-ai-crawler-v1"
],
"maxPages": 1
}

Report result preview

{
"rowType": "ai_crawler_indexability_report",
"billingEventName": "ai-crawler-indexability-report",
"sourceUrl": "https://example.com/robots.txt",
"checkedUrlCount": 3,
"issueCount": 1,
"reportStatus": "action_needed"
}

Use indexability-export-generated when a source-linked checklist must move into an implementation workflow. The policy event is the low-cost evidence step; it is not a ranking or citation guarantee.

Run the next report

  • Generate ai-crawler-indexability-report when robots.txt, llms.txt, canonical, or noindex findings need a site-level handoff.
  • Generate indexability-export-generated when SEO, content, or AI-readiness teams need a source-linked implementation checklist.
  • Pair the findings with Site QA Content Report Scraper for content issues on the same URL set.
  • Pair the findings with Site QA Broken Link Report Scraper for link-health issues on the same URL set.
  • Audit public support and policy paths with Site QA Contact Page Audit Scraper when crawler findings expose missing response channels.

Input Examples

Audit one page and origin policies

{
"urls": ["https://example.com/blog/launch"],
"aiCrawlerUserAgents": ["GPTBot", "ClaudeBot"],
"checkRobotsTxt": true,
"checkLlmsTxt": true,
"authorizedUseConfirmed": true,
"dryRun": false
}

Batch audit a site section

{
"urls": [
"https://example.com/",
"https://example.com/pricing",
"https://example.com/docs"
],
"maxPages": 25,
"emitPageRows": false,
"generateReport": true,
"authorizedUseConfirmed": true,
"dryRun": false
}

Generate a handoff export

{
"urls": ["https://example.com/landing-page"],
"aiCrawlerUserAgents": ["GPTBot", "Google-Extended", "PerplexityBot"],
"emitExport": true,
"emitUnchanged": false,
"authorizedUseConfirmed": true,
"dryRun": false
}

Sample Output

The complete sample starts with report and export rows, followed by issue and policy rows:

{
"actorName": "site-qa-indexability-ai-crawler-report-scraper",
"rowType": "ai_crawler_indexability_report",
"billingEventName": "ai-crawler-indexability-report",
"sourceUrl": "https://example.com/robots.txt",
"issueCount": 1,
"reportStatus": "action_needed"
}

Output Fields

  • rowType: ai_crawler_policy_observation, indexability_issue, ai_crawler_indexability_report, or indexability_export.
  • billingEventName: PAY_PER_EVENT event name used for the row.
  • sourceUrl: public URL or policy file that supports the row.
  • issueType: detected source-linked issue when applicable.
  • blockedUserAgents: crawler names with broad robots.txt blocks when detected.

Pricing And No-Change Runs

  • ai-crawler-policy-checked: $0.030 per public robots.txt or llms.txt policy observation.
  • indexability-issue-detected: $0.120 per source-linked indexability issue.
  • ai-crawler-indexability-report: $6.000 per site-level report.
  • indexability-export-generated: $8.000 per generated export.

When emitUnchanged=false, repeated unchanged runs emit zero dataset rows and zero charges after state is saved.

For recurring audits, the value path is indexability-issue-detected, ai-crawler-indexability-report, and indexability-export-generated. An unchanged poll does not create a report/export charge.

Compliance Guardrails

  • Public pages, robots.txt, and llms.txt only.
  • No login, paywall, CAPTCHA, private session, credentialed API, or bypass behavior.
  • Non-dry runs require authorizedUseConfirmed=true; use this only for sites you own, manage, or are allowed to audit.
  • This is an unofficial audit tool and is not affiliated with any crawler, search engine, or AI provider.
  • No ranking guarantee, AI citation guarantee, legal conclusion, or compliance certification.

Bundle Paths

See Also