Robots.txt & XML Sitemap Deep Analyzer
Pricing
from $1.30 / 1,000 robots.txt & sitemap audits
Go to Apify Store

Robots.txt & XML Sitemap Deep Analyzer
Parses robots.txt crawl rules and XML sitemaps. Detects disallowed paths, sitemap locations, AI web scraper permissions (GPTBot, ClaudeBot, CCBot), and index URL counts.
Pricing
from $1.30 / 1,000 robots.txt & sitemap audits
Rating
0.0
(0)
Developer
David Sandor
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Robots.txt & XML Sitemap Deep Analyzer 🚀
Parses robots.txt crawl rules and XML sitemaps. Detects disallowed paths, sitemap locations, AI web scraper permissions (GPTBot, ClaudeBot, CCBot), and index URL counts.
🌟 20+ Enterprise Enhancements (v2.0)
- RAG & LLM Ready: Pre-computed OpenAI token counts and chunked embeddings.
- Smart Keyword Filters: Include or exclude records by targeted keyword lists.
- Sentiment Scoring: Built-in lexical sentiment rating on text contents.
- Noise & Tracking Scrubber: Removes tracking query parameters and boilerplate banners.
- Pay-Per-Event (PPE): Ultra-cost-effective pricing per extracted item.
- Zero Cold Start: Sub-second execution with automated fallback guarantees.
💻 Integration Examples
Node.js (Apify Client)
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });const run = await client.actor('robots-txt-sitemap-crawler').call({// Pass customized inputs hereenableRagEnrichment: true});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log('Extracted Items:', items);
Python
from apify_client import ApifyClientclient = ApifyClient('YOUR_APIFY_TOKEN')run = client.actor('robots-txt-sitemap-crawler').call(run_input={ 'enableRagEnrichment': True })for item in client.dataset(run['defaultDatasetId']).iterate_items():print(item)
📄 Output Schema
Returns structured JSON, token counts, and RAG vector chunks.