URL Access Checker - robots.txt, Bot Protection & Scrapability
Pricing
from $4.20 / 1,000 url checkeds
URL Access Checker - robots.txt, Bot Protection & Scrapability
Before your AI agent fetches a URL: may it (robots.txt per crawler, RFC 9309) and can it (Cloudflare, DataDome, PerimeterX, AWS WAF, login, JavaScript)? One row per URL with a verdict, safeToFetch true/false and a one-sentence recommendation the agent can follow. It only looks, never bypasses.
Pricing
from $4.20 / 1,000 url checkeds
Rating
0.0
(0)
Developer
NeverEmpty
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Before your agent fetches a URL, ask this: may it fetch the page (robots.txt, per crawler name) and can it (does a plain request get the page, or a Cloudflare, DataDome, PerimeterX, Akamai, AWS WAF or Kasada wall, a sign-in or a 429)? One call, one row per URL, with a verdict, a safeToFetch boolean and a one-sentence recommendation an AI agent can follow as is.
It only looks. It never solves or bypasses a challenge page and never requests a page again after a challenge, and it only requests a page when that site's robots.txt allows it. (One exception, for robots.txt only: if the direct route is refused, robots.txt is read once more through the proxy route you chose, so that you get the site's real rules instead of RFC 9309's "refused means everything allowed".)
Example
Input:
{"urls": ["https://www.instagram.com/nasa/", "https://medium.com/", "https://en.wikipedia.org/wiki/Robots.txt"]}
Output (production run, shortened, one row per URL):
[{ "url": "https://www.instagram.com/nasa/", "verdict": "robots-disallowed", "safeToFetch": false,"recommendation": "Do not fetch: robots.txt disallows this URL for any crawler (Disallow: / on line 288, group User-agent: *).","robotsAllowed": false, "robotsGroup": "*", "robotsMatchedRule": "Disallow: /", "robotsMatchedLine": 288, "reachChecked": false },{ "url": "https://medium.com/", "verdict": "blocked", "safeToFetch": false,"recommendation": "Allowed by robots.txt, but every request was stopped by a cloudflare challenge page (0/1 direct got the page). This Actor does not solve or bypass challenges; an ordinary HTTP fetch will not get this page.","robotsAllowed": true, "directOk": 0, "directAttempts": 1, "httpStatus": 403,"wall": "challenge", "wallVendor": "cloudflare", "wallEvidence": "cf-mitigated: challenge header" },{ "url": "https://en.wikipedia.org/wiki/Robots.txt", "verdict": "fetchable", "safeToFetch": true,"recommendation": "Fetch it: robots.txt allows this URL for any crawler. The page was returned (1/1 direct).","robotsAllowed": true, "robotsGroup": "*", "directOk": 1, "httpStatus": 200, "pageTitle": "robots.txt - Wikipedia","needsJavaScript": false }]
With no input at all it checks https://en.wikipedia.org/wiki/Robots.txt, so the Actor always returns a row. Add "userAgents": ["GPTBot"] (or your own crawler name) to judge robots.txt for that crawler.
What it answers, per URL
- robots.txt, as RFC 9309 says. For each crawler name you give (
GPTBot,ClaudeBot,Googlebot,*, or a full User-Agent string), it picks the group that names that crawler, or theUser-agent: *group when there is none, and applies the longest matchingAllow/Disallowrule (*and$wildcards, percent-encoding and non-ASCII paths normalised,Allowwins a tie,/robots.txtis always allowed). You get the matched rule, its line number, the group,Crawl-delayand theSitemaplines. A robots.txt that answers 404/410 means everything is allowed; 5xx, 429 or no answer means the whole site is treated as disallowed, as the RFC says. robots.txt itself is always requested with an honest bot User-Agent (some sites, such as instagram.com, answer a browser User-Agent with an HTML page instead of the robots file). - Reachability, measured. On the routes you choose (
direct= the Actor's own server,datacenter= Apify datacenter proxy,residential= Apify residential proxy with a new IP per try), it requests the page N times and reports how many tries returned the page, the HTTP status and the time for each try. Redirects are followed one hop at a time, and each hop's robots.txt is checked too. - The wall, named. A refused answer is classified from the vendor's own markers: headers such as
cf-mitigated: challenge,x-amzn-waf-action: captcha,x-datadome,x-kpsdk-ct, and page markers taken from real challenge pages. Types:challenge,captcha,blocked,forbidden(a 403 without a known vendor marker),login-required(401 or a redirect to a sign-in page),geo-restricted(451 or a region notice),rate-limited(429 withRetry-After). Vendors: Cloudflare, DataDome, PerimeterX (HUMAN), Akamai, AWS WAF, Kasada, Imperva, Sucuri, DDoS-Guard, Vercel, Amazon and Google CAPTCHA pages, hCaptcha, reCAPTCHA and Turnstile. Large ordinary pages that merely load a vendor's script are not called walls; that goes toprotectionSignalsinstead. - The page itself. Whether the returned HTML needs JavaScript to show its content (
needsJavaScriptwith the evidence, such as an empty#rootshell), whether the data is embedded anyway (__NEXT_DATA__, JSON-LD), and the page title.
Verdicts
| verdict | safeToFetch | meaning |
|---|---|---|
fetchable | true | robots.txt allows it for the first crawler name, and at least one try returned the page |
robots-disallowed | false | robots.txt disallows it for the first crawler name (or a redirect leads to a disallowed URL) |
robots-unreachable | false | robots.txt answered 5xx/429 or not at all; treat the site as disallowed for now |
robots-refused | false | the site refused to serve robots.txt (401, 403, 418 and so on), so its rules are unknown. RFC 9309 would treat that as "everything allowed", so the reachability result is still given; choose the datacenter or residential route to let the Actor re-read robots.txt from another IP |
allowed-not-tested | true | robots.txt allows it for the first crawler name, but reachability was not tested (turned off, or robots.txt does not allow this Actor itself to request the page) |
blocked | false | no try returned the page, and at least one hit a challenge page, CAPTCHA or bot block |
login-required | false | 401 or a redirect to a sign-in page |
geo-restricted | false | 451 or a region notice from the tested location |
rate-limited | false | 429; see retryAfterSecs |
not-found | false | 404 or 410 |
server-error | false | every try answered 5xx |
no-answer / http-error | false | robots.txt answered but the page did not, or answered with another status |
Input
| Field | Default | What it does |
|---|---|---|
urls | (example URL) | URLs to check, 1 to 500 per run. Duplicates (also when only the #fragment differs) are checked and charged once. |
userAgents | * | Crawler names or full User-Agent strings for robots.txt. The first one is the subject of verdict; all appear in robotsByAgent. |
includeAiCrawlers | false | Also judge GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, anthropic-ai, Google-Extended, Googlebot, Bingbot, PerplexityBot, Perplexity-User, CCBot, Bytespider, Applebot-Extended, meta-externalagent and Amazonbot. |
checkReach | true | Request the page. Off gives a robots.txt-only answer. |
routes | direct | Any of direct, datacenter, residential. |
attemptsPerRoute | 1 | 1 to 5 tries per route, 0.7 s apart. |
residentialCountry | any | Two-letter country code for residential tries (for example US, DE, JP). |
fetchAs | browser | browser sends a Chrome User-Agent header; bot sends URLAccessChecker/1.0. No TLS or browser fingerprint spoofing either way. |
onlyChanges | false | Monitor mode: return only URLs whose robots.txt answer changed since the last run with the same watchName. |
watchName | Name of the remembered state for monitor mode. | |
resetMonitoringState | false | Forget the remembered answers before this run. |
maxConcurrency | 4 | URLs checked in parallel (1 to 8). |
requestTimeoutSecs | 20 | Time limit per request (5 to 60 s). |
Output columns
url, verdict, safeToFetch, recommendation, robotsAgent, robotsAllowed, robotsGroup, robotsMatchedRule, robotsMatchedLine, robotsReason, crawlDelaySecs, robotsByAgent (one entry per crawler name: agent, allowed, group, rule, line, crawlDelaySecs), robotsTxtUrl, robotsTxtStatus, robotsTxtAvailability (found, not-found, unreachable), robotsTxtFetchedVia, sitemaps, reachChecked, reachSkippedReason, fetchedAs, directOk/directAttempts, datacenterOk/datacenterAttempts, residentialOk/residentialAttempts, attempts (each try: route, attempt, httpStatus, ok, wall, vendor, ms, finalUrl, error), wall, wallVendor, wallEvidence, retryAfterSecs, protectionSignals, httpStatus, finalUrl, redirects, contentType, pageTitle, needsJavaScript, jsEvidence, visibleTextChars, dataInHtml, monitor columns (changeType, changedAgents, previousRobotsByAgent, previousCheckedAt, watchName), status, note, checkedAt.
Measured
One production run on 60 commonly linked URLs (news, social, shopping, travel, real estate, developer sites), 3 routes x 2 tries each, 203 seconds, 88 MB of memory:
| tries that returned the page | |
|---|---|
| direct | 56 of 104 |
| datacenter proxy | 57 of 104 |
| residential proxy (US) | 69 of 104 |
Verdicts: 35 fetchable, 15 blocked (Cloudflare 6, DataDome 4, AWS WAF 2, PerimeterX 1, Kasada 1, hCaptcha 1), 8 robots-disallowed (the User-agent: * group of reddit, x.com, instagram, linkedin, facebook and imdb, and path rules on yellowpages and traveloka), 2 not-found. 7 of the 60 pages need JavaScript to show their content (YouTube, youtu.be, Spotify, Vimeo, note.com, TikTok, StubHub).
The wall markers are checked against 31 real pages captured from Apify's routes and a home line, including large ordinary pages that load a vendor's script without blocking (Zillow, leboncoin, StubHub, Ticketmaster, Washington Post) and ordinary 404 pages that carry Cloudflare's beacon script.
Monitor mode
Set watchName and onlyChanges: true and schedule the run. The first run returns every URL. Later runs return only URLs whose robots.txt answer changed for any of your crawler names (for example a site that adds User-agent: GPTBot / Disallow: /), with changedAgents and the previous answer. A crawler name you add later is reported once for every URL. URLs whose answer did not change are not requested at all (no proxy tries). A robots.txt that could not be read reliably this time (5xx, 429, refused with 401/403/418, or an HTML page) is never compared: that URL comes back as a free not-compared row and is compared again on the next run.
Pricing
Pay per event:
- URL checked: each URL that got an answer from the site (robots.txt or the page), including "do not fetch" answers.
- Run start: once per run that returned at least one answer (or, in monitor mode, compared at least one URL; a monitor run where nothing changed costs only this).
- Datacenter try / Residential try: each try on that proxy route that the site answered, to cover the proxy traffic. A try the proxy could not deliver (even after a second IP) is free.
Free: invalid URLs, private or internal addresses, sites that do not answer at all, not-compared and no-change rows in monitor mode, and the rows that say a run hit its maximum total charge. A run whose maximum total charge cannot fit the start fee plus one answer (and its proxy tries) requests nothing and costs nothing.
Use it from an AI agent (MCP)
Through the Apify MCP server the input can be just {"urls": ["https://example.com/page"]}; add "userAgents": ["<your crawler name>"] to judge robots.txt for your own crawler. Read verdict and safeToFetch, and pass recommendation to the model as is.
Limits
- It does not solve, bypass or retry challenge pages or CAPTCHAs, does not sign in, and does not spoof TLS or browser fingerprints. A
blockedverdict means a plain HTTP client was stopped; a real browser may still get through. - The page is requested only when robots.txt allows this Actor (
URLAccessChecker, otherwiseUser-agent: *). Otherwise you get the robots.txt answer only, withreachSkippedReason. - Results are for the tested routes, time and location. Bot protection often decides per IP and per moment; use more tries per route for a steadier picture.
- robots.txt group matching is by exact product token (RFC 9309). Vendor-specific fallbacks (for example Googlebot-News falling back to Googlebot) are not applied.
- Reads at most 1.5 MB of a page on the direct route and 256 KB on proxy routes; non-text answers on proxy routes stop at 16 KB.
- Not affiliated with any of the sites or vendors named here.
Support
Open an issue on the Issues tab with the URL and the run ID.