URL Access Checker - robots.txt, Bot Protection & Scrapability avatar

URL Access Checker - robots.txt, Bot Protection & Scrapability

Pricing

from $4.20 / 1,000 url checkeds

Go to Apify Store
URL Access Checker - robots.txt, Bot Protection & Scrapability

URL Access Checker - robots.txt, Bot Protection & Scrapability

Before your AI agent fetches a URL: may it (robots.txt per crawler, RFC 9309) and can it (Cloudflare, DataDome, PerimeterX, AWS WAF, login, JavaScript)? One row per URL with a verdict, safeToFetch true/false and a one-sentence recommendation the agent can follow. It only looks, never bypasses.

Pricing

from $4.20 / 1,000 url checkeds

Rating

0.0

(0)

Developer

NeverEmpty

NeverEmpty

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Before your agent fetches a URL, ask this: may it fetch the page (robots.txt, per crawler name) and can it (does a plain request get the page, or a Cloudflare, DataDome, PerimeterX, Akamai, AWS WAF or Kasada wall, a sign-in or a 429)? One call, one row per URL, with a verdict, a safeToFetch boolean and a one-sentence recommendation an AI agent can follow as is.

It only looks. It never solves or bypasses a challenge page and never requests a page again after a challenge, and it only requests a page when that site's robots.txt allows it. (One exception, for robots.txt only: if the direct route is refused, robots.txt is read once more through the proxy route you chose, so that you get the site's real rules instead of RFC 9309's "refused means everything allowed".)

Example

Input:

{
"urls": ["https://www.instagram.com/nasa/", "https://medium.com/", "https://en.wikipedia.org/wiki/Robots.txt"]
}

Output (production run, shortened, one row per URL):

[
{ "url": "https://www.instagram.com/nasa/", "verdict": "robots-disallowed", "safeToFetch": false,
"recommendation": "Do not fetch: robots.txt disallows this URL for any crawler (Disallow: / on line 288, group User-agent: *).",
"robotsAllowed": false, "robotsGroup": "*", "robotsMatchedRule": "Disallow: /", "robotsMatchedLine": 288, "reachChecked": false },
{ "url": "https://medium.com/", "verdict": "blocked", "safeToFetch": false,
"recommendation": "Allowed by robots.txt, but every request was stopped by a cloudflare challenge page (0/1 direct got the page). This Actor does not solve or bypass challenges; an ordinary HTTP fetch will not get this page.",
"robotsAllowed": true, "directOk": 0, "directAttempts": 1, "httpStatus": 403,
"wall": "challenge", "wallVendor": "cloudflare", "wallEvidence": "cf-mitigated: challenge header" },
{ "url": "https://en.wikipedia.org/wiki/Robots.txt", "verdict": "fetchable", "safeToFetch": true,
"recommendation": "Fetch it: robots.txt allows this URL for any crawler. The page was returned (1/1 direct).",
"robotsAllowed": true, "robotsGroup": "*", "directOk": 1, "httpStatus": 200, "pageTitle": "robots.txt - Wikipedia",
"needsJavaScript": false }
]

With no input at all it checks https://en.wikipedia.org/wiki/Robots.txt, so the Actor always returns a row. Add "userAgents": ["GPTBot"] (or your own crawler name) to judge robots.txt for that crawler.

What it answers, per URL

  1. robots.txt, as RFC 9309 says. For each crawler name you give (GPTBot, ClaudeBot, Googlebot, *, or a full User-Agent string), it picks the group that names that crawler, or the User-agent: * group when there is none, and applies the longest matching Allow/Disallow rule (* and $ wildcards, percent-encoding and non-ASCII paths normalised, Allow wins a tie, /robots.txt is always allowed). You get the matched rule, its line number, the group, Crawl-delay and the Sitemap lines. A robots.txt that answers 404/410 means everything is allowed; 5xx, 429 or no answer means the whole site is treated as disallowed, as the RFC says. robots.txt itself is always requested with an honest bot User-Agent (some sites, such as instagram.com, answer a browser User-Agent with an HTML page instead of the robots file).
  2. Reachability, measured. On the routes you choose (direct = the Actor's own server, datacenter = Apify datacenter proxy, residential = Apify residential proxy with a new IP per try), it requests the page N times and reports how many tries returned the page, the HTTP status and the time for each try. Redirects are followed one hop at a time, and each hop's robots.txt is checked too.
  3. The wall, named. A refused answer is classified from the vendor's own markers: headers such as cf-mitigated: challenge, x-amzn-waf-action: captcha, x-datadome, x-kpsdk-ct, and page markers taken from real challenge pages. Types: challenge, captcha, blocked, forbidden (a 403 without a known vendor marker), login-required (401 or a redirect to a sign-in page), geo-restricted (451 or a region notice), rate-limited (429 with Retry-After). Vendors: Cloudflare, DataDome, PerimeterX (HUMAN), Akamai, AWS WAF, Kasada, Imperva, Sucuri, DDoS-Guard, Vercel, Amazon and Google CAPTCHA pages, hCaptcha, reCAPTCHA and Turnstile. Large ordinary pages that merely load a vendor's script are not called walls; that goes to protectionSignals instead.
  4. The page itself. Whether the returned HTML needs JavaScript to show its content (needsJavaScript with the evidence, such as an empty #root shell), whether the data is embedded anyway (__NEXT_DATA__, JSON-LD), and the page title.

Verdicts

verdictsafeToFetchmeaning
fetchabletruerobots.txt allows it for the first crawler name, and at least one try returned the page
robots-disallowedfalserobots.txt disallows it for the first crawler name (or a redirect leads to a disallowed URL)
robots-unreachablefalserobots.txt answered 5xx/429 or not at all; treat the site as disallowed for now
robots-refusedfalsethe site refused to serve robots.txt (401, 403, 418 and so on), so its rules are unknown. RFC 9309 would treat that as "everything allowed", so the reachability result is still given; choose the datacenter or residential route to let the Actor re-read robots.txt from another IP
allowed-not-testedtruerobots.txt allows it for the first crawler name, but reachability was not tested (turned off, or robots.txt does not allow this Actor itself to request the page)
blockedfalseno try returned the page, and at least one hit a challenge page, CAPTCHA or bot block
login-requiredfalse401 or a redirect to a sign-in page
geo-restrictedfalse451 or a region notice from the tested location
rate-limitedfalse429; see retryAfterSecs
not-foundfalse404 or 410
server-errorfalseevery try answered 5xx
no-answer / http-errorfalserobots.txt answered but the page did not, or answered with another status

Input

FieldDefaultWhat it does
urls(example URL)URLs to check, 1 to 500 per run. Duplicates (also when only the #fragment differs) are checked and charged once.
userAgents*Crawler names or full User-Agent strings for robots.txt. The first one is the subject of verdict; all appear in robotsByAgent.
includeAiCrawlersfalseAlso judge GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, anthropic-ai, Google-Extended, Googlebot, Bingbot, PerplexityBot, Perplexity-User, CCBot, Bytespider, Applebot-Extended, meta-externalagent and Amazonbot.
checkReachtrueRequest the page. Off gives a robots.txt-only answer.
routesdirectAny of direct, datacenter, residential.
attemptsPerRoute11 to 5 tries per route, 0.7 s apart.
residentialCountryanyTwo-letter country code for residential tries (for example US, DE, JP).
fetchAsbrowserbrowser sends a Chrome User-Agent header; bot sends URLAccessChecker/1.0. No TLS or browser fingerprint spoofing either way.
onlyChangesfalseMonitor mode: return only URLs whose robots.txt answer changed since the last run with the same watchName.
watchNameName of the remembered state for monitor mode.
resetMonitoringStatefalseForget the remembered answers before this run.
maxConcurrency4URLs checked in parallel (1 to 8).
requestTimeoutSecs20Time limit per request (5 to 60 s).

Output columns

url, verdict, safeToFetch, recommendation, robotsAgent, robotsAllowed, robotsGroup, robotsMatchedRule, robotsMatchedLine, robotsReason, crawlDelaySecs, robotsByAgent (one entry per crawler name: agent, allowed, group, rule, line, crawlDelaySecs), robotsTxtUrl, robotsTxtStatus, robotsTxtAvailability (found, not-found, unreachable), robotsTxtFetchedVia, sitemaps, reachChecked, reachSkippedReason, fetchedAs, directOk/directAttempts, datacenterOk/datacenterAttempts, residentialOk/residentialAttempts, attempts (each try: route, attempt, httpStatus, ok, wall, vendor, ms, finalUrl, error), wall, wallVendor, wallEvidence, retryAfterSecs, protectionSignals, httpStatus, finalUrl, redirects, contentType, pageTitle, needsJavaScript, jsEvidence, visibleTextChars, dataInHtml, monitor columns (changeType, changedAgents, previousRobotsByAgent, previousCheckedAt, watchName), status, note, checkedAt.

Measured

One production run on 60 commonly linked URLs (news, social, shopping, travel, real estate, developer sites), 3 routes x 2 tries each, 203 seconds, 88 MB of memory:

tries that returned the page
direct56 of 104
datacenter proxy57 of 104
residential proxy (US)69 of 104

Verdicts: 35 fetchable, 15 blocked (Cloudflare 6, DataDome 4, AWS WAF 2, PerimeterX 1, Kasada 1, hCaptcha 1), 8 robots-disallowed (the User-agent: * group of reddit, x.com, instagram, linkedin, facebook and imdb, and path rules on yellowpages and traveloka), 2 not-found. 7 of the 60 pages need JavaScript to show their content (YouTube, youtu.be, Spotify, Vimeo, note.com, TikTok, StubHub).

The wall markers are checked against 31 real pages captured from Apify's routes and a home line, including large ordinary pages that load a vendor's script without blocking (Zillow, leboncoin, StubHub, Ticketmaster, Washington Post) and ordinary 404 pages that carry Cloudflare's beacon script.

Monitor mode

Set watchName and onlyChanges: true and schedule the run. The first run returns every URL. Later runs return only URLs whose robots.txt answer changed for any of your crawler names (for example a site that adds User-agent: GPTBot / Disallow: /), with changedAgents and the previous answer. A crawler name you add later is reported once for every URL. URLs whose answer did not change are not requested at all (no proxy tries). A robots.txt that could not be read reliably this time (5xx, 429, refused with 401/403/418, or an HTML page) is never compared: that URL comes back as a free not-compared row and is compared again on the next run.

Pricing

Pay per event:

  • URL checked: each URL that got an answer from the site (robots.txt or the page), including "do not fetch" answers.
  • Run start: once per run that returned at least one answer (or, in monitor mode, compared at least one URL; a monitor run where nothing changed costs only this).
  • Datacenter try / Residential try: each try on that proxy route that the site answered, to cover the proxy traffic. A try the proxy could not deliver (even after a second IP) is free.

Free: invalid URLs, private or internal addresses, sites that do not answer at all, not-compared and no-change rows in monitor mode, and the rows that say a run hit its maximum total charge. A run whose maximum total charge cannot fit the start fee plus one answer (and its proxy tries) requests nothing and costs nothing.

Use it from an AI agent (MCP)

Through the Apify MCP server the input can be just {"urls": ["https://example.com/page"]}; add "userAgents": ["<your crawler name>"] to judge robots.txt for your own crawler. Read verdict and safeToFetch, and pass recommendation to the model as is.

Limits

  • It does not solve, bypass or retry challenge pages or CAPTCHAs, does not sign in, and does not spoof TLS or browser fingerprints. A blocked verdict means a plain HTTP client was stopped; a real browser may still get through.
  • The page is requested only when robots.txt allows this Actor (URLAccessChecker, otherwise User-agent: *). Otherwise you get the robots.txt answer only, with reachSkippedReason.
  • Results are for the tested routes, time and location. Bot protection often decides per IP and per moment; use more tries per route for a steadier picture.
  • robots.txt group matching is by exact product token (RFC 9309). Vendor-specific fallbacks (for example Googlebot-News falling back to Googlebot) are not applied.
  • Reads at most 1.5 MB of a page on the direct route and 256 KB on proxy routes; non-text answers on proxy routes stop at 16 KB.
  • Not affiliated with any of the sites or vendors named here.

Support

Open an issue on the Issues tab with the URL and the run ID.