# URL Access Checker - robots.txt, Bot Protection & Scrapability (`neverempty/url-access-checker`) Actor

Before your AI agent fetches a URL: may it (robots.txt per crawler, RFC 9309) and can it (Cloudflare, DataDome, PerimeterX, AWS WAF, login, JavaScript)? One row per URL with a verdict, safeToFetch true/false and a one-sentence recommendation the agent can follow. It only looks, never bypasses.

- **URL**: https://apify.com/neverempty/url-access-checker.md
- **Developed by:** [NeverEmpty](https://apify.com/neverempty) (community)
- **Categories:** Developer tools, AI, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.20 / 1,000 url checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## URL Access Checker - robots.txt, Bot Protection & Scrapability

**Before your agent fetches a URL, ask this: may it fetch the page (robots.txt, per crawler name) and can it (does a plain request get the page, or a Cloudflare, DataDome, PerimeterX, Akamai, AWS WAF or Kasada wall, a sign-in or a 429)?** One call, one row per URL, with a `verdict`, a `safeToFetch` boolean and a one-sentence `recommendation` an AI agent can follow as is.

It only looks. It never solves or bypasses a challenge page and never requests a page again after a challenge, and it only requests a page when that site's robots.txt allows it. (One exception, for robots.txt only: if the direct route is refused, robots.txt is read once more through the proxy route you chose, so that you get the site's real rules instead of RFC 9309's "refused means everything allowed".)

### Example

Input:

```json
{
  "urls": ["https://www.instagram.com/nasa/", "https://medium.com/", "https://en.wikipedia.org/wiki/Robots.txt"]
}
```

Output (production run, shortened, one row per URL):

```json
[
  { "url": "https://www.instagram.com/nasa/", "verdict": "robots-disallowed", "safeToFetch": false,
    "recommendation": "Do not fetch: robots.txt disallows this URL for any crawler (Disallow: / on line 288, group User-agent: *).",
    "robotsAllowed": false, "robotsGroup": "*", "robotsMatchedRule": "Disallow: /", "robotsMatchedLine": 288, "reachChecked": false },
  { "url": "https://medium.com/", "verdict": "blocked", "safeToFetch": false,
    "recommendation": "Allowed by robots.txt, but every request was stopped by a cloudflare challenge page (0/1 direct got the page). This Actor does not solve or bypass challenges; an ordinary HTTP fetch will not get this page.",
    "robotsAllowed": true, "directOk": 0, "directAttempts": 1, "httpStatus": 403,
    "wall": "challenge", "wallVendor": "cloudflare", "wallEvidence": "cf-mitigated: challenge header" },
  { "url": "https://en.wikipedia.org/wiki/Robots.txt", "verdict": "fetchable", "safeToFetch": true,
    "recommendation": "Fetch it: robots.txt allows this URL for any crawler. The page was returned (1/1 direct).",
    "robotsAllowed": true, "robotsGroup": "*", "directOk": 1, "httpStatus": 200, "pageTitle": "robots.txt - Wikipedia",
    "needsJavaScript": false }
]
```

With no input at all it checks `https://en.wikipedia.org/wiki/Robots.txt`, so the Actor always returns a row. Add `"userAgents": ["GPTBot"]` (or your own crawler name) to judge robots.txt for that crawler.

### What it answers, per URL

1. **robots.txt, as RFC 9309 says.** For each crawler name you give (`GPTBot`, `ClaudeBot`, `Googlebot`, `*`, or a full User-Agent string), it picks the group that names that crawler, or the `User-agent: *` group when there is none, and applies the longest matching `Allow`/`Disallow` rule (`*` and `$` wildcards, percent-encoding and non-ASCII paths normalised, `Allow` wins a tie, `/robots.txt` is always allowed). You get the matched rule, its line number, the group, `Crawl-delay` and the `Sitemap` lines. A robots.txt that answers 404/410 means everything is allowed; 5xx, 429 or no answer means the whole site is treated as disallowed, as the RFC says. robots.txt itself is always requested with an honest bot User-Agent (some sites, such as instagram.com, answer a browser User-Agent with an HTML page instead of the robots file).
2. **Reachability, measured.** On the routes you choose (`direct` = the Actor's own server, `datacenter` = Apify datacenter proxy, `residential` = Apify residential proxy with a new IP per try), it requests the page N times and reports how many tries returned the page, the HTTP status and the time for each try. Redirects are followed one hop at a time, and each hop's robots.txt is checked too.
3. **The wall, named.** A refused answer is classified from the vendor's own markers: headers such as `cf-mitigated: challenge`, `x-amzn-waf-action: captcha`, `x-datadome`, `x-kpsdk-ct`, and page markers taken from real challenge pages. Types: `challenge`, `captcha`, `blocked`, `forbidden` (a 403 without a known vendor marker), `login-required` (401 or a redirect to a sign-in page), `geo-restricted` (451 or a region notice), `rate-limited` (429 with `Retry-After`). Vendors: Cloudflare, DataDome, PerimeterX (HUMAN), Akamai, AWS WAF, Kasada, Imperva, Sucuri, DDoS-Guard, Vercel, Amazon and Google CAPTCHA pages, hCaptcha, reCAPTCHA and Turnstile. Large ordinary pages that merely load a vendor's script are **not** called walls; that goes to `protectionSignals` instead.
4. **The page itself.** Whether the returned HTML needs JavaScript to show its content (`needsJavaScript` with the evidence, such as an empty `#root` shell), whether the data is embedded anyway (`__NEXT_DATA__`, JSON-LD), and the page title.

### Verdicts

| verdict | safeToFetch | meaning |
|---|---|---|
| `fetchable` | true | robots.txt allows it for the first crawler name, and at least one try returned the page |
| `robots-disallowed` | false | robots.txt disallows it for the first crawler name (or a redirect leads to a disallowed URL) |
| `robots-unreachable` | false | robots.txt answered 5xx/429 or not at all; treat the site as disallowed for now |
| `robots-refused` | false | the site refused to serve robots.txt (401, 403, 418 and so on), so its rules are unknown. RFC 9309 would treat that as "everything allowed", so the reachability result is still given; choose the datacenter or residential route to let the Actor re-read robots.txt from another IP |
| `allowed-not-tested` | true | robots.txt allows it for the first crawler name, but reachability was not tested (turned off, or robots.txt does not allow this Actor itself to request the page) |
| `blocked` | false | no try returned the page, and at least one hit a challenge page, CAPTCHA or bot block |
| `login-required` | false | 401 or a redirect to a sign-in page |
| `geo-restricted` | false | 451 or a region notice from the tested location |
| `rate-limited` | false | 429; see `retryAfterSecs` |
| `not-found` | false | 404 or 410 |
| `server-error` | false | every try answered 5xx |
| `no-answer` / `http-error` | false | robots.txt answered but the page did not, or answered with another status |

### Input

| Field | Default | What it does |
|---|---|---|
| `urls` | (example URL) | URLs to check, 1 to 500 per run. Duplicates (also when only the #fragment differs) are checked and charged once. |
| `userAgents` | `*` | Crawler names or full User-Agent strings for robots.txt. The first one is the subject of `verdict`; all appear in `robotsByAgent`. |
| `includeAiCrawlers` | false | Also judge GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, anthropic-ai, Google-Extended, Googlebot, Bingbot, PerplexityBot, Perplexity-User, CCBot, Bytespider, Applebot-Extended, meta-externalagent and Amazonbot. |
| `checkReach` | true | Request the page. Off gives a robots.txt-only answer. |
| `routes` | `direct` | Any of `direct`, `datacenter`, `residential`. |
| `attemptsPerRoute` | 1 | 1 to 5 tries per route, 0.7 s apart. |
| `residentialCountry` | any | Two-letter country code for residential tries (for example `US`, `DE`, `JP`). |
| `fetchAs` | `browser` | `browser` sends a Chrome User-Agent header; `bot` sends `URLAccessChecker/1.0`. No TLS or browser fingerprint spoofing either way. |
| `onlyChanges` | false | Monitor mode: return only URLs whose robots.txt answer changed since the last run with the same `watchName`. |
| `watchName` | | Name of the remembered state for monitor mode. |
| `resetMonitoringState` | false | Forget the remembered answers before this run. |
| `maxConcurrency` | 4 | URLs checked in parallel (1 to 8). |
| `requestTimeoutSecs` | 20 | Time limit per request (5 to 60 s). |

### Output columns

`url`, `verdict`, `safeToFetch`, `recommendation`, `robotsAgent`, `robotsAllowed`, `robotsGroup`, `robotsMatchedRule`, `robotsMatchedLine`, `robotsReason`, `crawlDelaySecs`, `robotsByAgent` (one entry per crawler name: `agent`, `allowed`, `group`, `rule`, `line`, `crawlDelaySecs`), `robotsTxtUrl`, `robotsTxtStatus`, `robotsTxtAvailability` (`found`, `not-found`, `unreachable`), `robotsTxtFetchedVia`, `sitemaps`, `reachChecked`, `reachSkippedReason`, `fetchedAs`, `directOk`/`directAttempts`, `datacenterOk`/`datacenterAttempts`, `residentialOk`/`residentialAttempts`, `attempts` (each try: `route`, `attempt`, `httpStatus`, `ok`, `wall`, `vendor`, `ms`, `finalUrl`, `error`), `wall`, `wallVendor`, `wallEvidence`, `retryAfterSecs`, `protectionSignals`, `httpStatus`, `finalUrl`, `redirects`, `contentType`, `pageTitle`, `needsJavaScript`, `jsEvidence`, `visibleTextChars`, `dataInHtml`, monitor columns (`changeType`, `changedAgents`, `previousRobotsByAgent`, `previousCheckedAt`, `watchName`), `status`, `note`, `checkedAt`.

### Measured

One production run on 60 commonly linked URLs (news, social, shopping, travel, real estate, developer sites), 3 routes x 2 tries each, 203 seconds, 88 MB of memory:

| | tries that returned the page |
|---|---|
| direct | 56 of 104 |
| datacenter proxy | 57 of 104 |
| residential proxy (US) | 69 of 104 |

Verdicts: 35 fetchable, 15 blocked (Cloudflare 6, DataDome 4, AWS WAF 2, PerimeterX 1, Kasada 1, hCaptcha 1), 8 robots-disallowed (the `User-agent: *` group of reddit, x.com, instagram, linkedin, facebook and imdb, and path rules on yellowpages and traveloka), 2 not-found. 7 of the 60 pages need JavaScript to show their content (YouTube, youtu.be, Spotify, Vimeo, note.com, TikTok, StubHub).

The wall markers are checked against 31 real pages captured from Apify's routes and a home line, including large ordinary pages that load a vendor's script without blocking (Zillow, leboncoin, StubHub, Ticketmaster, Washington Post) and ordinary 404 pages that carry Cloudflare's beacon script.

### Monitor mode

Set `watchName` and `onlyChanges: true` and schedule the run. The first run returns every URL. Later runs return only URLs whose robots.txt answer changed for any of your crawler names (for example a site that adds `User-agent: GPTBot` / `Disallow: /`), with `changedAgents` and the previous answer. A crawler name you add later is reported once for every URL. URLs whose answer did not change are not requested at all (no proxy tries). A robots.txt that could not be read reliably this time (5xx, 429, refused with 401/403/418, or an HTML page) is never compared: that URL comes back as a free `not-compared` row and is compared again on the next run.

### Pricing

Pay per event:

- **URL checked**: each URL that got an answer from the site (robots.txt or the page), including "do not fetch" answers.
- **Run start**: once per run that returned at least one answer (or, in monitor mode, compared at least one URL; a monitor run where nothing changed costs only this).
- **Datacenter try** / **Residential try**: each try on that proxy route that the site answered, to cover the proxy traffic. A try the proxy could not deliver (even after a second IP) is free.

Free: invalid URLs, private or internal addresses, sites that do not answer at all, `not-compared` and `no-change` rows in monitor mode, and the rows that say a run hit its maximum total charge. A run whose maximum total charge cannot fit the start fee plus one answer (and its proxy tries) requests nothing and costs nothing.

### Use it from an AI agent (MCP)

Through the Apify MCP server the input can be just `{"urls": ["https://example.com/page"]}`; add `"userAgents": ["<your crawler name>"]` to judge robots.txt for your own crawler. Read `verdict` and `safeToFetch`, and pass `recommendation` to the model as is.

### Limits

- It does not solve, bypass or retry challenge pages or CAPTCHAs, does not sign in, and does not spoof TLS or browser fingerprints. A `blocked` verdict means a plain HTTP client was stopped; a real browser may still get through.
- The page is requested only when robots.txt allows this Actor (`URLAccessChecker`, otherwise `User-agent: *`). Otherwise you get the robots.txt answer only, with `reachSkippedReason`.
- Results are for the tested routes, time and location. Bot protection often decides per IP and per moment; use more tries per route for a steadier picture.
- robots.txt group matching is by exact product token (RFC 9309). Vendor-specific fallbacks (for example Googlebot-News falling back to Googlebot) are not applied.
- Reads at most 1.5 MB of a page on the direct route and 256 KB on proxy routes; non-text answers on proxy routes stop at 16 KB.
- Not affiliated with any of the sites or vendors named here.

### Support

Open an issue on the Issues tab with the URL and the run ID.

# Actor input Schema

## `urls` (type: `array`):

URLs to check, one per line (1 to 500 per run). A URL without http:// or https:// is read as https://. The same URL given twice (also when only the #fragment differs) is checked and charged once. If this is empty, the example URL https://en.wikipedia.org/wiki/Robots.txt is checked.

## `userAgents` (type: `array`):

Crawler names (robots.txt product tokens such as GPTBot, ClaudeBot or Googlebot, or a full User-Agent string) to judge robots.txt for. The first one is the subject of verdict, safeToFetch and recommendation; every name gets its own entry in robotsByAgent. Following RFC 9309, a name that has its own group in robots.txt follows only that group; otherwise the User-agent: \* group applies. Empty means \* (any crawler).

## `includeAiCrawlers` (type: `boolean`):

Add GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, anthropic-ai, Google-Extended, Googlebot, Bingbot, PerplexityBot, Perplexity-User, CCBot, Bytespider, Applebot-Extended, meta-externalagent and Amazonbot to robotsByAgent. It only reads robots.txt, so it costs nothing extra.

## `checkReach` (type: `boolean`):

Request the page through the chosen routes and report status codes, walls and whether it needs JavaScript. The page is requested only when robots.txt allows this Actor (URLAccessChecker, else User-agent: \*); otherwise only the robots.txt answer is returned. Turn off for a robots.txt-only answer.

## `routes` (type: `array`):

Where the page is requested from: direct (the Actor's own server, no proxy), datacenter (Apify datacenter proxy) and residential (Apify residential proxy, a new IP for every try). Datacenter and residential tries are charged per try that the site answered, to cover the proxy traffic.

## `attemptsPerRoute` (type: `integer`):

How many times the page is requested on each route (1 to 5). More tries show how reliable a route is (for example 2 of 3). Tries are spaced 0.7 seconds apart. A challenge page is recorded, never solved or retried.

## `residentialCountry` (type: `string`):

Optional two-letter country code for residential tries (for example US, DE, JP), useful to see geo-restricted pages. Empty means any country.

## `fetchAs` (type: `string`):

browser sends an ordinary Chrome User-Agent header (what most scrapers and fetch tools send); bot sends an honest URLAccessChecker bot header. robots.txt itself is always requested with the bot header. There is no TLS or browser fingerprint spoofing either way.

## `onlyChanges` (type: `boolean`):

Return only URLs whose robots.txt answer (allowed or not, and the matched rule, for each crawler name) changed since the last run with the same watch name, plus URLs new to the watch. The first run returns every URL as the starting point. A run where robots.txt could not be read does not count as a change. A run in which nothing changed returns a free row saying so and charges only the run start fee (and any proxy tries).

## `watchName` (type: `string`):

Name of the remembered state used to compare runs (letters, digits, dot, dash, underscore; up to 40). Setting it (or turning on monitor mode) fills changeType and the previous robots.txt answers. Use a different name for each list of URLs you track on its own schedule. With monitor mode on and no name, the name default is used.

## `resetMonitoringState` (type: `boolean`):

Start this watch over: forget the remembered answers before this run, so every URL is returned as a first check.

## `maxConcurrency` (type: `integer`):

How many URLs are checked in parallel (1 to 8). Tries for one URL always run one after another.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for one answer (5 to 60 seconds).

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Robots.txt",
    "https://www.instagram.com/nasa/",
    "https://github.com/settings/profile"
  ],
  "userAgents": [
    "*"
  ],
  "includeAiCrawlers": false,
  "checkReach": true,
  "routes": [
    "direct"
  ],
  "attemptsPerRoute": 1,
  "fetchAs": "browser",
  "onlyChanges": false,
  "resetMonitoringState": false,
  "maxConcurrency": 4,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

One row per URL: verdict (fetchable, robots-disallowed, blocked, login-required, rate-limited, not-found and so on), safeToFetch, a one-sentence recommendation, the robots.txt answer per crawler with the matched rule, line and group, Crawl-delay and sitemaps, how many tries per route returned the page, the wall type and vendor with the evidence, whether the page needs JavaScript, and the page title. An invalid URL or a site that does not answer at all comes back as a free row that says why.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Robots.txt",
        "https://www.instagram.com/nasa/",
        "https://github.com/settings/profile"
    ],
    "userAgents": [
        "*"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neverempty/url-access-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Robots.txt",
        "https://www.instagram.com/nasa/",
        "https://github.com/settings/profile",
    ],
    "userAgents": ["*"],
}

# Run the Actor and wait for it to finish
run = client.actor("neverempty/url-access-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/Robots.txt",
    "https://www.instagram.com/nasa/",
    "https://github.com/settings/profile"
  ],
  "userAgents": [
    "*"
  ]
}' |
apify call neverempty/url-access-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neverempty/url-access-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gWSZd0IauHZIqZ4v4/builds/bYehqMlXnD03qr0zQ/openapi.json
