# Web Unblocker - Anti-Bot Scraper (Cloudflare, DataDome) (`get_anything/web-unblocker`) Actor

Fetch any URL even behind anti-bot walls (Cloudflare, DataDome, PerimeterX, Akamai, Imperva). Returns HTML plus optional Markdown/text and a screenshot, detects which protection guards the site, and retries via a hardened browser on a fresh residential IP when blocked. No API key.

- **URL**: https://apify.com/get_anything/web-unblocker.md
- **Developed by:** [Get Anything](https://apify.com/get_anything) (community)
- **Categories:** Developer tools
- **Stats:** 26 total users, 18 monthly users, 97.3% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Web Unblocker — Anti-Bot Scraper (Cloudflare, DataDome, PerimeterX, Akamai)

Fetch **any URL and get the page back**, even when it sits behind an anti-bot wall. Point it at a URL and it returns the final HTML (and optionally clean Markdown/text and a screenshot), tells you **which protection guards the site**, and **retries through a hardened browser on a fresh IP** when it gets blocked.

Built for the most common complaint in web scraping: *"I'm blocked from this website, what are my options?"*

### How it works

1. **Fast path** — an HTTP request with `curl_cffi` (real Chrome TLS/JA3 impersonation). Cheap and instant for unprotected pages.
2. **Browser render** — if the fast path is blocked or challenged, the page is rendered in **Camoufox** (a hardened Firefox with humanized fingerprints) which clears most JavaScript anti-bot challenges from Apify IPs.
3. **Retry on a new IP** — if it's *still* blocked, it retries on a fresh browser context with a new residential proxy IP, up to `maxRetries` times.

### Anti-bot detection

Every result tells you what was protecting the site, matched from response headers + body fingerprints:

- **Cloudflare** (`cf-ray`, "Just a moment", challenge-platform)
- **DataDome** (`x-datadome`, captcha-delivery)
- **PerimeterX / HUMAN** (`_px*` cookies, px-captcha)
- **Akamai Bot Manager** (`_abck`, `ak_bmsc`)
- **Imperva / Incapsula** (`incap_ses`, `x-iinfo`)
- **Generic CAPTCHA** (reCAPTCHA / hCaptcha)

### What you get

| Field | Description |
|-------|-------------|
| `url` / `finalUrl` | Requested URL and the URL after redirects |
| `success` | Whether a real (non-challenge) page was returned |
| `statusCode` | HTTP status of the final response |
| `protectionDetected` | Anti-bot vendor detected, or `null` |
| `bypassed` | `true` when a protection was present **and** cleared |
| `method` | `http` or `browser` |
| `html` | Full page HTML *(when output includes HTML)* |
| `markdown` / `text` | Cleaned content with boilerplate removed *(optional)* |
| `wordCount` | Word count of cleaned text *(optional)* |
| `screenshotUrl` | Link to a full-page JPEG in the key-value store *(optional)* |
| `scrapedAt` | ISO 8601 timestamp |

### Input

```json
{
  "startUrls": [{ "url": "https://example.com/protected" }],
  "renderJs": "auto",
  "outputFormat": "html",
  "waitForSelector": "",
  "waitMs": 2000,
  "screenshot": false,
  "maxRetries": 2,
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `startUrls` | array | — | URLs to fetch |
| `renderJs` | string | `auto` | `auto` (HTTP first, browser if blocked), `always`, or `never` |
| `outputFormat` | string | `html` | `html`, `markdown`, `text`, or `all` |
| `includeLinks` | boolean | true | Keep hyperlinks in Markdown |
| `maxChars` | integer | 0 | Truncate Markdown/text (0 = no limit) |
| `waitForSelector` | string | — | CSS selector to wait for (browser render) |
| `waitMs` | integer | 2000 | Wait after load when no selector (browser render) |
| `screenshot` | boolean | false | Save a full-page screenshot |
| `maxRetries` | integer | 2 | Retries on a fresh IP when blocked |
| `proxyConfiguration` | object | Residential | Proxy — residential strongly recommended |

### Use cases

- **Unblock a site you keep getting 403/429 on** and parse the HTML yourself.
- **Check what anti-bot a site runs** before you invest in building a scraper.
- **LLM/RAG ingestion** of pages that need JS rendering, as clean Markdown.
- **Monitoring & QA** of protected pages with screenshots.

### 🤖 Use with Claude or ChatGPT (MCP)

Run this actor straight from Claude, ChatGPT, Cursor or any MCP client via the [Apify MCP server](https://mcp.apify.com). In **Claude Desktop**: Settings → Connectors → Add custom connector → `https://mcp.apify.com`, then ask it to fetch a URL. Or expose just this tool:

```json
{ "mcpServers": { "apify": { "url": "https://mcp.apify.com?tools=get_anything/web-unblocker" } } }
```

Full guide: [Connect Apify actors to Claude & ChatGPT](https://dev.to/get_anything/connect-any-apify-scraper-to-claude-or-chatgpt-in-2-minutes-mcp-37he).

### Notes

- Only publicly accessible pages are fetched; no logins or paywalled content.
- Residential proxy is what makes retries effective — each retry rotates to a new IP.
- For maximum bypass reliability set `renderJs` to `always`.

# Actor input Schema

## `startUrls` (type: `array`):

One or more URLs to fetch. Each is returned even when it sits behind an anti-bot wall.

## `renderJs` (type: `string`):

Auto tries a fast HTTP fetch first and only spins up the hardened browser if the page is blocked or challenged.

## `outputFormat` (type: `string`):

What to return per page. Raw HTML is best if you want to parse it yourself; Markdown/text are cleaned with boilerplate removed.

## `includeLinks` (type: `boolean`):

Preserve inline hyperlinks when producing Markdown.

## `maxChars` (type: `integer`):

Truncate cleaned Markdown/text to this many characters (0 = no limit). Does not affect raw HTML.

## `waitForSelector` (type: `string`):

Optional CSS selector to wait for before capturing (browser render only), e.g. '#price', '.product'.

## `waitMs` (type: `integer`):

How long to wait after the page loads before capturing, when no selector is given (browser render only).

## `screenshot` (type: `boolean`):

Save a full-page JPEG screenshot to the key-value store and add its URL to each item (browser render only).

## `maxRetries` (type: `integer`):

If a page is still blocked after browser render, retry on a fresh browser context + new proxy IP up to this many times.

## `proxyConfiguration` (type: `object`):

Proxy used for fetching. Residential proxies are strongly recommended for protected sites - retries rotate to a new IP.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://nowsecure.nl"
    }
  ],
  "renderJs": "auto",
  "outputFormat": "html",
  "includeLinks": true,
  "maxChars": 0,
  "waitForSelector": "",
  "waitMs": 2000,
  "screenshot": false,
  "maxRetries": 2,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One item per URL in the default dataset - HTML and/or Markdown/text, the detected anti-bot protection, and whether it was bypassed.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://nowsecure.nl"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("get_anything/web-unblocker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://nowsecure.nl" }] }

# Run the Actor and wait for it to finish
run = client.actor("get_anything/web-unblocker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://nowsecure.nl"
    }
  ]
}' |
apify call get_anything/web-unblocker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,get_anything/web-unblocker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vTzUCWxm0zjUk1cXo/builds/EUFUZRzvg72VaPBUO/openapi.json
