# Reliable Playwright Web Scraper - Python page function (`eszetael_lab/reliable-playwright-scraper`) Actor

Crawl any website with a headless browser and your own Python page function. Reliable runs, honest per-result pricing, clean flat output. Respects robots.txt.

- **URL**: https://apify.com/eszetael\_lab/reliable-playwright-scraper.md
- **Developed by:** [Radosław Szal](https://apify.com/eszetael_lab) (community)
- **Categories:** AI, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js/docs.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python/docs.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

## Reliable Playwright Web Scraper

> 🔗 Part of the **[Apify actors collection](https://github.com/Eszetael/apify-actors)** — 6 reliable, tested scrapers & data tools.

Crawl any site with a headless Chromium browser and your own Python page function.  
Reliable runs, honest per‑result pricing, respects `robots.txt` by default.

### Quick start

```json
{
  "startUrls": [{ "url": "https://apify.com" }],
  "pageFunction": "async def page_function(page, context, request):\n    return {\n        \"url\": page.url,\n        \"title\": await page.title()\n    }"
}
````

### Page function

Signature:

```python
async def page_function(page, context, request):
    # page      – Playwright Page instance
    # context   – Actor context (key‑value store, dataset, etc.)
    # request   – Current Request object (url, userData, …)
    # Return a dict or a list of dicts; each dict becomes one result item.
```

### Crawling

- **linkSelector** – CSS selector for links to follow.
- **includeGlobs** – Glob patterns that URLs must match to be enqueued.
- **maxRequestsPerCrawl** – Hard limit on total requests for the run.

### Input

| Field | Type | Description |
|-------|------|-------------|
| startUrls | array\[object] | List of `{ "url": "…" }` to seed the crawl |
| pageFunction | string | Python source of the async page function |
| linkSelector | string | CSS selector for link discovery |
| includeGlobs | array\[string] | URL glob patterns to follow |
| maxRequestsPerCrawl | integer | Max requests per run |
| maxConcurrency | integer | Parallel browser contexts (default 5) |
| requestTimeoutSecs | integer | Per-page timeout in seconds (default 45) |
| waitUntil | string | Playwright `waitUntil` (`load`, `domcontentloaded`, `networkidle`) |
| respectRobotsTxt | boolean | Obey `robots.txt` (default `true`) |
| proxyConfiguration | object | Apify proxy config |
| maxItems | integer | Stop after this many result items |

### Pricing

- Pay per result item (flat rate).
- First *N* items free (see actor pricing page).
- No hidden fees; you only pay for data you actually receive.

### Notes

- **`robots.txt` is respected by default** (URLs disallowed for crawlers are skipped *before* they are visited). Turn `respectRobotsTxt` off only for sites you own.
- **Your page function runs inside your own run** (your account, your data) — like any code you run on the platform. A per-page timeout bounds runaway code.
- You are responsible for ensuring your scraping complies with applicable laws and the target site's terms of service.

# Actor input Schema

## `startUrls` (type: `array`):

URLs to start crawling from.

## `pageFunction` (type: `string`):

Async Python run on each page. Signature: async def page\_function(page, context, request). `page` is a Playwright Page, `request.url` is the current URL. Return a dict or list of dicts to save to the dataset (or None to skip).

## `linkSelector` (type: `string`):

CSS selector for links to enqueue and follow (e.g. a\[href]). Empty = scrape only the start URLs.

## `includeGlobs` (type: `array`):

Only enqueue URLs matching these glob patterns, e.g. https://example.com/product/\*

## `maxRequestsPerCrawl` (type: `integer`):

Hard cap on the number of pages crawled.

## `maxConcurrency` (type: `integer`):

Maximum pages loaded in parallel.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for a single page (load + page function) before giving up.

## `waitUntil` (type: `string`):

When to consider a page loaded before running the page function.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs disallowed by the site's robots.txt or TDM reservation. Keep on unless you own the target site.

## `proxyConfiguration` (type: `object`):

Proxy for the requests. Datacenter is fine for most sites; use RESIDENTIAL for heavily protected ones.

## `maxItems` (type: `integer`):

Cap on delivered dataset items (0 = up to the safety limit of 50000).

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://apify.com"
    }
  ],
  "pageFunction": "async def page_function(page, context, request):\n    return {\n        'url': request.url,\n        'title': await page.title(),\n    }",
  "maxRequestsPerCrawl": 100,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 45,
  "waitUntil": "load",
  "respectRobotsTxt": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxItems": 0
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://apify.com"
        }
    ],
    "pageFunction": `async def page_function(page, context, request):
    return {
        'url': request.url,
        'title': await page.title(),
    }`
};

// Run the Actor and wait for it to finish
const run = await client.actor("eszetael_lab/reliable-playwright-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://apify.com" }],
    "pageFunction": """async def page_function(page, context, request):
    return {
        'url': request.url,
        'title': await page.title(),
    }""",
}

# Run the Actor and wait for it to finish
run = client.actor("eszetael_lab/reliable-playwright-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://apify.com"
    }
  ],
  "pageFunction": "async def page_function(page, context, request):\\n    return {\\n        '\''url'\'': request.url,\\n        '\''title'\'': await page.title(),\\n    }"
}' |
apify call eszetael_lab/reliable-playwright-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=eszetael_lab/reliable-playwright-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "Reliable Playwright Web Scraper - Python page function",
        "description": "Crawl any website with a headless browser and your own Python page function. Reliable runs, honest per-result pricing, clean flat output. Respects robots.txt.",
        "version": "0.2",
        "x-build-id": "wZdSu3uT9VX2KK9CI"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/eszetael_lab~reliable-playwright-scraper/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-eszetael_lab-reliable-playwright-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/eszetael_lab~reliable-playwright-scraper/runs": {
            "post": {
                "operationId": "runs-sync-eszetael_lab-reliable-playwright-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/eszetael_lab~reliable-playwright-scraper/run-sync": {
            "post": {
                "operationId": "run-sync-eszetael_lab-reliable-playwright-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "required": [
                    "startUrls",
                    "pageFunction"
                ],
                "properties": {
                    "startUrls": {
                        "title": "Start URLs",
                        "type": "array",
                        "description": "URLs to start crawling from.",
                        "items": {
                            "type": "object",
                            "required": [
                                "url"
                            ],
                            "properties": {
                                "url": {
                                    "type": "string",
                                    "title": "URL of a web page",
                                    "format": "uri"
                                }
                            }
                        }
                    },
                    "pageFunction": {
                        "title": "Page function",
                        "type": "string",
                        "description": "Async Python run on each page. Signature: async def page_function(page, context, request). `page` is a Playwright Page, `request.url` is the current URL. Return a dict or list of dicts to save to the dataset (or None to skip)."
                    },
                    "linkSelector": {
                        "title": "Link selector",
                        "type": "string",
                        "description": "CSS selector for links to enqueue and follow (e.g. a[href]). Empty = scrape only the start URLs."
                    },
                    "includeGlobs": {
                        "title": "Include globs",
                        "type": "array",
                        "description": "Only enqueue URLs matching these glob patterns, e.g. https://example.com/product/*",
                        "items": {
                            "type": "object",
                            "required": [
                                "glob"
                            ],
                            "properties": {
                                "glob": {
                                    "type": "string",
                                    "title": "Glob of a web page"
                                }
                            }
                        }
                    },
                    "maxRequestsPerCrawl": {
                        "title": "Max pages to crawl",
                        "type": "integer",
                        "description": "Hard cap on the number of pages crawled.",
                        "default": 100
                    },
                    "maxConcurrency": {
                        "title": "Max concurrency",
                        "type": "integer",
                        "description": "Maximum pages loaded in parallel.",
                        "default": 5
                    },
                    "requestTimeoutSecs": {
                        "title": "Request timeout (seconds)",
                        "type": "integer",
                        "description": "How long to wait for a single page (load + page function) before giving up.",
                        "default": 45
                    },
                    "waitUntil": {
                        "title": "Wait until",
                        "enum": [
                            "load",
                            "domcontentloaded",
                            "networkidle"
                        ],
                        "type": "string",
                        "description": "When to consider a page loaded before running the page function.",
                        "default": "load"
                    },
                    "respectRobotsTxt": {
                        "title": "Respect robots.txt",
                        "type": "boolean",
                        "description": "Skip URLs disallowed by the site's robots.txt or TDM reservation. Keep on unless you own the target site.",
                        "default": true
                    },
                    "proxyConfiguration": {
                        "title": "Proxy configuration",
                        "type": "object",
                        "description": "Proxy for the requests. Datacenter is fine for most sites; use RESIDENTIAL for heavily protected ones.",
                        "default": {
                            "useApifyProxy": true
                        }
                    },
                    "maxItems": {
                        "title": "Max items",
                        "type": "integer",
                        "description": "Cap on delivered dataset items (0 = up to the safety limit of 50000).",
                        "default": 0
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
