# Playwright Scraper — Custom DOM Records (`automation-lab/public-web-custom-browser-page-function`) Actor

Extract custom JSON records from rendered public pages with an isolated DOM page function. Bound pages, depth, time and output; optionally follow same-origin links. No Node hooks, login or challenge bypass.

- **URL**: https://apify.com/automation-lab/public-web-custom-browser-page-function.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.68 / 1,000 item extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Playwright Scraper — Custom DOM Records

This playwright scraper turns rendered public web pages into your own JSON records. Supply start URLs and a JavaScript DOM `pageFunction`; optionally follow same-origin links. It is for developers building repeatable extraction jobs, catalog sampling and public-page analysis without running server-side customer code.

### Who is it for?

Developers defining custom DOM records for repeatable public-page extraction, data engineers feeding JSON pipelines, and analysts sampling publicly accessible catalogs. It is not a login automation or protected-site unlocker.

### Why use this Actor?

- Keep extraction logic in a small reusable DOM function.
- Extract several objects per page, not just a page title or raw HTML dump.
- Bound pages, depth, records, time and result size.
- Evaluate your function in an isolated **offline DOM snapshot**, without Node hooks, page cookies or original site scripts.
- Use the same record envelope across unrelated public websites.

This is a deliberately narrower tool than a general Playwright code runner. It cannot log in, click through a challenge, execute Node hooks or access private networks. It does not promise access to every website.

### Getting started

1. Enter a public HTTP(S) start URL.
2. Supply a JavaScript function expression such as the example below.
3. Start with `maxPages: 1` and `maxDepth: 0`.
4. Inspect the `data` object in the default dataset.
5. Add a same-origin `linkSelector` and increase depth only after checking the site structure.

```json
{
  "startUrls": [{ "url": "https://quotes.toscrape.com/" }],
  "pageFunction": "({ document, url }) => ({ title: document.title, heading: document.querySelector('h1')?.textContent?.trim() ?? null, url })",
  "maxPages": 1,
  "maxDepth": 0,
  "maxItems": 20,
  "pageTimeoutSecs": 15
}
```

### How the DOM page function works

Chromium renders the public page, waiting for network idle. The Actor then copies its HTML into a separate offline browser context. Scripts, frames, embedded objects, base tags, refresh metadata and inline event handlers are removed. Your function receives `{ document, url }`, where `url` is the loaded source URL, not the snapshot's `about:blank` location.

Return a JSON object, an array of JSON objects, or `null`/`undefined` for no records. Each record must still be an object after JSON serialization: sparse arrays, Date records and `toJSON` conversions to primitives are rejected. The entire page output must fit 256 KiB before records are truncated to the remaining limit. Async functions are supported, but external network calls are blocked. Use `new URL(relativeHref, url)` to resolve relative links.

```javascript
({ document, url }) => Array.from(document.querySelectorAll('.quote')).map(quote => ({
  text: quote.querySelector('.text')?.textContent,
  author: quote.querySelector('.author')?.textContent,
  url
}))
```

The function does not receive Playwright's `page`, Node's `require`/`process`, Actor storage, tokens, browser cookies or a live authenticated session. Shadow DOM, canvas pixels, transient JavaScript state and interactions with the original page are not preserved in the HTML snapshot.

### Inputs

| Input | Meaning |
| --- | --- |
| `startUrls` | Required list of 1–20 public HTTP(S) URLs; strings or `{url}` objects; optional `method: "GET"` only |
| `pageFunction` | Required JavaScript function expression, at most 20,000 characters |
| `linkSelector` | CSS selector for href-bearing elements on the live rendered page; default empty |
| `maxPages` | Global unique visited-page cap, default 3; range 1–20 |
| `maxDepth` | Start URLs are depth 0; default 0 disables recursion; range 0–3 |
| `maxItems` | Global saved-record cap, default 20; range 1–1000 |
| `pageTimeoutSecs` | Per-page wall-clock deadline, default 15; range 1–30 seconds |

Zero and unlimited limits are not supported. Unknown inputs fail validation instead of being silently ignored. Per-URL headers, payloads, userData and non-GET methods are rejected, even if a generic hosted editor exposes these controls. Recursion requires a nonempty `linkSelector` whenever depth is positive; at depth zero the selector is unused.

### Same-origin crawling

URLs are deduplicated globally after removing fragments. Query strings remain distinct. Crawling is breadth-first; each root establishes its own origin. Discovered links must match the origin exactly, including scheme and port. Cross-origin redirects fail rather than expanding scope.

Only the first 1,000 matching links per page are considered, and at most 1,000 requests can be queued. The run stops at the first global page or record cap. An oversized final record array is truncated to the remaining record limit; that truncation is intentional and does not assert complete upstream coverage.

### Output records

Each returned object becomes a default-dataset record:

| Field | Meaning |
| --- | --- |
| `sourceUrl` | Final public source URL |
| `depth` | Number of followed links from a start URL |
| `data` | Your custom JSON object, with your chosen keys |
| `scrapedAt` | UTC extraction timestamp |

Observed example from the default local run:

```json
{
  "sourceUrl": "https://example.com/",
  "depth": 0,
  "data": {"title": "Example Domain", "heading": null, "url": "https://example.com/"},
  "scrapedAt": "2026-10-08T06:06:20.358Z"
}
```

`data` is intentionally flexible. It is not a fixed product, article or contact schema. Return only the fields needed for your job. Dataset exports are available through Apify's normal JSON, CSV and spreadsheet export features; nested objects are most naturally consumed as JSON.

### Crawl summary

The key-value store's `SUMMARY` record reports visited pages, saved records, queued requests remaining, configured limits and guarded transfer bytes. It is written after successful completion. It is not a claim that every link or website record was collected.

A successful run returning no objects has an empty dataset and a summary. A failed run can retain earlier records, but those records are partial and should not be treated as complete results.

### How much does it cost to extract custom DOM records?

There are two event charges: one `start` fee per run and one `item` charge for each saved custom object. Pages returning null have no item charge. There is no separate page, recursion, proxy, API or AI charge from this Actor.

The configuration is $0.012 per start plus the following item prices:

| Apify spend tier | Per saved item |
| --- | --- |
| FREE | $0.00322 |
| BRONZE | $0.0028 |
| SILVER | $0.002184 |
| GOLD | $0.00168 |
| PLATINUM | $0.00168 |
| DIAMOND | $0.00168 |

BRONZE estimated total bills, including the start fee: 1 saved record = $0.0148; 10 saved records = $0.0400; 100 saved records = $0.2920. Spend tiers depend on qualifying aggregate monthly Store spend, not this Actor's per-run volume. Your platform billing settings may impose an additional run spending cap.

Charges and payout illustrations are estimates, not guaranteed invoices or earnings. Refunds, fraud, disputes, taxes, corrections and clawbacks can affect final amounts.

### Safety and runtime limits

Only public IPv4 HTTP(S) destinations on ports 80/443 are supported. URL credentials, private/reserved IP ranges, localhost and IPv6-only sites are rejected. DNS addresses are checked and pinned for outbound browser connections; redirects and subresources traverse the same guard.

Concurrency is one page. Images, media and fonts are blocked. A run has a 60-second browser-work deadline, a 2 MiB rendered-DOM cap, a 256 KiB serialized-output cap per page and a 20 MiB guarded-transfer cap. Deadlines include extraction: an infinite JavaScript loop terminates the browser and fails the run. No automatic paid proxy, retry or challenge bypass is enabled.

Logins, detected password forms and known challenge shapes fail. Detection is not a guarantee that every access restriction is recognized. You must choose sources you are authorized to access.

### Integrations

- Send custom records to a database or ETL pipeline using the dataset API.
- Connect Apify dataset exports to a spreadsheet for manual review.
- Use your own scheduled Apify runs to collect comparable public-page snapshots.
- Compare successive datasets in your own application; this Actor does not provide built-in change detection or alerts.

### API usage

Replace the token placeholder with your own Apify token. Never embed secrets in the JavaScript function.

```bash
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~public-web-custom-browser-page-function/runs' \
  -H 'Authorization: Bearer YOUR_APIFY_TOKEN' \
  -H 'Content-Type: application/json' \
  -d '{"startUrls":[{"url":"https://example.com"}],"pageFunction":"({document}) => ({title:document.title})","maxPages":1}'
```

```javascript
const response = await fetch('https://api.apify.com/v2/acts/automation-lab~public-web-custom-browser-page-function/runs', {
  method: 'POST',
  headers: { Authorization: `Bearer ${process.env.APIFY_TOKEN}`, 'Content-Type': 'application/json' },
  body: JSON.stringify({ startUrls: [{url:'https://example.com'}], pageFunction: '({document}) => ({title:document.title})', maxPages: 1 })
});
const {data: run} = await response.json();
console.log(run.id, run.defaultDatasetId);
```

```python
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/public-web-custom-browser-page-function').call(
    run_input={'startUrls': [{'url':'https://quotes.toscrape.com/'}],
               'pageFunction': '({document}) => ({title:document.title})', 'maxPages': 1}
)
print(run['id'], run['defaultDatasetId'])
```

Retain the run ID, poll that same run until terminal, then retrieve bounded pages from its default dataset. A start response is metadata, not extracted data.

### MCP setup

For Claude Code:

```bash
claude mcp add --transport http apify \
  'https://mcp.apify.com?tools=automation-lab/public-web-custom-browser-page-function'
```

For Claude Desktop, Cursor and VS Code, use the equivalent HTTP MCP configuration supported by your client:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=automation-lab/public-web-custom-browser-page-function"
    }
  }
}
```

Authentication is handled by the client and your Apify account. Discover actual tool names and schemas with scoped `tools/list` before invoking anything. Actor selection also exposes run/data utility tools; it is not exactly one tool.

Example prompt: "Extract the title from https://example.com with one page and a DOM pageFunction. Start once, retain the run ID, and return the resulting custom record."

For an authorized invocation, start once and retain the run ID/status/storage IDs. Server `waitSecs` is 0–45 seconds (default 30), not a completion guarantee. Set the client request timeout above the selected server wait plus transport margin, within a total 120-second deadline. Follow the same run using `get-actor-run` with bounded 2/4/8-second backoff (cap 10). After a client timeout, recover the known run instead of restarting; if its ID is unknown, report uncertainty rather than blindly rerun. At the deadline report pending with its ID and stop polling. Terminal failures mean partial results, not full success.

After success, use `get-dataset-items` with the returned dataset ID, explicit `limit: 20`, `offset: 0` and needed `fields` (for example `sourceUrl,depth,data,scrapedAt`, in the format exposed by tools/list). Advance the source offset by each untransformed page and follow pagination metadata, not filtered/transformed row counts. Admit at most 100 rows and 64 KiB of serialized UTF-8 content to model context across all pages, whichever comes first. Enforce the byte ceiling host-side even for an oversized row/page; disclose omissions, returned counts, continuation offsets and partial/truncated status. A client without response interception cannot claim that hard byte guarantee. Keep full exports outside model context. No numeric context/token-savings guarantee is made.

### Legality and data handling

Respect source terms, robots policies, permissions and applicable laws. Do not use this Actor to collect sensitive/private information, log in, evade challenges or probe internal networks. Public visibility alone does not establish permission for every reuse.

The Actor uses no AI model during extraction and sends no page data to an AI provider. Chromium visits your selected public sources, and Apify stores input, results and logs under your account's storage/retention settings. Browser contexts and cookies are ephemeral and discarded after each page. There is no cross-run cache. Delete runs, datasets and key-value stores through Apify when no longer required; this Actor does not override platform retention or backup policy.

Failed operations send sanitized diagnostic input, exceptions and actor/build/run IDs to our private GlitchTip service for repair. Secret fields and URL queries are removed, and reports are retained for 30 days. JavaScript source is input: **do not embed credentials, private identifiers or personal data in it**. Custom objects are stored exactly as returned, so choose fields carefully. User-function exception text is replaced with a generic failure to avoid copying extracted page content into logs.

This Actor is independently developed and is not affiliated with or endorsed by Microsoft, Playwright, Apify or the websites you visit. Apify's standard end-user terms apply; no additional custom terms are imposed.

### FAQ and troubleshooting

**Can I use server-side Playwright hooks?** No. The function receives a DOM snapshot, not `page`, `browser` or Node APIs.

**Why did `fetch` fail inside my function?** Network is deliberately disabled in the extraction context. Obtain values from the rendered DOM.

**Why is a heading null?** The selector did not match the snapshot. Check the source and choose a suitable CSS selector. Null values inside custom objects are allowed.

**Why did my run fail?** Check public-network eligibility, HTTP status, cross-origin redirects, challenge/login pages, time/size caps and JSON return shape. Syntax errors and oversized/function/BigInt/circular outputs fail.

**Are empty results charged?** The start fee still applies; there is no item charge unless an object is saved.

**Does it bypass bot protection?** No. Use an accessible public source; no challenge bypass or residential fallback is provided.

**Where can I get help?** Open an issue through this Actor's Apify Store Issues tab, including a safe minimal input and your run link. Never include credentials or sensitive data.

### Related automation

This is a standalone developer-configured utility. No fixed-source Actor is interchangeable with its arbitrary DOM contract, so no unrelated portfolio product is presented as a substitute.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/public-web-custom-browser-page-function/changelog.md

# Actor input Schema

## `startUrls` (type: `array`):

1–20 credential-free HTTP(S) URLs on public IPv4 destinations, ports 80/443 only. Deduplicated globally without fragments; redirects must stay on the original origin. GET only: entries support url and optional method GET. Headers, payload, userData and non-GET methods are rejected even if a generic editor exposes them.

## `pageFunction` (type: `string`):

JavaScript function expression receiving { document, url }. Runs in an isolated offline browser context over the rendered DOM snapshot, not a Playwright/Node hook. Return one JSON object, an array of objects, or null. Max 20,000 characters; 256 KiB serialized output per page.

## `linkSelector` (type: `string`):

Optional CSS selector for elements with href on the live rendered page. Required when maxDepth > 0; ignored at depth 0. Only same-origin public links are queued; up to 1,000 links per page and queued requests.

## `maxPages` (type: `integer`):

Global maximum unique visited pages across all start URLs and discovered links, before extraction. Default 3, range 1–20; zero/unlimited not supported.

## `maxDepth` (type: `integer`):

Start URLs are depth 0. Default 0 disables recursion; depths 1–3 require linkSelector. Breadth-first traversal is bounded by maxPages and maxItems.

## `maxItems` (type: `integer`):

Global maximum saved custom records, default 20 (1–1000). Excess records from the final page are truncated; null returns no records. Stops scheduling at the limit.

## `pageTimeoutSecs` (type: `integer`):

Wall-clock deadline for navigation, rendering, extraction and cleanup per page, including synchronous JavaScript loops. Default 15 seconds, range 1–30; exhaustion fails the run with any saved rows marked partial.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://quotes.toscrape.com/"
    }
  ],
  "pageFunction": "({ document, url }) => ({ title: document.title, heading: document.querySelector(\"h1\")?.textContent?.trim() ?? null, url })",
  "linkSelector": "",
  "maxPages": 3,
  "maxDepth": 0,
  "maxItems": 20,
  "pageTimeoutSecs": 15
}
```

# Actor output Schema

## `dataset` (type: `string`):

Each custom object in the default dataset.

## `summary` (type: `string`):

Visited pages, output counts and configured limits.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://quotes.toscrape.com/"
        }
    ],
    "pageFunction": ({ document, url }) => ({ title: document.title, heading: document.querySelector("h1")?.textContent?.trim() ?? null, url })
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/public-web-custom-browser-page-function").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://quotes.toscrape.com/" }],
    "pageFunction": "({ document, url }) => ({ title: document.title, heading: document.querySelector(\"h1\")?.textContent?.trim() ?? null, url })",
}

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/public-web-custom-browser-page-function").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://quotes.toscrape.com/"
    }
  ],
  "pageFunction": "({ document, url }) => ({ title: document.title, heading: document.querySelector(\\"h1\\")?.textContent?.trim() ?? null, url })"
}' |
apify call automation-lab/public-web-custom-browser-page-function --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/public-web-custom-browser-page-function"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9m5fd6AJzDDtZjYa9/builds/0h4ncqC9VwrJ6dvLn/openapi.json
