Playwright Scraper — Custom DOM Records
Pricing
from $1.68 / 1,000 item extracteds
Playwright Scraper — Custom DOM Records
Extract custom JSON records from rendered public pages with an isolated DOM page function. Bound pages, depth, time and output; optionally follow same-origin links. No Node hooks, login or challenge bypass.
Pricing
from $1.68 / 1,000 item extracteds
Rating
0.0
(0)
Developer
Automation Lab
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
This playwright scraper turns rendered public web pages into your own JSON records. Supply start URLs and a JavaScript DOM pageFunction; optionally follow same-origin links. It is for developers building repeatable extraction jobs, catalog sampling and public-page analysis without running server-side customer code.
Who is it for?
Developers defining custom DOM records for repeatable public-page extraction, data engineers feeding JSON pipelines, and analysts sampling publicly accessible catalogs. It is not a login automation or protected-site unlocker.
Why use this Actor?
- Keep extraction logic in a small reusable DOM function.
- Extract several objects per page, not just a page title or raw HTML dump.
- Bound pages, depth, records, time and result size.
- Evaluate your function in an isolated offline DOM snapshot, without Node hooks, page cookies or original site scripts.
- Use the same record envelope across unrelated public websites.
This is a deliberately narrower tool than a general Playwright code runner. It cannot log in, click through a challenge, execute Node hooks or access private networks. It does not promise access to every website.
Getting started
- Enter a public HTTP(S) start URL.
- Supply a JavaScript function expression such as the example below.
- Start with
maxPages: 1andmaxDepth: 0. - Inspect the
dataobject in the default dataset. - Add a same-origin
linkSelectorand increase depth only after checking the site structure.
{"startUrls": [{ "url": "https://quotes.toscrape.com/" }],"pageFunction": "({ document, url }) => ({ title: document.title, heading: document.querySelector('h1')?.textContent?.trim() ?? null, url })","maxPages": 1,"maxDepth": 0,"maxItems": 20,"pageTimeoutSecs": 15}
How the DOM page function works
Chromium renders the public page, waiting for network idle. The Actor then copies its HTML into a separate offline browser context. Scripts, frames, embedded objects, base tags, refresh metadata and inline event handlers are removed. Your function receives { document, url }, where url is the loaded source URL, not the snapshot's about:blank location.
Return a JSON object, an array of JSON objects, or null/undefined for no records. Each record must still be an object after JSON serialization: sparse arrays, Date records and toJSON conversions to primitives are rejected. The entire page output must fit 256 KiB before records are truncated to the remaining limit. Async functions are supported, but external network calls are blocked. Use new URL(relativeHref, url) to resolve relative links.
({ document, url }) => Array.from(document.querySelectorAll('.quote')).map(quote => ({text: quote.querySelector('.text')?.textContent,author: quote.querySelector('.author')?.textContent,url}))
The function does not receive Playwright's page, Node's require/process, Actor storage, tokens, browser cookies or a live authenticated session. Shadow DOM, canvas pixels, transient JavaScript state and interactions with the original page are not preserved in the HTML snapshot.
Inputs
| Input | Meaning |
|---|---|
startUrls | Required list of 1–20 public HTTP(S) URLs; strings or {url} objects; optional method: "GET" only |
pageFunction | Required JavaScript function expression, at most 20,000 characters |
linkSelector | CSS selector for href-bearing elements on the live rendered page; default empty |
maxPages | Global unique visited-page cap, default 3; range 1–20 |
maxDepth | Start URLs are depth 0; default 0 disables recursion; range 0–3 |
maxItems | Global saved-record cap, default 20; range 1–1000 |
pageTimeoutSecs | Per-page wall-clock deadline, default 15; range 1–30 seconds |
Zero and unlimited limits are not supported. Unknown inputs fail validation instead of being silently ignored. Per-URL headers, payloads, userData and non-GET methods are rejected, even if a generic hosted editor exposes these controls. Recursion requires a nonempty linkSelector whenever depth is positive; at depth zero the selector is unused.
Same-origin crawling
URLs are deduplicated globally after removing fragments. Query strings remain distinct. Crawling is breadth-first; each root establishes its own origin. Discovered links must match the origin exactly, including scheme and port. Cross-origin redirects fail rather than expanding scope.
Only the first 1,000 matching links per page are considered, and at most 1,000 requests can be queued. The run stops at the first global page or record cap. An oversized final record array is truncated to the remaining record limit; that truncation is intentional and does not assert complete upstream coverage.
Output records
Each returned object becomes a default-dataset record:
| Field | Meaning |
|---|---|
sourceUrl | Final public source URL |
depth | Number of followed links from a start URL |
data | Your custom JSON object, with your chosen keys |
scrapedAt | UTC extraction timestamp |
Observed example from the default local run:
{"sourceUrl": "https://example.com/","depth": 0,"data": {"title": "Example Domain", "heading": null, "url": "https://example.com/"},"scrapedAt": "2026-10-08T06:06:20.358Z"}
data is intentionally flexible. It is not a fixed product, article or contact schema. Return only the fields needed for your job. Dataset exports are available through Apify's normal JSON, CSV and spreadsheet export features; nested objects are most naturally consumed as JSON.
Crawl summary
The key-value store's SUMMARY record reports visited pages, saved records, queued requests remaining, configured limits and guarded transfer bytes. It is written after successful completion. It is not a claim that every link or website record was collected.
A successful run returning no objects has an empty dataset and a summary. A failed run can retain earlier records, but those records are partial and should not be treated as complete results.
How much does it cost to extract custom DOM records?
There are two event charges: one start fee per run and one item charge for each saved custom object. Pages returning null have no item charge. There is no separate page, recursion, proxy, API or AI charge from this Actor.
The configuration is $0.012 per start plus the following item prices:
| Apify spend tier | Per saved item |
|---|---|
| FREE | $0.00322 |
| BRONZE | $0.0028 |
| SILVER | $0.002184 |
| GOLD | $0.00168 |
| PLATINUM | $0.00168 |
| DIAMOND | $0.00168 |
BRONZE estimated total bills, including the start fee: 1 saved record = $0.0148; 10 saved records = $0.0400; 100 saved records = $0.2920. Spend tiers depend on qualifying aggregate monthly Store spend, not this Actor's per-run volume. Your platform billing settings may impose an additional run spending cap.
Charges and payout illustrations are estimates, not guaranteed invoices or earnings. Refunds, fraud, disputes, taxes, corrections and clawbacks can affect final amounts.
Safety and runtime limits
Only public IPv4 HTTP(S) destinations on ports 80/443 are supported. URL credentials, private/reserved IP ranges, localhost and IPv6-only sites are rejected. DNS addresses are checked and pinned for outbound browser connections; redirects and subresources traverse the same guard.
Concurrency is one page. Images, media and fonts are blocked. A run has a 60-second browser-work deadline, a 2 MiB rendered-DOM cap, a 256 KiB serialized-output cap per page and a 20 MiB guarded-transfer cap. Deadlines include extraction: an infinite JavaScript loop terminates the browser and fails the run. No automatic paid proxy, retry or challenge bypass is enabled.
Logins, detected password forms and known challenge shapes fail. Detection is not a guarantee that every access restriction is recognized. You must choose sources you are authorized to access.
Integrations
- Send custom records to a database or ETL pipeline using the dataset API.
- Connect Apify dataset exports to a spreadsheet for manual review.
- Use your own scheduled Apify runs to collect comparable public-page snapshots.
- Compare successive datasets in your own application; this Actor does not provide built-in change detection or alerts.
API usage
Replace the token placeholder with your own Apify token. Never embed secrets in the JavaScript function.
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~public-web-custom-browser-page-function/runs' \-H 'Authorization: Bearer YOUR_APIFY_TOKEN' \-H 'Content-Type: application/json' \-d '{"startUrls":[{"url":"https://example.com"}],"pageFunction":"({document}) => ({title:document.title})","maxPages":1}'
const response = await fetch('https://api.apify.com/v2/acts/automation-lab~public-web-custom-browser-page-function/runs', {method: 'POST',headers: { Authorization: `Bearer ${process.env.APIFY_TOKEN}`, 'Content-Type': 'application/json' },body: JSON.stringify({ startUrls: [{url:'https://example.com'}], pageFunction: '({document}) => ({title:document.title})', maxPages: 1 })});const {data: run} = await response.json();console.log(run.id, run.defaultDatasetId);
import osfrom apify_client import ApifyClientclient = ApifyClient(os.environ['APIFY_TOKEN'])run = client.actor('automation-lab/public-web-custom-browser-page-function').call(run_input={'startUrls': [{'url':'https://quotes.toscrape.com/'}],'pageFunction': '({document}) => ({title:document.title})', 'maxPages': 1})print(run['id'], run['defaultDatasetId'])
Retain the run ID, poll that same run until terminal, then retrieve bounded pages from its default dataset. A start response is metadata, not extracted data.
MCP setup
For Claude Code:
claude mcp add --transport http apify \'https://mcp.apify.com?tools=automation-lab/public-web-custom-browser-page-function'
For Claude Desktop, Cursor and VS Code, use the equivalent HTTP MCP configuration supported by your client:
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=automation-lab/public-web-custom-browser-page-function"}}}
Authentication is handled by the client and your Apify account. Discover actual tool names and schemas with scoped tools/list before invoking anything. Actor selection also exposes run/data utility tools; it is not exactly one tool.
Example prompt: "Extract the title from https://example.com with one page and a DOM pageFunction. Start once, retain the run ID, and return the resulting custom record."
For an authorized invocation, start once and retain the run ID/status/storage IDs. Server waitSecs is 0–45 seconds (default 30), not a completion guarantee. Set the client request timeout above the selected server wait plus transport margin, within a total 120-second deadline. Follow the same run using get-actor-run with bounded 2/4/8-second backoff (cap 10). After a client timeout, recover the known run instead of restarting; if its ID is unknown, report uncertainty rather than blindly rerun. At the deadline report pending with its ID and stop polling. Terminal failures mean partial results, not full success.
After success, use get-dataset-items with the returned dataset ID, explicit limit: 20, offset: 0 and needed fields (for example sourceUrl,depth,data,scrapedAt, in the format exposed by tools/list). Advance the source offset by each untransformed page and follow pagination metadata, not filtered/transformed row counts. Admit at most 100 rows and 64 KiB of serialized UTF-8 content to model context across all pages, whichever comes first. Enforce the byte ceiling host-side even for an oversized row/page; disclose omissions, returned counts, continuation offsets and partial/truncated status. A client without response interception cannot claim that hard byte guarantee. Keep full exports outside model context. No numeric context/token-savings guarantee is made.
Legality and data handling
Respect source terms, robots policies, permissions and applicable laws. Do not use this Actor to collect sensitive/private information, log in, evade challenges or probe internal networks. Public visibility alone does not establish permission for every reuse.
The Actor uses no AI model during extraction and sends no page data to an AI provider. Chromium visits your selected public sources, and Apify stores input, results and logs under your account's storage/retention settings. Browser contexts and cookies are ephemeral and discarded after each page. There is no cross-run cache. Delete runs, datasets and key-value stores through Apify when no longer required; this Actor does not override platform retention or backup policy.
Failed operations send sanitized diagnostic input, exceptions and actor/build/run IDs to our private GlitchTip service for repair. Secret fields and URL queries are removed, and reports are retained for 30 days. JavaScript source is input: do not embed credentials, private identifiers or personal data in it. Custom objects are stored exactly as returned, so choose fields carefully. User-function exception text is replaced with a generic failure to avoid copying extracted page content into logs.
This Actor is independently developed and is not affiliated with or endorsed by Microsoft, Playwright, Apify or the websites you visit. Apify's standard end-user terms apply; no additional custom terms are imposed.
FAQ and troubleshooting
Can I use server-side Playwright hooks? No. The function receives a DOM snapshot, not page, browser or Node APIs.
Why did fetch fail inside my function? Network is deliberately disabled in the extraction context. Obtain values from the rendered DOM.
Why is a heading null? The selector did not match the snapshot. Check the source and choose a suitable CSS selector. Null values inside custom objects are allowed.
Why did my run fail? Check public-network eligibility, HTTP status, cross-origin redirects, challenge/login pages, time/size caps and JSON return shape. Syntax errors and oversized/function/BigInt/circular outputs fail.
Are empty results charged? The start fee still applies; there is no item charge unless an object is saved.
Does it bypass bot protection? No. Use an accessible public source; no challenge bypass or residential fallback is provided.
Where can I get help? Open an issue through this Actor's Apify Store Issues tab, including a safe minimal input and your run link. Never include credentials or sensitive data.
Related automation
This is a standalone developer-configured utility. No fixed-source Actor is interchangeable with its arbitrary DOM contract, so no unrelated portfolio product is presented as a substitute.