Turn web pages, or another Actor's dataset, into clean JSON with an LLM. Give a JSON schema or list the fields in plain English. Output is validated against your schema; pages that can't be fetched or don't match aren't charged. $5 per 1,000 pages plus the model's tokens.
Versions follow MAJOR.MINOR.PATCH (src/version.py); Apify shows MAJOR.MINOR from .actor/actor.json.
Every run logs its version and records it in the RUN_STATS key-value record.
About future changes in output: what a model reads off a page can differ between models, and between runs
of the same model. Every answer is checked against your schema, and one that doesn't match comes back with
valid: false and the reasons, uncharged.
1.0.0 (unreleased)
First release: web pages, or the pages in another actor's dataset, to JSON that matches your schema.
Input:urls, or datasetId (+ optional datasetUrlField) to read page URLs from any dataset and chain
after another actor with {{resource.defaultDatasetId}} (the shared reader of page-to-markdown and seven more
actors: url, pageUrl, link, website, loadedUrl found automatically; Google Maps links never used;
tracking parameters dropped; each row carries sourceTitle, sourcePlaceId, sourceIndex).
What to extract:fields in plain English (each becomes a camelCase key; a hint after a colon goes to the
model) or a JSON Schema in schema (overrides fields), plus optional instructions.
Reading pages: the shared guarded client (robots.txt, AI-crawler opt-outs, private-network guard, ports 80
and 443, HTML only, 5 MB cap), no browser: pages that need JavaScript are reported, not sent to the model.
pageContentfullPage (default) keeps headers, footers and sidebars, minus scripts, forms, hidden elements and
navigation menus; mainContent sends only the main content. The page's schema.org JSON-LD goes along.
maxTokensPerPage (default 8,000) is a hard cap: a longer page keeps its start and end, truncated: true.
The model: one call per page through Apify's OpenRouter proxy with the run's own token, so tokens are billed
to the account running the actor; model takes any OpenRouter id (default openai/gpt-4.1-mini). The schema
goes as a structured-output format, strict when it qualifies; a provider that refuses it in strict mode gets it
again as a hint. Every answer is validated here: valid, errors (at most 5), data kept for inspection.
A proxy that refuses the run (401/402/403) stops it at once, with the reason.
Charging: rows are free; the custom page-extracted event is charged once per page whose answer validated,
after the row is in the dataset. Failed, blocked and invalid pages are rows with the reason, never charged.
maxPages and the maximum cost per run are honoured before any page is fetched: a slot is claimed per page and
given back when the page doesn't validate.
Failure isolation: one page (or one bug on it) never affects another. A run where no page validated fails with
the first reason; the rows stay in the dataset.
Sources and terms (ACTOR-CHECKLIST section 0): the pages are the user's own URLs, read like page-to-markdown
reads them, AI opt-outs honoured. The model endpoint is Apify's own, documented for actors
(https://docs.apify.com/sdk/python/docs/guides/ai-agents , read 2026-09-29: "The token usage is billed against the
Apify account running the Actor, so no provider API key is required"; "The Actor authenticates with the proxy
using the APIFY_TOKEN that the platform injects into every run"). No credential or identity of ours is used.