AI Schema Web Extractor: BYOK + Browser Fallback
Pricing
from $4.00 / 1,000 http page extracteds
AI Schema Web Extractor: BYOK + Browser Fallback
Extract schema-validated JSON from URLs with your own LLM key. Starts on fast HTTP and renders JavaScript only when the page needs it, with source evidence for the fields.
Pricing
from $4.00 / 1,000 http page extracteds
Rating
0.0
(0)
Developer
Vadim Bezrukov
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
You have a list of pages and a JSON Schema. Paste both, add your own OpenAI or Anthropic key, and this Actor returns one row per URL with schema-checked JSON. It fetches the page over HTTP first. It opens a browser only when that static page is too thin to extract, for example an empty app shell. Each non-empty field can carry a short quote from the page, and the Actor checks that the quote is actually there.
The default input is test mode. That run returns a labeled fixture for the sample schema. It does not fetch a page and it does not call a model, so you can see the row shape before you spend a key.
Try a live batch
Turn Test mode off. Set URLs, keep or replace the JSON Schema, choose openai or anthropic, and paste API key. Leave Rendering on AUTO.
A useful first result is a row with status SUCCESS or PARTIAL, data matching your schema, and _evidence_verified true or false. PARTIAL means the JSON validated, but at least one quote was missing from the page text. NO_MATCH means the page was fetched and the model returned a valid empty result. Those are different from FETCH_FAILED and EXTRACTION_FAILED.
Send data from the SUCCESS and PARTIAL rows to your database, CRM or agent. For the next run, keep the same schema and pass the next batch of URLs. This Actor does not watch a page for changes. The reason to run it again is the next batch, or the same schema inside a workflow that already has new links.
An agent can call the direct MCP tool at https://mcp.apify.com?tools=automa-flow/ai-schema-web-extractor. A concrete ask is: "Extract name and price from these product URLs into the JSON Schema I provide, using my own model key, and return only schema-valid rows." Read RUN_SUMMARY before paging through the dataset. Anonymous hosted MCP can search for the Actor. Calling it needs your own Apify account.
What you are charged
You pay the model provider directly. The Actor charges only after a schema-valid SUCCESS or PARTIAL row is stored.
| Event | When | Price |
|---|---|---|
page-extracted-http | Valid row after HTTP | $0.004 |
page-extracted-browser | Valid row after browser fallback | $0.006 |
That is from $4 per 1,000 successful schema-valid pages, with platform usage included. Fetches that fail, blocked pages, NO_MATCH, invalid model output, retries, repair attempts and test mode are not charged. A repair that then validates is still one page event. A forced browser run fails before it fetches anything when the charge limit is below $0.006, and the status says the limit must cover one rendered page. Set maxTotalChargeUsd before an API or agent run: $0.004 for each HTTP page and $0.006 for each browser page you expect to keep.
What a row contains
source_url is the URL you sent. source_id is stable for that normalized URL. scraped_at is when this run observed the page. fetch_mode is HTTP or BROWSER. fallback_reason says why AUTO opened a browser, or is null. data is the schema value, or null when extraction did not produce one. _evidence holds the quotes that were found in the page. _attempts counts model calls for that URL, including the single repair. _input_chars is how much cleaned text was sent. _schema_version is 1.
fingerprint hashes the status, the normalized URL and data, so a later run can be compared without treating scraped_at or a tracking parameter as a change.
Limits
Process only pages you are authorized to access and use. This Actor extracts user-specified web content and does not grant rights to third-party content.
v1 accepts 1 to 500 explicit URLs. It does not crawl a site, submit forms, log in, solve CAPTCHA or rotate residential proxies. A login wall or challenge is BLOCKED, not a puzzle to get around. robots.txt is enforced for the shared fetcher user agent AutomaFlowContentCrawler/1.0, including a page-level noai directive.
HTTP requests pin DNS to a public address and refuse private, link-local and metadata targets, including redirects. The browser sends its traffic through a local proxy that does the same lookup and connects to that public address, so Chromium does not resolve those hosts itself.
The model allowlist is gpt-5.6-terra, gpt-5.6-luna, gpt-5.6-sol, claude-haiku-4-5 and claude-sonnet-5. There is no gpt-5.6-mini id. Terra is the current mini-tier equivalent. The Actor does not accept a custom base URL.
Evidence is "this quote appears in the cleaned page", not a probability. The model can still be wrong when the quote is real but the field mapping is not. One repair runs when JSON or schema validation fails. A second failure is EXTRACTION_FAILED and is not billed.
Test mode fills a simple example for your schema and labels the row with test_mode: true and warning TEST_MODE_FIXTURE. If your schema is too constrained for that example, the row is EXTRACTION_FAILED and still does not touch the network.
Statuses
SUCCESS and PARTIAL are the rows to keep. NO_MATCH is a valid empty extraction. FETCH_FAILED means the page was not usable (HTTP_404, SOFT_404, INTERACTION_REQUIRED, HTTP_429, HTTP_5XX, TIMEOUT, UNSAFE_URL, CONTENT_EMPTY, CONTENT_TOO_LARGE, BROWSER_UNAVAILABLE, ROBOTS_UNAVAILABLE). SOFT_404 means the server returned a page, but the page says it was not found, so the model is not called. INTERACTION_REQUIRED means the visible text asks for a click or expansion, so the model is not called. ROBOTS_UNAVAILABLE means robots.txt could not be checked, so the page was not fetched. BLOCKED means a robots disallow, a challenge or a login wall. EXTRACTION_FAILED means the model key, an unpaid model account, the provider response or schema validation failed after the page was fetched. SKIPPED means the charge limit, an uncertain charge or the run clock stopped before that URL. One bad URL does not drop the rest of the batch. A bad API key or an unpaid model account stops later model calls. A refused or uncertain page charge skips the remaining URLs.
RUN_SUMMARY.result is COMPLETE when every unique URL returns SUCCESS or a verified NO_MATCH. Mixed results, unverified evidence or skipped work are PARTIAL, with the collected rows retained. NO_EXTRACTABLE_PAGES means every URL got a final answer about the page itself: HTTP_404, SOFT_404, INTERACTION_REQUIRED, UNSAFE_URL, CONTENT_TOO_LARGE, an unsupported content type, or a robots.txt or noai block. That run succeeds, because retrying would give the same rows. If no URL produced a usable row and at least one failed for another reason (a timeout, HTTP_5XX, HTTP_429, a challenge or login wall, an unreadable robots.txt, or a model error), the run is FAILED; its rows and summary remain available. Exhausted OpenAI credits and account limits stop further model calls, just like invalid credentials.
The input form rejects a URL that does not start with http:// or https://, a blank URL and an unknown model id, so no run starts. A live run without an API key, an OpenAI key with a Claude model (or the reverse), or a schema keyword this Actor does not support (for example if, not or a remote $ref) fails at startup. The run status names the field, the code and the fix, for example Invalid input. apiKey (API_KEY_REQUIRED): a live run needs your OpenAI or Anthropic key in apiKey. Nothing is fetched or charged.
The provider follows your key and model. An Anthropic key (sk-ant-...) or a Claude model uses Anthropic even when llmProvider is left at openai, and an OpenAI key uses OpenAI. The run log notes the switch.
Local references such as #/$defs/Price or #/definitions/Product, the form Pydantic model_json_schema() and zod-to-json-schema export, are inlined before the first fetch. Recursive references are rejected at startup.
Browser HTTP errors are reported before extraction. A navigation failure affects that URL; later pages can still render. HTTP pages are decoded using their declared header or HTML charset, while browser text is kept as Unicode.
If something looks wrong, open an Apify Issue with the run ID, the expected field and a sanitized page. Do not paste the API key.
Updates
See the changelog for release history.