# AI Schema Web Extractor: BYOK + Browser Fallback (`automa-flow/ai-schema-web-extractor`) Actor

Extract schema-validated JSON from URLs with your own LLM key. Starts on fast HTTP and renders JavaScript only when the page needs it, with source evidence for the fields.

- **URL**: https://apify.com/automa-flow/ai-schema-web-extractor.md
- **Developed by:** [Vadim Bezrukov](https://apify.com/automa-flow) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 http page extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## AI Schema Web Extractor: BYOK + Browser Fallback

You have a list of pages and a JSON Schema. Paste both, add your own OpenAI or Anthropic key, and this Actor returns one row per URL with schema-checked JSON. It fetches the page over HTTP first. It opens a browser only when that static page is too thin to extract, for example an empty app shell. Each non-empty field can carry a short quote from the page, and the Actor checks that the quote is actually there.

The default input is test mode. That run returns a labeled fixture for the sample schema. It does not fetch a page and it does not call a model, so you can see the row shape before you spend a key.

### Try a live batch

Turn **Test mode** off. Set **URLs**, keep or replace the **JSON Schema**, choose **openai** or **anthropic**, and paste **API key**. Leave **Rendering** on AUTO.

A useful first result is a row with `status` `SUCCESS` or `PARTIAL`, `data` matching your schema, and `_evidence_verified` true or false. `PARTIAL` means the JSON validated, but at least one quote was missing from the page text. `NO_MATCH` means the page was fetched and the model returned a valid empty result. Those are different from `FETCH_FAILED` and `EXTRACTION_FAILED`.

Send `data` from the `SUCCESS` and `PARTIAL` rows to your database, CRM or agent. For the next run, keep the same schema and pass the next batch of URLs. This Actor does not watch a page for changes. The reason to run it again is the next batch, or the same schema inside a workflow that already has new links.

An agent can call the direct MCP tool at `https://mcp.apify.com?tools=automa-flow/ai-schema-web-extractor`. A concrete ask is: "Extract name and price from these product URLs into the JSON Schema I provide, using my own model key, and return only schema-valid rows." Read `RUN_SUMMARY` before paging through the dataset. Anonymous hosted MCP can search for the Actor. Calling it needs your own Apify account.

### What you are charged

You pay the model provider directly. The Actor charges only after a schema-valid `SUCCESS` or `PARTIAL` row is stored.

| Event | When | Price |
| --- | --- | --- |
| `page-extracted-http` | Valid row after HTTP | $0.004 |
| `page-extracted-browser` | Valid row after browser fallback | $0.006 |

That is from $4 per 1,000 successful schema-valid pages, with platform usage included. Fetches that fail, blocked pages, `NO_MATCH`, invalid model output, retries, repair attempts and test mode are not charged. A repair that then validates is still one page event. A forced browser run fails before it fetches anything when the charge limit is below $0.006, and the status says the limit must cover one rendered page. Set `maxTotalChargeUsd` before an API or agent run: $0.004 for each HTTP page and $0.006 for each browser page you expect to keep.

### What a row contains

`source_url` is the URL you sent. `source_id` is stable for that normalized URL. `scraped_at` is when this run observed the page. `fetch_mode` is `HTTP` or `BROWSER`. `fallback_reason` says why AUTO opened a browser, or is null. `data` is the schema value, or null when extraction did not produce one. `_evidence` holds the quotes that were found in the page. `_attempts` counts model calls for that URL, including the single repair. `_input_chars` is how much cleaned text was sent. `_schema_version` is 1.

`fingerprint` hashes the status, the normalized URL and `data`, so a later run can be compared without treating `scraped_at` or a tracking parameter as a change.

### Limits

Process only pages you are authorized to access and use. This Actor extracts user-specified web content and does not grant rights to third-party content.

v1 accepts 1 to 500 explicit URLs. It does not crawl a site, submit forms, log in, solve CAPTCHA or rotate residential proxies. A login wall or challenge is `BLOCKED`, not a puzzle to get around. `robots.txt` is enforced for the shared fetcher user agent `AutomaFlowContentCrawler/1.0`, including a page-level `noai` directive.

HTTP requests pin DNS to a public address and refuse private, link-local and metadata targets, including redirects. The browser sends its traffic through a local proxy that does the same lookup and connects to that public address, so Chromium does not resolve those hosts itself.

The model allowlist is `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.6-sol`, `claude-haiku-4-5` and `claude-sonnet-5`. There is no `gpt-5.6-mini` id. Terra is the current mini-tier equivalent. The Actor does not accept a custom base URL.

Evidence is "this quote appears in the cleaned page", not a probability. The model can still be wrong when the quote is real but the field mapping is not. One repair runs when JSON or schema validation fails. A second failure is `EXTRACTION_FAILED` and is not billed.

Test mode fills a simple example for your schema and labels the row with `test_mode: true` and warning `TEST_MODE_FIXTURE`. If your schema is too constrained for that example, the row is `EXTRACTION_FAILED` and still does not touch the network.

### Statuses

`SUCCESS` and `PARTIAL` are the rows to keep. `NO_MATCH` is a valid empty extraction. `FETCH_FAILED` means the page was not usable (`HTTP_404`, `SOFT_404`, `INTERACTION_REQUIRED`, `HTTP_429`, `HTTP_5XX`, `TIMEOUT`, `UNSAFE_URL`, `CONTENT_EMPTY`, `CONTENT_TOO_LARGE`, `BROWSER_UNAVAILABLE`, `ROBOTS_UNAVAILABLE`). `SOFT_404` means the server returned a page, but the page says it was not found, so the model is not called. `INTERACTION_REQUIRED` means the visible text asks for a click or expansion, so the model is not called. `ROBOTS_UNAVAILABLE` means robots.txt could not be checked, so the page was not fetched. `BLOCKED` means a robots disallow, a challenge or a login wall. `EXTRACTION_FAILED` means the model key, an unpaid model account, the provider response or schema validation failed after the page was fetched. `SKIPPED` means the charge limit, an uncertain charge or the run clock stopped before that URL. One bad URL does not drop the rest of the batch. A bad API key or an unpaid model account stops later model calls. A refused or uncertain page charge skips the remaining URLs.

`RUN_SUMMARY.result` is `COMPLETE` when every unique URL returns `SUCCESS` or a verified `NO_MATCH`. Mixed results, unverified evidence or skipped work are `PARTIAL`, with the collected rows retained. `NO_EXTRACTABLE_PAGES` means every URL got a final answer about the page itself: `HTTP_404`, `SOFT_404`, `INTERACTION_REQUIRED`, `UNSAFE_URL`, `CONTENT_TOO_LARGE`, an unsupported content type, or a robots.txt or `noai` block. That run succeeds, because retrying would give the same rows. If no URL produced a usable row and at least one failed for another reason (a timeout, `HTTP_5XX`, `HTTP_429`, a challenge or login wall, an unreadable robots.txt, or a model error), the run is `FAILED`; its rows and summary remain available. Exhausted OpenAI credits and account limits stop further model calls, just like invalid credentials.

The input form rejects a URL that does not start with `http://` or `https://`, a blank URL and an unknown model id, so no run starts. A live run without an API key, an OpenAI key with a Claude model (or the reverse), or a schema keyword this Actor does not support (for example `if`, `not` or a remote `$ref`) fails at startup. The run status names the field, the code and the fix, for example `Invalid input. apiKey (API_KEY_REQUIRED): a live run needs your OpenAI or Anthropic key in apiKey.` Nothing is fetched or charged.

The provider follows your key and model. An Anthropic key (`sk-ant-...`) or a Claude model uses Anthropic even when `llmProvider` is left at `openai`, and an OpenAI key uses OpenAI. The run log notes the switch.

Local references such as `#/$defs/Price` or `#/definitions/Product`, the form Pydantic `model_json_schema()` and zod-to-json-schema export, are inlined before the first fetch. Recursive references are rejected at startup.

Browser HTTP errors are reported before extraction. A navigation failure affects that URL; later pages can still render. HTTP pages are decoded using their declared header or HTML charset, while browser text is kept as Unicode.

If something looks wrong, open an Apify Issue with the run ID, the expected field and a sanitized page. Do not paste the API key.

### Updates

See the [changelog](https://apify.com/automa-flow/ai-schema-web-extractor/changelog) for release history.

# Changelog

This Actor's version history is a separate document: https://apify.com/automa-flow/ai-schema-web-extractor/changelog.md

# Actor input Schema

## `urls` (type: `array`):

1 to 500 public http or https pages you are allowed to process. Exact normalized duplicates are dropped. Localhost, private networks and file downloads are reported per URL.

## `schema` (type: `object`):

JSON Schema for one object, or an array of objects. Draft 2020-12 and draft-07 only. Local $ref such as #/$defs/Price (Pydantic or Zod export) is inlined. Remote or recursive $ref and conditional keywords are rejected before any page is fetched.

## `instructions` (type: `string`):

Optional note for the model, such as which product on the page to extract. Webpage text is never treated as instructions.

## `llmProvider` (type: `string`):

Bring your own key. openai and anthropic are the supported providers. An Anthropic key (sk-ant-...) or a Claude model uses anthropic, and an OpenAI key uses openai, whatever this field says. The Actor does not sell model tokens.

## `apiKey` (type: `string`):

Your OpenAI or Anthropic key. Required when test mode is off, otherwise the run stops at startup with API\_KEY\_REQUIRED. Leave empty in test mode. The key is not stored, logged or written to the dataset.

## `model` (type: `string`):

Optional allowlisted model id. OpenAI defaults to gpt-5.6-terra. Anthropic defaults to claude-haiku-4-5. Arbitrary endpoint URLs are rejected.

## `renderingMode` (type: `string`):

AUTO fetches HTTP first and opens a browser only when the static page is too thin. HTTP\_ONLY never renders. BROWSER renders every URL once, still without login or CAPTCHA bypass. BROWSER stops before fetch unless the charge limit covers at least $0.006.

## `includeEvidence` (type: `boolean`):

Ask the model for a source quote per extracted field and check that the quote appears in the page text. This is a presence check, not a confidence score.

## `maxInputChars` (type: `integer`):

Maximum characters of cleaned page text sent to the model. Structured blocks are included first. The row records how many characters were sent.

## `testMode` (type: `boolean`):

On, the Actor returns a labeled non-live fixture for your schema and does not fetch pages or call a model. Turn this off for a live extraction.

## Actor input object example

```json
{
  "urls": [
    "https://example.com/product/widget"
  ],
  "schema": {
    "type": "object",
    "additionalProperties": false,
    "properties": {
      "name": {
        "type": "string"
      },
      "price": {
        "type": [
          "number",
          "null"
        ]
      },
      "currency": {
        "type": [
          "string",
          "null"
        ]
      }
    },
    "required": [
      "name"
    ]
  },
  "instructions": "Extract the main product shown on the page.",
  "llmProvider": "openai",
  "renderingMode": "AUTO",
  "includeEvidence": true,
  "maxInputChars": 40000,
  "testMode": true
}
```

# Actor output Schema

## `rows` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

## `billingReceipt` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com/product/widget"
    ],
    "schema": {
        "type": "object",
        "additionalProperties": false,
        "properties": {
            "name": {
                "type": "string"
            },
            "price": {
                "type": [
                    "number",
                    "null"
                ]
            },
            "currency": {
                "type": [
                    "string",
                    "null"
                ]
            }
        },
        "required": [
            "name"
        ]
    },
    "instructions": "Extract the main product shown on the page.",
    "llmProvider": "openai",
    "renderingMode": "AUTO",
    "includeEvidence": true,
    "maxInputChars": 40000,
    "testMode": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("automa-flow/ai-schema-web-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://example.com/product/widget"],
    "schema": {
        "type": "object",
        "additionalProperties": False,
        "properties": {
            "name": { "type": "string" },
            "price": { "type": [
                    "number",
                    "null",
                ] },
            "currency": { "type": [
                    "string",
                    "null",
                ] },
        },
        "required": ["name"],
    },
    "instructions": "Extract the main product shown on the page.",
    "llmProvider": "openai",
    "renderingMode": "AUTO",
    "includeEvidence": True,
    "maxInputChars": 40000,
    "testMode": True,
}

# Run the Actor and wait for it to finish
run = client.actor("automa-flow/ai-schema-web-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com/product/widget"
  ],
  "schema": {
    "type": "object",
    "additionalProperties": false,
    "properties": {
      "name": {
        "type": "string"
      },
      "price": {
        "type": [
          "number",
          "null"
        ]
      },
      "currency": {
        "type": [
          "string",
          "null"
        ]
      }
    },
    "required": [
      "name"
    ]
  },
  "instructions": "Extract the main product shown on the page.",
  "llmProvider": "openai",
  "renderingMode": "AUTO",
  "includeEvidence": true,
  "maxInputChars": 40000,
  "testMode": true
}' |
apify call automa-flow/ai-schema-web-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automa-flow/ai-schema-web-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qLdCkihYxCzDuJ1rE/builds/ulFxICH4PewmnynJd/openapi.json
