# Changelog of AI Schema Web Extractor: BYOK + Browser Fallback (`automa-flow/ai-schema-web-extractor`) Actor

- **URL**: https://apify.com/automa-flow/ai-schema-web-extractor/changelog.md
- **Full Actor documentation**: https://apify.com/automa-flow/ai-schema-web-extractor.md

## Changelog

### 2026-09-27: Fewer failed runs and clearer start errors

#### Breaking changes

- A run where every URL gets a final answer about the page itself now succeeds instead of failing. That covers a missing page (`HTTP_404`, `SOFT_404`), a page that needs a click, a robots.txt or `noai` block, an unsafe or unreadable address, and a file type or size this Actor does not read. `RUN_SUMMARY.result` is `NO_EXTRACTABLE_PAGES`, each row keeps its `error_code`, and nothing is charged. A run still fails when pages time out, return server errors or rate limits, show a challenge or login wall, or when the model key or account fails. If a webhook or script treated a failed run as "the page is gone", read `RUN_SUMMARY.result` instead.

#### Features

- Local `$ref` in your JSON Schema works. References such as `#/$defs/Price` or `#/definitions/Product`, as exported by Pydantic `model_json_schema()` and zod-to-json-schema, are inlined before the first page is fetched. Nested optional objects are accepted too. Remote and recursive references are still rejected at startup.

#### Improvements

- The model provider follows your key and model. An Anthropic key (`sk-ant-...`) or a Claude model uses Anthropic even when `llmProvider` is left at `openai`, and an OpenAI key uses OpenAI. The run log notes the switch. A key and a model from different providers still stop at startup, with a message that says which is which.

#### Fixes

- Start errors name the field, the code and the fix, for example `Invalid input. apiKey (API_KEY_REQUIRED): a live run needs your OpenAI or Anthropic key in apiKey.` A missing key, a provider mismatch or an unsupported schema previously showed only `value_error`.

### 2026-09-26: More reliable requests and spending checks

#### Fixes

- Spending checks and billing receipts use the effective event prices supplied for your run, including discounts and zero-priced events.
- Compressed HTTP responses are limited while they are decoded, so a small compressed download cannot expand past the response limit in memory.

### 2026-09-24: Schema-checked JSON from a URL batch (0.1)

#### Features

Send up to 500 URLs and a JSON Schema, and bring your own OpenAI or Anthropic key. Each URL becomes one Dataset row with a status, `data` that matches your schema, and short source quotes that the Actor found in the page text. Pages are fetched over HTTP first. A browser opens only when the static page is too thin to extract, for example an empty app shell. The default input is test mode: it returns one labeled fixture row without fetching a page or calling a model.

A `SUCCESS` or `PARTIAL` row is charged once: $0.004 after HTTP, $0.006 after the browser. `PARTIAL` means the JSON validated but at least one quote was not found on the page. `NO_MATCH`, fetch failures, blocked pages, invalid model output, retries and test mode are not charged. You pay the model provider directly.

A page that says it was not found, even with HTTP 200, is `FETCH_FAILED` / `SOFT_404`. A page whose visible text asks for a click or expansion is `FETCH_FAILED` / `INTERACTION_REQUIRED`. Neither is sent to the model. If robots.txt cannot be checked, the URL is `FETCH_FAILED` / `ROBOTS_UNAVAILABLE` and the page is not fetched. A robots disallow, a challenge or a login wall is `BLOCKED`. A URL query parameter that carries a token or another secret is masked in the Dataset row, and a username or password in the URL is removed. A URL the Actor cannot parse, including a non-numeric port or a broken address, fails that one row and the rest of the batch continues. The stored address is a neutral marker, so a password in the broken URL is not saved. That page is not fetched. A redirect is checked against robots.txt before the page is extracted, including after a browser navigation. An `X-Robots-Tag: noai` header blocks the page the same way a page-level `noai` directive does, including a JSON response. If the API key appears in the instructions, the schema or the page, the error names that place and the model is not called. A bad key, exhausted credits or an account limit stop later model calls. Temporary rate limits use bounded retries. The page title prefers the document title or the main heading when the page also lists cards, products or quotations. Pages in legacy character encodings keep their original text and quotes.

The input form rejects a URL that does not start with `http://` or `https://`, a blank URL and an unknown model id before a run starts. A live run without an API key, a model from the other provider, or a schema keyword this Actor does not support (for example `$ref`) fails at startup. The status names the field and the error code, and it does not repeat the rejected value. Nothing is fetched or charged. A forced browser run fails before fetching when the charge limit is below $0.006. When every URL in a batch fails, the run fails and its rows and `RUN_SUMMARY` stay available. A mixed batch keeps its completed rows and reports `PARTIAL`.

The default memory is 1024 MB. Set `maxTotalChargeUsd` before an API or agent run: $0.004 for each HTTP page and $0.006 for each browser page you expect to keep.

This is the first version, so there is no previous input to migrate.
