# Changelog of AI Data Extractor: Web Pages to Structured JSON (`humble-echidna/ai-extract`) Actor

- **URL**: https://apify.com/humble-echidna/ai-extract/changelog.md
- **Full Actor documentation**: https://apify.com/humble-echidna/ai-extract.md

## Changelog

Versions follow MAJOR.MINOR.PATCH (`src/version.py`); Apify shows MAJOR.MINOR from `.actor/actor.json`.
Every run logs its version and records it in the `RUN_STATS` key-value record.

> **About future changes in output:** what a model reads off a page can differ between models, and between runs
> of the same model. Every answer is checked against your schema, and one that doesn't match comes back with
> `valid: false` and the reasons, uncharged.

### 1.0.0 (unreleased)

First release: web pages, or the pages in another actor's dataset, to JSON that matches your schema.

- **Input:** `urls`, or `datasetId` (+ optional `datasetUrlField`) to read page URLs from any dataset and chain
  after another actor with `{{resource.defaultDatasetId}}` (the shared reader of page-to-markdown and seven more
  actors: `url`, `pageUrl`, `link`, `website`, `loadedUrl` found automatically; Google Maps links never used;
  tracking parameters dropped; each row carries `sourceTitle`, `sourcePlaceId`, `sourceIndex`).
- **What to extract:** `fields` in plain English (each becomes a camelCase key; a hint after a colon goes to the
  model) or a JSON Schema in `schema` (overrides `fields`), plus optional `instructions`.
- **Reading pages:** the shared guarded client (robots.txt, AI-crawler opt-outs, private-network guard, ports 80
  and 443, HTML only, 5 MB cap), no browser: pages that need JavaScript are reported, not sent to the model.
  `pageContent` `fullPage` (default) keeps headers, footers and sidebars, minus scripts, forms, hidden elements and
  navigation menus; `mainContent` sends only the main content. The page's schema.org JSON-LD goes along.
  `maxTokensPerPage` (default 8,000) is a hard cap: a longer page keeps its start and end, `truncated: true`.
- **The model:** one call per page through Apify's OpenRouter proxy with the run's own token, so tokens are billed
  to the account running the actor; `model` takes any OpenRouter id (default `openai/gpt-4.1-mini`). The schema
  goes as a structured-output format, strict when it qualifies; a provider that refuses it in strict mode gets it
  again as a hint. Every answer is validated here: `valid`, `errors` (at most 5), `data` kept for inspection.
  A proxy that refuses the run (401/402/403) stops it at once, with the reason.
- **Charging:** rows are free; the custom `page-extracted` event is charged once per page whose answer validated,
  after the row is in the dataset. Failed, blocked and invalid pages are rows with the reason, never charged.
  `maxPages` and the maximum cost per run are honoured before any page is fetched: a slot is claimed per page and
  given back when the page doesn't validate.
- Failure isolation: one page (or one bug on it) never affects another. A run where no page validated fails with
  the first reason; the rows stay in the dataset.
- **Sources and terms (ACTOR-CHECKLIST section 0):** the pages are the user's own URLs, read like page-to-markdown
  reads them, AI opt-outs honoured. The model endpoint is Apify's own, documented for actors
  (https://docs.apify.com/sdk/python/docs/guides/ai-agents, read 2026-09-29: "The token usage is billed against the
  Apify account running the Actor, so no provider API key is required"; "The Actor authenticates with the proxy
  using the APIFY\_TOKEN that the platform injects into every run"). No credential or identity of ours is used.
