# AI Data Extractor: Web Pages to Structured JSON (`humble-echidna/ai-extract`) Actor

Turn web pages, or another Actor's dataset, into clean JSON with an LLM. Give a JSON schema or list the fields in plain English. Output is validated against your schema; pages that can't be fetched or don't match aren't charged. $5 per 1,000 pages plus the model's tokens.

- **URL**: https://apify.com/humble-echidna/ai-extract.md
- **Developed by:** [Michael Costa](https://apify.com/humble-echidna) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 page extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does AI Data Extractor: Web Pages to Structured JSON do?

**AI Data Extractor** turns **web pages into structured JSON** with an LLM. Give it page URLs or another actor's
dataset, and say what you want: **plain-English fields** ("product name, price, in stock") or a **JSON Schema**.
Every answer is **validated against your schema**, and you pay only for pages that match.

It is **not** a crawler: it reads the pages you give it, one call per page. To get the pages first, chain it after
any scraper or crawler ([how](#chain-it-after-google-maps-or-any-scraper)).

[Input fields](https://apify.com/humble-echidna/ai-extract/input-schema) ·
[API](https://apify.com/humble-echidna/ai-extract/api) ·
[Use it from Claude, ChatGPT or Cursor (MCP)](#can-i-use-ai-data-extractor-from-an-ai-agent-mcp)

**Jump to:** [Fields](#what-data-does-ai-data-extractor-return) ·
[Price](#how-much-does-it-cost-to-extract-data-from-web-pages-with-ai) ·
[How to use](#how-to-extract-structured-data-from-web-pages) ·
[Chain after Google Maps](#chain-it-after-google-maps-or-any-scraper) ·
[Input](#input) · [Output](#output) ·
[AI agents (MCP)](#can-i-use-ai-data-extractor-from-an-ai-agent-mcp) · [Limits](#limits) · [FAQ](#faq)

**Try it in one click:** the input comes pre-filled with one book page from books.toscrape.com (a sandbox site made
for scraping) and six fields. That's 1 page: $0.005, plus the $0.00005 start fee, plus about $0.0005 of model tokens
on a paid Apify plan. **Then replace it with your own pages and fields.**

### What data does AI Data Extractor return?

One row per page you give it, including the pages that failed (with the reason, never charged).

| Field | Example | Notes |
|---|---|---|
| `data` | `{"title": "How Music Works", "price": 37.32, "currency": "GBP", "inStock": true}` | The extracted record: your fields as camelCase keys, or your schema's shape. |
| `valid` | `true` | `true` when `data` matches your schema. Only these pages are charged. |
| `errors` | `["price: 'cheap' is not of type 'number'"]` | Why a page has no valid data: blocked, needs JavaScript, HTTP error, or what didn't match. Empty when valid. |
| `url` | `https://books.toscrape.com/catalogue/how-music-works_979/index.html` | The page as you gave it. |
| `loadedUrl` | same | Where it ended up after redirects; `null` if it couldn't be fetched. |
| `title` | `How Music Works \| Books to Scrape - Sandbox` | The page title. |
| `model` | `openai/gpt-4.1-mini` | The model that read the page. |
| `inputTokens`, `outputTokens` | `1629`, `69` | Tokens the model reported: what the token charge is based on. |
| `approxTokens`, `truncated` | `812`, `false` | Page text sent (about 4 characters a token); `true` if the page was cut to fit `maxTokensPerPage`. |
| `sourceTitle`, `sourcePlaceId`, `sourceIndex` | `"Lamp Shop"`, `"ChIJ..."`, `0` | With a dataset as input: the item the URL came from, to join rows back to your leads. `null` for typed-in URLs. |

The full list is under [Output](#output).

### How much does it cost to extract data from web pages with AI?

Two parts, both on your Apify account:

1. **This actor: $5.00 per 1,000 pages extracted** ($0.005 a page), plus $0.00005 each time a run starts. Only pages
   whose answer matched your schema count.
2. **The model's tokens**, billed by Apify's own [OpenRouter actor](https://apify.com/apify/openrouter), which this
   actor calls with your run's token. On paid Apify plans that's OpenRouter's list price; on the Free plan the
   OpenRouter actor charges 10 times as much (its pricing, not ours). With the default model, `openai/gpt-4.1-mini`,
   ten book pages took 1,790 tokens a page on average: **$0.84 per 1,000 pages** on a paid plan ($8.45 on Free).
   Longer pages cost more; `maxTokensPerPage` caps it.

- **The example below:** 10 pages × $0.005 = $0.05, plus the start fee and about $0.0085 of tokens (paid plan).
- **A month, for example:** 2,000 leads from a Google Maps scraper, each business website read once for services and
  contacts: 2,000 × $0.005 = $10, plus roughly $2-4 of tokens (business home pages run larger than book pages).
- **Caps:** **Max pages per run** in the input, and **Maximum cost per run** in the run options (it covers this
  actor's charges; the token charges are the OpenRouter actor's, so cap those with **Max tokens per page** and
  **Max pages per run**). The run stops cleanly at whichever comes first. Each page is counted against the limit
  before it's fetched, and the count is given back if the page fails, so a capped run never reads a page it can't
  return.

**Never charged by this actor:** pages that fail, need JavaScript, are disallowed by robots.txt or opted out of AI
use, and answers that don't match your schema. (A page that reached the model still used tokens.)

### How to extract structured data from web pages

1. Open AI Data Extractor and click **Try for free** (or **Start** if you're signed in).
2. Put your pages in **Web page URLs**, one per line, or pick a dataset in **Or: page URLs from a dataset**.
3. Say what to extract: list the fields in **Fields to extract** (`product name, price, currency, in stock`), or
   paste a JSON Schema into **Or: JSON Schema for one record**. Add **Extra instructions** if the pages need
   explaining ("prices are in EUR unless the page says otherwise").
4. Click **Start**, then open the **Output** tab: `data` holds each page's record. Export as JSON, CSV or Excel.

#### Plain-English fields or a JSON Schema?

- **Fields** are quickest: each becomes a camelCase key (`in stock` → `inStock`) that holds text, a number,
  true/false, a list of texts or `null`. Add a hint after a colon: `price: the sale price, not the list price`.
  One field per line, or comma-separated on one line; at most 50.
- **A JSON Schema** gives you exact types, nesting and required fields, e.g. a list of objects per page. Answers
  that don't match (a price as text, a missing required field) come back with `valid: false`, the model's answer in
  `data` and the reasons in `errors`, and aren't charged.

### Chain it after Google Maps or any scraper

Point it at another actor's results and each item's page URL is read like a line of **Web page URLs**. The field is
found automatically (`url`, `pageUrl`, `link`, `website`, `loadedUrl`, then the same names one level down); a Google
Maps link is never used, so after a Maps scraper it reads each business's own `website`.

- **Once:** pick the dataset in **Or: page URLs from a dataset** (`datasetId`) and click **Start**.
- **Every time the other actor finishes:** on that actor (or its task), open **Integrations**, add **Run an actor
  or task**, pick AI Data Extractor, and set the input to include `"datasetId": "{{resource.defaultDatasetId}}"`
  plus your `fields`.

Each row carries `sourceTitle`, `sourcePlaceId` and `sourceIndex` from its item, so you can join the extracted
services and contacts back to your leads. It reads up to 20,000 items and 10,000 distinct URLs per run.

### Example: ten book pages

Input (a real run on Apify, 2026-09-30; nine book pages and the site's home page):

```json
{
  "urls": ["https://books.toscrape.com/catalogue/how-music-works_979/index.html", "..."],
  "fields": "title, price, currency, in stock, number available, UPC, product description (first sentence)"
}
```

All 10 pages validated in 11 seconds. One row, unshortened:

```json
{
  "url": "https://books.toscrape.com/catalogue/how-music-works_979/index.html",
  "loadedUrl": "https://books.toscrape.com/catalogue/how-music-works_979/index.html",
  "title": "How Music Works | Books to Scrape - Sandbox",
  "data": {
    "title": "How Music Works",
    "price": 37.32,
    "currency": "GBP",
    "inStock": true,
    "numberAvailable": 19,
    "upc": "327f68a59745c102",
    "productDescriptionFirstSentence": "How Music Works is David Byrne’s remarkable and buoyant celebration of a subject he has spent a lifetime thinking about."
  },
  "valid": true,
  "errors": [],
  "model": "openai/gpt-4.1-mini",
  "approxTokens": 812,
  "inputTokens": 1629,
  "outputTokens": 69,
  "truncated": false,
  "fetchedAt": "2026-09-30T03:01:14.281857Z",
  "sourceTitle": null,
  "sourcePlaceId": null,
  "sourceIndex": null
}
```

The home page, which isn't a product, came back with `price`, `currency` and `upc` as `null`: the model is told to
use `null` for what a page doesn't say, never to guess. Cost: $0.05 for the pages plus $0.00845 of tokens (845
OpenRouter token events at the paid-plan price).

### Ready-to-run examples

Each example opens with the input already filled in. Run it as it is, or change the input first.

- **[Extract product name, price and stock from product pages](https://apify.com/humble-echidna/ai-extract/examples/product-pages-to-json)**: Turn product pages into clean JSON: name, price, currency, stock and UPC, validated against the fields you list.
- **[Pull services and contact details from company websites](https://apify.com/humble-echidna/ai-extract/examples/company-websites-to-contacts)**: Read company websites and get the services they offer plus email, phone and address as JSON.

### Input

| Field | What it does |
|---|---|
| **Web page URLs** (`urls`) | The pages, one per line. Ignored when a dataset is set. |
| Or: page URLs from a dataset (`datasetId`) | Another actor's results; chain with `{{resource.defaultDatasetId}}`. |
| Field with the page URL (`datasetUrlField`) | Only with a dataset; empty = found automatically. |
| **Fields to extract** (`fields`) | Plain-English fields. Ignored when a schema is set. |
| Or: JSON Schema for one record (`schema`) | A JSON Schema with `"type": "object"` at the top. |
| Extra instructions (`instructions`) | Guidance for the model, up to 4,000 characters. |
| Model (`model`) | Any OpenRouter model id with structured outputs. Default `openai/gpt-4.1-mini`. |
| What the model reads (`pageContent`) | `fullPage` (default): the whole visible page minus scripts, forms and menus, so prices and footer contacts are in. `mainContent`: the article only, fewer tokens. |
| Max tokens per page (`maxTokensPerPage`) | Hard cap on the page text sent, default 8,000 (500-100,000). A longer page keeps its start and end. |
| Max pages per run (`maxPages`) | Stop after this many valid pages (1-10,000); empty = no limit. |

```json
{
  "urls": ["https://example.com/products/lamp", "https://example.com/products/chair"],
  "schema": {
    "type": "object",
    "properties": {
      "name": {"type": "string"},
      "price": {"type": "number"},
      "currency": {"type": "string"},
      "variants": {"type": "array", "items": {"type": "object", "properties": {"color": {"type": "string"}, "inStock": {"type": "boolean"}}}}
    },
    "required": ["name", "price"]
  },
  "model": "openai/gpt-4.1-mini",
  "maxPages": 100
}
```

**Which model?** The default, `openai/gpt-4.1-mini`, is accurate and cheap for extraction. For big batches of
simple pages, `openai/gpt-4.1-nano` or `google/gemini-2.5-flash-lite` cost about a quarter as much; for messy pages
or deep schemas, `openai/gpt-5.4-mini`. The model must support structured outputs.

### Output

One row per page, in the order they finish:

```json
{
  "url": "https://shop.example.com/mystery-box",
  "loadedUrl": "https://shop.example.com/mystery-box",
  "title": "Mystery Box | Example Shop",
  "data": {"name": "Mystery Box", "price": "cheap"},
  "valid": false,
  "errors": ["price: 'cheap' is not of type 'number'"],
  "model": "openai/gpt-4.1-mini",
  "approxTokens": 640,
  "inputTokens": 1180,
  "outputTokens": 18,
  "truncated": false,
  "fetchedAt": "2026-09-30T03:01:14Z",
  "sourceTitle": null,
  "sourcePlaceId": null,
  "sourceIndex": null
}
```

A page that couldn't be read has `data: null`, `loadedUrl: null` and the reason in `errors` (for example
`blocked by robots.txt` or `the page builds its content with JavaScript`). The run's `RUN_STATS` record counts pages
extracted, invalid and failed, tokens used, and the options that applied. The **Tokens** view of the dataset shows
tokens per page.

### Run it on a schedule, or from your own code

1. Save your input as a [task](https://docs.apify.com/platform/actors/running/tasks) (**Save as a new task**, top right
   of the actor page) and add it to a [schedule](https://docs.apify.com/platform/schedules) (Console → Schedules →
   **Create new**).
2. Collect results: download the dataset as JSON, CSV or Excel; fetch the latest run's results from the
   [API](https://apify.com/humble-echidna/ai-extract/api)
   (`GET https://api.apify.com/v2/actor-tasks/<task id>/runs/last/dataset/items?status=SUCCEEDED&format=json`, with
   your API token); let a [webhook](https://docs.apify.com/platform/integrations/webhooks) tell your system when a
   run succeeds; or connect it to Make, Zapier or n8n through
   [Apify's integrations](https://docs.apify.com/platform/integrations). It's built for those: they have no LLM of
   their own, and this gives them validated JSON with a stable shape.

#### Can I use AI Data Extractor from an AI agent (MCP)?

Yes, through [Apify's MCP server](https://docs.apify.com/platform/integrations/mcp), from Claude, ChatGPT, Cursor or
any other MCP client. Add this to your client's MCP configuration (or let the agent find it with the server's actor
search); your client signs you in to Apify:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=humble-echidna/ai-extract"
    }
  }
}
```

To use an [Apify API token](https://console.apify.com/settings/integrations) instead of signing in, add
`"headers": {"Authorization": "Bearer <APIFY_TOKEN>"}` next to `url`.

An agent reading a few pages can do it itself; this pays off for batches: pass
`{"urls": [...], "fields": "company name, services, email, phone", "maxPages": 200}` and get one validated record
per page without spending the agent's own context on the pages.

### Who it's for

Anyone who already has page URLs (a lead list, a product catalogue, a crawler's output) and needs the same few fields
from each, in a fixed shape: lead enrichment after a Google Maps scrape, product data from shops without an API,
directory and listing pages into a spreadsheet, automations in Make, Zapier or n8n.

### Why this one?

- **Validated, not just generated.** Every answer is checked against your schema here, whatever the model promised.
  One that doesn't match is marked `valid: false` with the reasons, and isn't charged.
- **Failed pages are free and explained.** Blocked, JavaScript-only, missing and invalid pages are rows with the
  reason in `errors`, not silent gaps, and none of them is charged.
- **$5 per 1,000 pages, tokens at the model's price on paid plans.** Pick any OpenRouter model; the token cost is
  what the model costs, and `inputTokens`/`outputTokens` show it per page.
- **Reads what matters on a page.** Headers, footers and sidebars stay in (that's where prices, stock and contact
  details live); scripts, menus and hidden text don't. The page's schema.org JSON-LD goes along too.
- **Chains after anything.** Dataset input with automatic URL-field detection, the same reader as our other actors.
- **Polite and safe.** It identifies itself honestly (User-Agent `HumbleEchidnaApify`), follows each site's
  robots.txt and Crawl-delay, skips sites that opt out of AI crawlers, and only requests public web addresses on the
  standard ports (80 and 443).

### Limits

- **No browser.** Pages that build their content with JavaScript are reported (`needs JavaScript`), not read.
  Scrape those with a browser-based actor first, or feed it their server-rendered versions.
- **No crawling.** One call per page you give it; for whole sites, chain it after a crawler or
  [Sitemap URL Extractor](https://apify.com/humble-echidna/sitemap-urls).
- **One record per page.** For a list page, ask for an array in your schema (`{"products": {"type": "array", ...}}`).
- **Long pages are cut** to `maxTokensPerPage` (start and end kept; `truncated: true`). Raise it for long pages.
- **Models can still be wrong** about what a page says, even with a valid answer. Validation checks the shape, not
  the facts; spot-check a sample before you rely on a new schema.
- **Only public pages** on ports 80 and 443; no logins, no proxies.

### FAQ

#### Why is a page marked "needs JavaScript"?

It has next to no text in its HTML and the markings of a client-side app, so without a browser there is nothing
real to read. It's not sent to the model and not charged.

#### Why is a page "blocked by robots.txt" or "opts out of AI use"?

The site's robots.txt disallows our crawler for that page, or disallows AI crawlers (GPTBot, ClaudeBot, CCBot,
Google-Extended and the like). Reading a page for an LLM is AI use, so those pages are skipped, not charged.

#### Why did a valid page come back with `null` values?

The model found nothing for those fields on the page, and it's told to use `null` rather than guess. Try
`pageContent: fullPage` if you'd switched to `mainContent`, raise `maxTokensPerPage` if the row says `truncated`, or
add a hint to the field (`price: the number next to "Our price"`).

#### Who pays for the model?

You do, through your own Apify account: the actor calls Apify's OpenRouter actor with your run's token, and Apify
bills the tokens to the account running the actor. No API key is needed and none of ours is used.

#### Is it legal to extract data from web pages?

It reads only the public pages you give it, the way a browser would, and follows robots.txt and AI opt-outs. What
you do with the data is up to you: don't collect personal data you have no lawful basis for, and respect the sites'
terms.

### Related actors

| Actor | Use it when |
|---|---|
| [Website & Page to Markdown](https://apify.com/humble-echidna/page-to-markdown) | You want the page text itself (for RAG), not fields. |
| [Sitemap URL Extractor](https://apify.com/humble-echidna/sitemap-urls) | You need every page URL of a site to feed in here. |
| [Dataset Transformer](https://apify.com/humble-echidna/dataset-transform) | You need to filter, rename or reshape the results afterwards. |

### Feedback and support

Found a bug, or a page type it reads badly? Open an issue on the **Issues** tab with the input you used.

### Versions

Current version: **1.0**. See the Changelog tab for what changed in each version.

# Changelog

This Actor's version history is a separate document: https://apify.com/humble-echidna/ai-extract/changelog.md

# Actor input Schema

## `urls` (type: `array`):

The pages to extract from, one per line, e.g. https://books.toscrape.com/catalogue/a-light-in-the-attic\_1000/index.html; a missing https:// is added. Each page is read once, no crawling. Pages the site's robots.txt disallows or opts out of AI crawlers are skipped, as are pages that need JavaScript to show their content. Ignored when datasetId is set.

## `datasetId` (type: `string`):

Default empty. One of your Apify datasets, picked here or given by id, e.g. a Google Maps or web scraper's results; to chain this actor after another in an integration, `{{resource.defaultDatasetId}}`. Each item's page URL is read like a line of urls (urls is ignored). Each row carries sourceTitle, sourcePlaceId and sourceIndex from its item. Reads at most 20,000 items and 10,000 URLs, read-only.

## `datasetUrlField` (type: `string`):

Only with datasetId: the item field that holds the page URL, e.g. `website`, or a dotted path such as `metadata.url`. Leave empty (the default) to find it automatically: the first of url, pageUrl, link, website and loadedUrl that has a web address in the first 100 items, then the same names one level down. Google Maps links are never used.

## `fields` (type: `string`):

What to pull from each page, comma-separated or one per line, e.g. `product name, price, currency, in stock`. Add a hint after a colon: `price: the sale price, not the list price`. Each becomes a camelCase key in data (productName, price, ...) holding text, a number, true/false, a list of texts or null. Ignored when schema is set. At most 50 fields.

## `schema` (type: `object`):

Default empty. A JSON Schema for the object to extract from each page, with "type": "object" at the top, e.g. {"type": "object", "properties": {"price": {"type": "number"}, "services": {"type": "array", "items": {"type": "string"}}}}. Every answer is validated against it; one that doesn't match is returned with valid false and errors, and not charged. Overrides fields.

## `instructions` (type: `string`):

Default empty. Optional guidance for the model, e.g. `Prices are in EUR unless the page says otherwise` or `Only list services the business offers, not blog topics`. At most 4,000 characters.

## `model` (type: `string`):

The model that reads each page, as an OpenRouter model id. Default `openai/gpt-4.1-mini`: accurate and cheap for extraction. Cheaper: `openai/gpt-4.1-nano` or `google/gemini-2.5-flash-lite`; harder pages: `openai/gpt-5.4-mini`. It must support structured outputs. Its tokens are billed to your Apify account by Apify's OpenRouter actor, at its prices.

## `pageContent` (type: `string`):

`fullPage` (default): the whole visible page as Markdown except scripts, forms and navigation menus, so prices, stock and footer contact details are included. `mainContent`: only the article or main content, as page-to-markdown returns it; fewer tokens, best for articles and docs. Either way the page's schema.org JSON-LD is sent too.

## `maxTokensPerPage` (type: `integer`):

Hard cap on the page text sent to the model, in tokens (about 4 characters each). Default 8000. A longer page keeps its start and its end (where footers hold addresses) and has truncated true. From 500 to 100000. Lower it to cut token cost on long pages.

## `maxPages` (type: `integer`):

Stop after this many pages are extracted, e.g. 20. From 1 to 10000; leave empty (the default) for no limit. Pages that fail or don't match the schema don't count. The run also stops cleanly at the maximum cost per run you set in the run options, whichever comes first.

## Actor input object example

```json
{
  "urls": [
    "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
  ],
  "fields": "title, price, currency, in stock, number available, UPC",
  "model": "openai/gpt-4.1-mini",
  "pageContent": "fullPage",
  "maxTokensPerPage": 8000
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `runStats` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
    ],
    "fields": "title, price, currency, in stock, number available, UPC",
    "model": "openai/gpt-4.1-mini"
};

// Run the Actor and wait for it to finish
const run = await client.actor("humble-echidna/ai-extract").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"],
    "fields": "title, price, currency, in stock, number available, UPC",
    "model": "openai/gpt-4.1-mini",
}

# Run the Actor and wait for it to finish
run = client.actor("humble-echidna/ai-extract").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
  ],
  "fields": "title, price, currency, in stock, number available, UPC",
  "model": "openai/gpt-4.1-mini"
}' |
apify call humble-echidna/ai-extract --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,humble-echidna/ai-extract"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6PjlfDBBQsLrpNsnP/builds/PUv9hyeDmKOn3S72N/openapi.json
