AI Data Extractor: Web Pages to Structured JSON
Pricing
from $5.00 / 1,000 page extracteds
AI Data Extractor: Web Pages to Structured JSON
Turn web pages, or another Actor's dataset, into clean JSON with an LLM. Give a JSON schema or list the fields in plain English. Output is validated against your schema; pages that can't be fetched or don't match aren't charged. $5 per 1,000 pages plus the model's tokens.
Pricing
from $5.00 / 1,000 page extracteds
Rating
0.0
(0)
Developer
Michael Costa
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
What does AI Data Extractor: Web Pages to Structured JSON do?
AI Data Extractor turns web pages into structured JSON with an LLM. Give it page URLs or another actor's dataset, and say what you want: plain-English fields ("product name, price, in stock") or a JSON Schema. Every answer is validated against your schema, and you pay only for pages that match.
It is not a crawler: it reads the pages you give it, one call per page. To get the pages first, chain it after any scraper or crawler (how).
Input fields · API · Use it from Claude, ChatGPT or Cursor (MCP)
Jump to: Fields · Price · How to use · Chain after Google Maps · Input · Output · AI agents (MCP) · Limits · FAQ
Try it in one click: the input comes pre-filled with one book page from books.toscrape.com (a sandbox site made for scraping) and six fields. That's 1 page: $0.005, plus the $0.00005 start fee, plus about $0.0005 of model tokens on a paid Apify plan. Then replace it with your own pages and fields.
What data does AI Data Extractor return?
One row per page you give it, including the pages that failed (with the reason, never charged).
| Field | Example | Notes |
|---|---|---|
data | {"title": "How Music Works", "price": 37.32, "currency": "GBP", "inStock": true} | The extracted record: your fields as camelCase keys, or your schema's shape. |
valid | true | true when data matches your schema. Only these pages are charged. |
errors | ["price: 'cheap' is not of type 'number'"] | Why a page has no valid data: blocked, needs JavaScript, HTTP error, or what didn't match. Empty when valid. |
url | https://books.toscrape.com/catalogue/how-music-works_979/index.html | The page as you gave it. |
loadedUrl | same | Where it ended up after redirects; null if it couldn't be fetched. |
title | How Music Works | Books to Scrape - Sandbox | The page title. |
model | openai/gpt-4.1-mini | The model that read the page. |
inputTokens, outputTokens | 1629, 69 | Tokens the model reported: what the token charge is based on. |
approxTokens, truncated | 812, false | Page text sent (about 4 characters a token); true if the page was cut to fit maxTokensPerPage. |
sourceTitle, sourcePlaceId, sourceIndex | "Lamp Shop", "ChIJ...", 0 | With a dataset as input: the item the URL came from, to join rows back to your leads. null for typed-in URLs. |
The full list is under Output.
How much does it cost to extract data from web pages with AI?
Two parts, both on your Apify account:
- This actor: $5.00 per 1,000 pages extracted ($0.005 a page), plus $0.00005 each time a run starts. Only pages whose answer matched your schema count.
- The model's tokens, billed by Apify's own OpenRouter actor, which this
actor calls with your run's token. On paid Apify plans that's OpenRouter's list price; on the Free plan the
OpenRouter actor charges 10 times as much (its pricing, not ours). With the default model,
openai/gpt-4.1-mini, ten book pages took 1,790 tokens a page on average: $0.84 per 1,000 pages on a paid plan ($8.45 on Free). Longer pages cost more;maxTokensPerPagecaps it.
- The example below: 10 pages × $0.005 = $0.05, plus the start fee and about $0.0085 of tokens (paid plan).
- A month, for example: 2,000 leads from a Google Maps scraper, each business website read once for services and contacts: 2,000 × $0.005 = $10, plus roughly $2-4 of tokens (business home pages run larger than book pages).
- Caps: Max pages per run in the input, and Maximum cost per run in the run options (it covers this actor's charges; the token charges are the OpenRouter actor's, so cap those with Max tokens per page and Max pages per run). The run stops cleanly at whichever comes first. Each page is counted against the limit before it's fetched, and the count is given back if the page fails, so a capped run never reads a page it can't return.
Never charged by this actor: pages that fail, need JavaScript, are disallowed by robots.txt or opted out of AI use, and answers that don't match your schema. (A page that reached the model still used tokens.)
How to extract structured data from web pages
- Open AI Data Extractor and click Try for free (or Start if you're signed in).
- Put your pages in Web page URLs, one per line, or pick a dataset in Or: page URLs from a dataset.
- Say what to extract: list the fields in Fields to extract (
product name, price, currency, in stock), or paste a JSON Schema into Or: JSON Schema for one record. Add Extra instructions if the pages need explaining ("prices are in EUR unless the page says otherwise"). - Click Start, then open the Output tab:
dataholds each page's record. Export as JSON, CSV or Excel.
Plain-English fields or a JSON Schema?
- Fields are quickest: each becomes a camelCase key (
in stock→inStock) that holds text, a number, true/false, a list of texts ornull. Add a hint after a colon:price: the sale price, not the list price. One field per line, or comma-separated on one line; at most 50. - A JSON Schema gives you exact types, nesting and required fields, e.g. a list of objects per page. Answers
that don't match (a price as text, a missing required field) come back with
valid: false, the model's answer indataand the reasons inerrors, and aren't charged.
Chain it after Google Maps or any scraper
Point it at another actor's results and each item's page URL is read like a line of Web page URLs. The field is
found automatically (url, pageUrl, link, website, loadedUrl, then the same names one level down); a Google
Maps link is never used, so after a Maps scraper it reads each business's own website.
- Once: pick the dataset in Or: page URLs from a dataset (
datasetId) and click Start. - Every time the other actor finishes: on that actor (or its task), open Integrations, add Run an actor
or task, pick AI Data Extractor, and set the input to include
"datasetId": "{{resource.defaultDatasetId}}"plus yourfields.
Each row carries sourceTitle, sourcePlaceId and sourceIndex from its item, so you can join the extracted
services and contacts back to your leads. It reads up to 20,000 items and 10,000 distinct URLs per run.
Example: ten book pages
Input (a real run on Apify, 2026-09-30; nine book pages and the site's home page):
{"urls": ["https://books.toscrape.com/catalogue/how-music-works_979/index.html", "..."],"fields": "title, price, currency, in stock, number available, UPC, product description (first sentence)"}
All 10 pages validated in 11 seconds. One row, unshortened:
{"url": "https://books.toscrape.com/catalogue/how-music-works_979/index.html","loadedUrl": "https://books.toscrape.com/catalogue/how-music-works_979/index.html","title": "How Music Works | Books to Scrape - Sandbox","data": {"title": "How Music Works","price": 37.32,"currency": "GBP","inStock": true,"numberAvailable": 19,"upc": "327f68a59745c102","productDescriptionFirstSentence": "How Music Works is David Byrne’s remarkable and buoyant celebration of a subject he has spent a lifetime thinking about."},"valid": true,"errors": [],"model": "openai/gpt-4.1-mini","approxTokens": 812,"inputTokens": 1629,"outputTokens": 69,"truncated": false,"fetchedAt": "2026-09-30T03:01:14.281857Z","sourceTitle": null,"sourcePlaceId": null,"sourceIndex": null}
The home page, which isn't a product, came back with price, currency and upc as null: the model is told to
use null for what a page doesn't say, never to guess. Cost: $0.05 for the pages plus $0.00845 of tokens (845
OpenRouter token events at the paid-plan price).
Ready-to-run examples
Each example opens with the input already filled in. Run it as it is, or change the input first.
- Extract product name, price and stock from product pages: Turn product pages into clean JSON: name, price, currency, stock and UPC, validated against the fields you list.
- Pull services and contact details from company websites: Read company websites and get the services they offer plus email, phone and address as JSON.
Input
| Field | What it does |
|---|---|
Web page URLs (urls) | The pages, one per line. Ignored when a dataset is set. |
Or: page URLs from a dataset (datasetId) | Another actor's results; chain with {{resource.defaultDatasetId}}. |
Field with the page URL (datasetUrlField) | Only with a dataset; empty = found automatically. |
Fields to extract (fields) | Plain-English fields. Ignored when a schema is set. |
Or: JSON Schema for one record (schema) | A JSON Schema with "type": "object" at the top. |
Extra instructions (instructions) | Guidance for the model, up to 4,000 characters. |
Model (model) | Any OpenRouter model id with structured outputs. Default openai/gpt-4.1-mini. |
What the model reads (pageContent) | fullPage (default): the whole visible page minus scripts, forms and menus, so prices and footer contacts are in. mainContent: the article only, fewer tokens. |
Max tokens per page (maxTokensPerPage) | Hard cap on the page text sent, default 8,000 (500-100,000). A longer page keeps its start and end. |
Max pages per run (maxPages) | Stop after this many valid pages (1-10,000); empty = no limit. |
{"urls": ["https://example.com/products/lamp", "https://example.com/products/chair"],"schema": {"type": "object","properties": {"name": {"type": "string"},"price": {"type": "number"},"currency": {"type": "string"},"variants": {"type": "array", "items": {"type": "object", "properties": {"color": {"type": "string"}, "inStock": {"type": "boolean"}}}}},"required": ["name", "price"]},"model": "openai/gpt-4.1-mini","maxPages": 100}
Which model? The default, openai/gpt-4.1-mini, is accurate and cheap for extraction. For big batches of
simple pages, openai/gpt-4.1-nano or google/gemini-2.5-flash-lite cost about a quarter as much; for messy pages
or deep schemas, openai/gpt-5.4-mini. The model must support structured outputs.
Output
One row per page, in the order they finish:
{"url": "https://shop.example.com/mystery-box","loadedUrl": "https://shop.example.com/mystery-box","title": "Mystery Box | Example Shop","data": {"name": "Mystery Box", "price": "cheap"},"valid": false,"errors": ["price: 'cheap' is not of type 'number'"],"model": "openai/gpt-4.1-mini","approxTokens": 640,"inputTokens": 1180,"outputTokens": 18,"truncated": false,"fetchedAt": "2026-09-30T03:01:14Z","sourceTitle": null,"sourcePlaceId": null,"sourceIndex": null}
A page that couldn't be read has data: null, loadedUrl: null and the reason in errors (for example
blocked by robots.txt or the page builds its content with JavaScript). The run's RUN_STATS record counts pages
extracted, invalid and failed, tokens used, and the options that applied. The Tokens view of the dataset shows
tokens per page.
Run it on a schedule, or from your own code
- Save your input as a task (Save as a new task, top right of the actor page) and add it to a schedule (Console → Schedules → Create new).
- Collect results: download the dataset as JSON, CSV or Excel; fetch the latest run's results from the
API
(
GET https://api.apify.com/v2/actor-tasks/<task id>/runs/last/dataset/items?status=SUCCEEDED&format=json, with your API token); let a webhook tell your system when a run succeeds; or connect it to Make, Zapier or n8n through Apify's integrations. It's built for those: they have no LLM of their own, and this gives them validated JSON with a stable shape.
Can I use AI Data Extractor from an AI agent (MCP)?
Yes, through Apify's MCP server, from Claude, ChatGPT, Cursor or any other MCP client. Add this to your client's MCP configuration (or let the agent find it with the server's actor search); your client signs you in to Apify:
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=humble-echidna/ai-extract"}}}
To use an Apify API token instead of signing in, add
"headers": {"Authorization": "Bearer <APIFY_TOKEN>"} next to url.
An agent reading a few pages can do it itself; this pays off for batches: pass
{"urls": [...], "fields": "company name, services, email, phone", "maxPages": 200} and get one validated record
per page without spending the agent's own context on the pages.
Who it's for
Anyone who already has page URLs (a lead list, a product catalogue, a crawler's output) and needs the same few fields from each, in a fixed shape: lead enrichment after a Google Maps scrape, product data from shops without an API, directory and listing pages into a spreadsheet, automations in Make, Zapier or n8n.
Why this one?
- Validated, not just generated. Every answer is checked against your schema here, whatever the model promised.
One that doesn't match is marked
valid: falsewith the reasons, and isn't charged. - Failed pages are free and explained. Blocked, JavaScript-only, missing and invalid pages are rows with the
reason in
errors, not silent gaps, and none of them is charged. - $5 per 1,000 pages, tokens at the model's price on paid plans. Pick any OpenRouter model; the token cost is
what the model costs, and
inputTokens/outputTokensshow it per page. - Reads what matters on a page. Headers, footers and sidebars stay in (that's where prices, stock and contact details live); scripts, menus and hidden text don't. The page's schema.org JSON-LD goes along too.
- Chains after anything. Dataset input with automatic URL-field detection, the same reader as our other actors.
- Polite and safe. It identifies itself honestly (User-Agent
HumbleEchidnaApify), follows each site's robots.txt and Crawl-delay, skips sites that opt out of AI crawlers, and only requests public web addresses on the standard ports (80 and 443).
Limits
- No browser. Pages that build their content with JavaScript are reported (
needs JavaScript), not read. Scrape those with a browser-based actor first, or feed it their server-rendered versions. - No crawling. One call per page you give it; for whole sites, chain it after a crawler or Sitemap URL Extractor.
- One record per page. For a list page, ask for an array in your schema (
{"products": {"type": "array", ...}}). - Long pages are cut to
maxTokensPerPage(start and end kept;truncated: true). Raise it for long pages. - Models can still be wrong about what a page says, even with a valid answer. Validation checks the shape, not the facts; spot-check a sample before you rely on a new schema.
- Only public pages on ports 80 and 443; no logins, no proxies.
FAQ
Why is a page marked "needs JavaScript"?
It has next to no text in its HTML and the markings of a client-side app, so without a browser there is nothing real to read. It's not sent to the model and not charged.
Why is a page "blocked by robots.txt" or "opts out of AI use"?
The site's robots.txt disallows our crawler for that page, or disallows AI crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended and the like). Reading a page for an LLM is AI use, so those pages are skipped, not charged.
Why did a valid page come back with null values?
The model found nothing for those fields on the page, and it's told to use null rather than guess. Try
pageContent: fullPage if you'd switched to mainContent, raise maxTokensPerPage if the row says truncated, or
add a hint to the field (price: the number next to "Our price").
Who pays for the model?
You do, through your own Apify account: the actor calls Apify's OpenRouter actor with your run's token, and Apify bills the tokens to the account running the actor. No API key is needed and none of ours is used.
Is it legal to extract data from web pages?
It reads only the public pages you give it, the way a browser would, and follows robots.txt and AI opt-outs. What you do with the data is up to you: don't collect personal data you have no lawful basis for, and respect the sites' terms.
Related actors
| Actor | Use it when |
|---|---|
| Website & Page to Markdown | You want the page text itself (for RAG), not fields. |
| Sitemap URL Extractor | You need every page URL of a site to feed in here. |
| Dataset Transformer | You need to filter, rename or reshape the results afterwards. |
Feedback and support
Found a bug, or a page type it reads badly? Open an issue on the Issues tab with the input you used.
Versions
Current version: 1.0. See the Changelog tab for what changed in each version.