AI Web Scraper: Any Website to JSON by Schema
Pricing
from $30.00 / 1,000 page extracteds
AI Web Scraper: Any Website to JSON by Schema
Give page URLs and the fields you want with their types. An AI web scraper reads each page and returns JSON that matches your schema: type coercion, null for missing fields, a validity flag and a check of whether each value is on the page.
Pricing
from $30.00 / 1,000 page extracteds
Rating
0.0
(0)
Developer
Turgay NANTA
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
0
Monthly active users
4 days ago
Last modified
Categories
Share
AI Web Scraper: Any Website to JSON with Your Own Schema
Writing a scraper for every new website means selectors, maintenance and broken runs whenever a page layout changes. Generic AI scrapers skip the selectors but give you output you cannot trust: a price that comes back as text, a field that silently disappears, a value the model made up.
Give this AI web scraper a list of page URLs and the fields you want, with their types. It reads each page and returns one JSON record per URL that matches your schema exactly: the right types, null for anything that is not on the page, a validity flag per record, and a check of whether each value was actually found on the page.
- You define the output:
{"title": "string", "price": "number", "in_stock": "boolean"}and that is what you get, nothing more. - Type coercion built in:
"$1,299.90"becomes1299.9,"20,000"becomes20000,"in stock"becomestrue,"a, b, c"becomes a list. - Missing means
null, never a silent gap. Required fields that are missing mark the record as invalid. - Grounding check: every text and number value is labelled
high(found on the page) orunverified(not found, check it), so likely model inventions are visible. - No selectors, no code: works on product pages, company pages, listings, articles and profiles.
- Five output languages for text values: English, Turkish, German, Spanish, French.
- Fair billing: a page that cannot be fetched or extracted is not charged.
Example: one product page, five fields
Input:
{"urls": ["https://shop.example.com/products/trail-backpack-30l"],"schema": {"title": "string", "price": "number", "currency": "string", "in_stock": "boolean", "rating": "number"}}
Output record:
{"title": "Trail Backpack 30L","price": 89.9,"currency": "EUR","in_stock": true,"rating": null,"_url": "https://shop.example.com/products/trail-backpack-30l","_schema_valid": true,"_field_status": {"title": "ok", "price": "coerced", "currency": "ok", "in_stock": "coerced", "rating": "missing"},"_confidence": {"title": "high", "price": "high", "currency": "high", "in_stock": "filled", "rating": "null"}}
Illustrative example; the page and values are made up to show the format. rating is null because the page has no rating, and it is reported as missing instead of being dropped.
Cost of this example: $0.0451 for the page (start fee + one extracted page + five validated fields), plus the platform usage of the page fetch. See Pricing.
At a glance
| Input | Page URLs (urls, also url or startUrls) + your field schema (schema) |
| Output | One record per URL with your fields + _url, _schema_valid, _field_status, _confidence; a final summary row |
| Field types | string, number, integer, boolean, array, object |
| Page fetching | Apify Website Content Crawler, one page per URL, JavaScript-capable |
| Extraction | Language model at temperature 0, instructed to use only what is on the page |
| Languages | Output text values in English, Turkish, German, Spanish or French |
| Pricing | Per extracted page + per validated field; see Pricing |
| Runs on | Apify Console, API, schedules, n8n, Make, Zapier, AI agents via MCP |
Quick start
- Paste one or more page URLs into URLs.
- Edit Schema: write the field names you want and their types. The form comes with a product example (
title,price,currency,in_stock,rating). - Pick the Output language for text values.
- Click Start. Open the Dataset tab: one row per URL, plus a summary row at the end.
An empty run (no URLs) fetches nothing, charges nothing and returns one row telling you what to fill in.
Why "your schema" matters
Most AI extraction tools let you describe fields in plain language. That is convenient, but the output contract is loose: a number can arrive as "1.299,90 TL", a boolean as "Yes", and a field the model could not find may just be missing from the object. Every downstream system (a database insert, a spreadsheet formula, a price comparison) then needs its own cleanup code.
This Actor puts a strict layer between the model and your data:
| Problem | What the Actor does |
|---|---|
| Numbers as text with currency, spaces, separators | Cleans symbols and resolves thousands vs decimal separators (1,299.90, 1.299,90, 20,000, 12,5) |
| Booleans as words | Recognizes common words in English and Turkish (yes, in stock, available, sold out, out of stock...) |
| Lists as comma-separated text | Splits into an array |
| Integers with decimals | Rounds to the nearest integer for integer fields |
| Field not on the page | Writes null and reports missing (or required_missing) |
| Value the model could not convert | Writes null instead of a wrong type |
| Extra fields the model added | Dropped; only your schema fields are returned |
| Value not actually on the page | Labelled unverified in _confidence |
Use cases
E-commerce and pricing teams
Extract title, price, currency, in_stock and sku from competitor product pages on different shops, each with its own layout, into one consistent table. No per-shop scraper to maintain.
Lead generation and sales operations
From a list of company websites, extract company_name, industry, city, contact_email and services (array). Mark company_name as required so incomplete records are flagged.
Market and investment research
Pull founded_year (integer), headquarters, employee_count (number) and products (array) from company "About" pages into a comparison sheet.
Real estate and listings
From listing pages, extract price, area_m2, rooms, address and features. Coercion handles local number formats; unverified labels show values to double-check.
Recruiting and HR
From job posting pages, extract job_title, location, remote (boolean), salary_min, salary_max and requirements (array).
Content and SEO teams
From articles, extract headline, author, published_date, summary and topics, in the output language you need, even when the source is in another language.
Developers and data engineers
Use it as the extraction step in a pipeline: you already know the URLs, you need typed JSON that loads into a database without cleanup code.
AI agents
An agent that needs specific facts from a known page can call the Actor with a small schema and get typed values plus a grounding label for each one.
Worked example: a competitor price table
You sell outdoor gear and want a weekly price table for 40 competitor products across 6 shops.
- Collect the product URLs in a sheet and paste them into
urls. - Define the schema:
{"title": {"type": "string", "required": true},"price": {"type": "number", "required": true},"currency": "string","in_stock": "boolean","shipping_cost": "number"}
- Run. Each URL returns one record. Records where the title or price could not be found have
_schema_valid: false. - Check the summary row.
validity_ratetells you how many pages produced complete records;field_qualityshows, per field, how often it wasok,coerced,missingorrequired_missing. - Filter on
_confidence.price == "high"for the values you want to trust automatically; review theunverifiedones. - Save as a task and schedule it weekly. Export to Google Sheets and chart price changes.
Cost for 40 pages with 5 fields: $0.0001 + 40 x $0.03 + 200 x $0.003 = $1.8001, plus the platform usage of the page fetches.
How it works
- Collect URLs. From
urls,urlandstartUrls(all accepted, duplicates removed, order kept). - Fetch. For each URL, the Actor runs Apify's Website Content Crawler for that single page (adaptive browser mode, so JavaScript-rendered pages work) and takes the page text. It does not follow links.
- Prompt. The first 12,000 characters of the page text and your field list go to the language model with strict rules: use only what is on the page, write
nullfor missing fields, return the stated types, return only JSON, write text values in the chosen language. - Parse. The JSON object is extracted from the reply (code fences are removed). A malformed reply is logged, not hidden.
- Validate and coerce. Each schema field is coerced to its type. The record keeps only your fields. Each field gets a status:
ok,coerced,missingorrequired_missing. - Ground. Each text or number value is searched for on the page text:
highwhen found,unverifiedwhen not. Booleans, arrays and objects are labelledfilled; empty valuesnull. - Write. One record per URL, then a summary row across all pages.
Pages are processed one after another. If one page fails, the others continue, and the failed page still gets a record (all fields null) so you can see which URL failed.
Input
| Field | Type | Default | Description |
|---|---|---|---|
urls | array of strings | none | Page URLs to extract from; one record per URL. url (single string) and startUrls (Apify format [{"url": ...}]) are also accepted. |
schema | object or array | product example | The fields you want and their types (formats below). |
language | string | English | Language for text values: English, Türkçe, Deutsch, Español, Français. |
model | string | claude-haiku-4-5-20251001 | Advanced. The default is fast and economical; claude-sonnet-4-6 for complex pages. |
urls has no default on purpose: fetching a page starts a Website Content Crawler run on your account, so nothing is fetched until you give a URL.
Schema formats
Simple map (name to type):
{"title": "string", "price": "number", "in_stock": "boolean"}
Map with required fields:
{"title": {"type": "string", "required": true}, "price": {"type": "number", "required": true}, "tags": "array"}
List of objects:
[{"name": "title", "type": "string", "required": true}, {"name": "rating", "type": "number"}]
List of strings (name or name:type):
["title", "price:number", "in_stock:boolean"]
Type aliases are accepted: str/text = string, int = integer, float/double/decimal = number, bool = boolean, list = array, dict/json = object. Unknown types fall back to string.
Output
One record per URL
| Key | Content |
|---|---|
| your fields | The values, in your types; null when missing or not convertible |
_url | The page URL |
_schema_valid | true when every required field has a value |
_field_status | Per field: ok, coerced (value was converted to your type), missing, required_missing |
_confidence | Per field: high (found on the page), unverified (not found on the page, possible invention), filled (boolean, array or object), null (empty) |
Summary row
The last row has _type: "summary" and contains:
| Key | Content |
|---|---|
page_count | Pages processed |
schema_valid_records | Records with every required field filled |
validity_rate | schema_valid_records / page_count |
field_quality | Per field: counts of ok, coerced, missing, required_missing |
How grounding decides
- Text: the first 40 characters of the value (at least 3) must appear in the page text, case-insensitive.
- Numbers: the digits of the value (at least 2) must appear in the page's digit sequence, so
1299.9matches1.299,90on the page. Single-digit numbers are alwaysunverified, because they match almost anything. - Translated text (when the output language differs from the page) will often be
unverified, because the translated words are not on the page. That is expected.
Grounding is a deterministic text match, not proof. Treat unverified as "check this", not as "wrong".
Pricing
This Actor charges per event. Current prices are on the Pricing tab; at the time of writing:
| Event | When | Price |
|---|---|---|
actor-start | Once per run | $0.0001 |
page-extracted | Per page successfully extracted | $0.03 |
field-validated | Per schema field, per successfully extracted page | $0.003 |
Pages that could not be fetched or whose extraction failed are not charged (only the start fee applies to them).
Examples
| Run | Calculation | Total |
|---|---|---|
| 1 page, 5 fields | $0.0001 + $0.03 + 5 x $0.003 | $0.0451 |
| 10 pages, 5 fields | $0.0001 + 10 x $0.03 + 50 x $0.003 | $0.4501 |
| 100 pages, 5 fields | $0.0001 + 100 x $0.03 + 500 x $0.003 | $4.5001 |
| 100 pages, 10 fields | $0.0001 + 100 x $0.03 + 1,000 x $0.003 | $6.0001 |
| Empty input | nothing fetched | $0 |
Page fetching is billed separately: each page is fetched by a Website Content Crawler run on your Apify account, and the platform usage of that run is billed to you at Apify's normal rates. It is a single-page crawl per URL.
Integrations
API
curl -X POST "https://api.apify.com/v2/acts/enezli~ai-website-to-dataset/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"urls":["https://example.com/product/1"],"schema":{"title":"string","price":"number"}}'
Python
from apify_client import ApifyClientclient = ApifyClient("YOUR_TOKEN")run = client.actor("enezli/ai-website-to-dataset").call(run_input={"urls": ["https://example.com/product/1", "https://example.com/product/2"],"schema": {"title": {"type": "string", "required": True}, "price": "number", "in_stock": "boolean"},})for rec in client.dataset(run["defaultDatasetId"]).iterate_items():if rec.get("_type") == "summary":print("validity:", rec["validity_rate"])else:print(rec["_url"], rec["title"], rec["price"], rec["_confidence"])
Schedules
Save the input as a task and add a schedule to re-extract the same pages daily or weekly (price tracking, stock checks, content changes).
n8n, Make, Zapier
Run Actor → Get dataset items → skip the summary row → write to Google Sheets, Airtable, a database or a webhook. Use _schema_valid to route incomplete records to a review step.
AI agents (MCP)
Available through the Apify MCP server. An assistant can call it with a URL and a small schema and receive typed values with grounding labels.
Chaining with other Actors
If another Actor gives you URLs (a sitemap scraper, a search results scraper), pass its URL list into urls to turn each page into a structured record.
Tips for better results
- Keep schemas small and specific. 5-10 well-named fields extract better than 30 vague ones.
- Use clear field names.
price_eurorfounded_yeartells the model more thanvalue1. - Mark the fields you cannot live without as required, then filter on
_schema_valid. - Use
integerfor counts and years,numberfor prices and ratings. - Use
arrayfor lists (features, tags, services) andbooleanfor yes/no facts. - Pick the stronger model for long, dense or messy pages.
- Keep the output language equal to the page language when you care about grounding; translated values are harder to verify.
- Check the summary row first. A low
validity_rateusually means the pages do not contain the fields you asked for.
Limitations
- One page per URL. The Actor does not crawl a site or follow links. Give it the exact pages you want.
- First 12,000 characters of page text. Values that appear only further down a very long page may be missed.
- Visible text only. Values that exist only in images, PDFs linked from the page, or behind a login are not available.
- Model output size. Very large schemas or long arrays can exceed the model's reply size; the page then gets an empty record and is not charged.
- Grounding is heuristic. A correct value phrased differently from the page can be
unverified; a short common word can match by chance. - Sites that block crawlers may return empty pages; those pages are not charged.
- Sequential processing. Large URL lists take proportionally longer; split them across parallel runs if speed matters.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| One row saying "No URL given" | urls was empty | Add page URLs |
A record with every field null | The page could not be fetched or extracted (see the log) | Open the URL in a browser; try again later or check if the site blocks crawlers. The page was not charged |
Many missing fields | The page does not contain those facts, or they are far down the page | Check the page; use a more specific URL |
_schema_valid: false | A required field is missing | Expected when the page lacks it; make the field optional if that is acceptable |
Numbers labelled unverified | The value was computed or reformatted by the model | Check a few manually; the value may still be correct |
Text labelled unverified | Output language differs from the page, or the model paraphrased | Use the page language, or treat as "review" |
| Run failed: "No valid schema provided" | The schema was empty or unreadable | Use one of the formats in Schema formats |
FAQ
Do I need to write CSS selectors or code? No. You write field names and types. The model finds the values on the page.
Can it crawl a whole website? No. It extracts from the exact URLs you give it, one page each. For site-wide crawling, collect the URLs first with a crawler or sitemap tool, then pass them here.
Does it work on JavaScript-heavy pages? Pages are fetched by Website Content Crawler in adaptive browser mode, which renders JavaScript when needed.
Will it invent values?
The model is instructed to use only page content and write null otherwise, and it runs at temperature 0. Because no model is perfect, every value is also checked against the page text and labelled high or unverified.
What happens with extra fields the model returns? They are dropped. The record contains only your schema fields, always in the same order and types.
Can it translate? Yes. Text values are written in the output language you choose, even when the page is in another language. Field names are never translated.
Am I charged for pages that failed? No. Pages that could not be fetched or extracted are not charged. The page fetch itself uses platform resources on your account.
How many URLs can I send? There is no fixed limit in the Actor. Pages are processed one after another within the run timeout; for large lists, split them across several runs.
Which number formats are understood?
US and European separators (1,299.90, 1.299,90), thousands only (20,000, 1.234.567), decimal commas (12,5), and currency or unit symbols around the number.
Is my data stored? Only in your own Apify storage. Page text is sent to the language model provider solely to extract your fields.
Can I use it from an AI agent? Yes, through the Apify MCP server or the API.
What's new
- Fair billing: pages that cannot be fetched or extracted are no longer charged.
- Better number parsing: thousands separators such as
20,000and1.234.567are read correctly. - Empty input is free: a run without URLs fetches nothing and explains what to fill in.
Trademarks
This Actor is not affiliated with, endorsed by, or connected to any website you extract data from. All product and company names are trademarks of their respective owners. Make sure your use of extracted data complies with the terms of the websites you process and with applicable law.


