AI Web Scraper: Any Website to JSON by Schema avatar

AI Web Scraper: Any Website to JSON by Schema

Pricing

from $30.00 / 1,000 page extracteds

Go to Apify Store
AI Web Scraper: Any Website to JSON by Schema

AI Web Scraper: Any Website to JSON by Schema

Give page URLs and the fields you want with their types. An AI web scraper reads each page and returns JSON that matches your schema: type coercion, null for missing fields, a validity flag and a check of whether each value is on the page.

Pricing

from $30.00 / 1,000 page extracteds

Rating

0.0

(0)

Developer

Turgay NANTA

Turgay NANTA

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

4 days ago

Last modified

Share

AI Web Scraper: Any Website to JSON with Your Own Schema

Writing a scraper for every new website means selectors, maintenance and broken runs whenever a page layout changes. Generic AI scrapers skip the selectors but give you output you cannot trust: a price that comes back as text, a field that silently disappears, a value the model made up.

Give this AI web scraper a list of page URLs and the fields you want, with their types. It reads each page and returns one JSON record per URL that matches your schema exactly: the right types, null for anything that is not on the page, a validity flag per record, and a check of whether each value was actually found on the page.

  • You define the output: {"title": "string", "price": "number", "in_stock": "boolean"} and that is what you get, nothing more.
  • Type coercion built in: "$1,299.90" becomes 1299.9, "20,000" becomes 20000, "in stock" becomes true, "a, b, c" becomes a list.
  • Missing means null, never a silent gap. Required fields that are missing mark the record as invalid.
  • Grounding check: every text and number value is labelled high (found on the page) or unverified (not found, check it), so likely model inventions are visible.
  • No selectors, no code: works on product pages, company pages, listings, articles and profiles.
  • Five output languages for text values: English, Turkish, German, Spanish, French.
  • Fair billing: a page that cannot be fetched or extracted is not charged.

Example: one product page, five fields

Input:

{
"urls": ["https://shop.example.com/products/trail-backpack-30l"],
"schema": {"title": "string", "price": "number", "currency": "string", "in_stock": "boolean", "rating": "number"}
}

Output record:

{
"title": "Trail Backpack 30L",
"price": 89.9,
"currency": "EUR",
"in_stock": true,
"rating": null,
"_url": "https://shop.example.com/products/trail-backpack-30l",
"_schema_valid": true,
"_field_status": {"title": "ok", "price": "coerced", "currency": "ok", "in_stock": "coerced", "rating": "missing"},
"_confidence": {"title": "high", "price": "high", "currency": "high", "in_stock": "filled", "rating": "null"}
}

Illustrative example; the page and values are made up to show the format. rating is null because the page has no rating, and it is reported as missing instead of being dropped.

Cost of this example: $0.0451 for the page (start fee + one extracted page + five validated fields), plus the platform usage of the page fetch. See Pricing.


At a glance

InputPage URLs (urls, also url or startUrls) + your field schema (schema)
OutputOne record per URL with your fields + _url, _schema_valid, _field_status, _confidence; a final summary row
Field typesstring, number, integer, boolean, array, object
Page fetchingApify Website Content Crawler, one page per URL, JavaScript-capable
ExtractionLanguage model at temperature 0, instructed to use only what is on the page
LanguagesOutput text values in English, Turkish, German, Spanish or French
PricingPer extracted page + per validated field; see Pricing
Runs onApify Console, API, schedules, n8n, Make, Zapier, AI agents via MCP

Quick start

  1. Paste one or more page URLs into URLs.
  2. Edit Schema: write the field names you want and their types. The form comes with a product example (title, price, currency, in_stock, rating).
  3. Pick the Output language for text values.
  4. Click Start. Open the Dataset tab: one row per URL, plus a summary row at the end.

An empty run (no URLs) fetches nothing, charges nothing and returns one row telling you what to fill in.


Why "your schema" matters

Most AI extraction tools let you describe fields in plain language. That is convenient, but the output contract is loose: a number can arrive as "1.299,90 TL", a boolean as "Yes", and a field the model could not find may just be missing from the object. Every downstream system (a database insert, a spreadsheet formula, a price comparison) then needs its own cleanup code.

This Actor puts a strict layer between the model and your data:

ProblemWhat the Actor does
Numbers as text with currency, spaces, separatorsCleans symbols and resolves thousands vs decimal separators (1,299.90, 1.299,90, 20,000, 12,5)
Booleans as wordsRecognizes common words in English and Turkish (yes, in stock, available, sold out, out of stock...)
Lists as comma-separated textSplits into an array
Integers with decimalsRounds to the nearest integer for integer fields
Field not on the pageWrites null and reports missing (or required_missing)
Value the model could not convertWrites null instead of a wrong type
Extra fields the model addedDropped; only your schema fields are returned
Value not actually on the pageLabelled unverified in _confidence

Use cases

E-commerce and pricing teams

Extract title, price, currency, in_stock and sku from competitor product pages on different shops, each with its own layout, into one consistent table. No per-shop scraper to maintain.

Lead generation and sales operations

From a list of company websites, extract company_name, industry, city, contact_email and services (array). Mark company_name as required so incomplete records are flagged.

Market and investment research

Pull founded_year (integer), headquarters, employee_count (number) and products (array) from company "About" pages into a comparison sheet.

Real estate and listings

From listing pages, extract price, area_m2, rooms, address and features. Coercion handles local number formats; unverified labels show values to double-check.

Recruiting and HR

From job posting pages, extract job_title, location, remote (boolean), salary_min, salary_max and requirements (array).

Content and SEO teams

From articles, extract headline, author, published_date, summary and topics, in the output language you need, even when the source is in another language.

Developers and data engineers

Use it as the extraction step in a pipeline: you already know the URLs, you need typed JSON that loads into a database without cleanup code.

AI agents

An agent that needs specific facts from a known page can call the Actor with a small schema and get typed values plus a grounding label for each one.


Worked example: a competitor price table

You sell outdoor gear and want a weekly price table for 40 competitor products across 6 shops.

  1. Collect the product URLs in a sheet and paste them into urls.
  2. Define the schema:
    {
    "title": {"type": "string", "required": true},
    "price": {"type": "number", "required": true},
    "currency": "string",
    "in_stock": "boolean",
    "shipping_cost": "number"
    }
  3. Run. Each URL returns one record. Records where the title or price could not be found have _schema_valid: false.
  4. Check the summary row. validity_rate tells you how many pages produced complete records; field_quality shows, per field, how often it was ok, coerced, missing or required_missing.
  5. Filter on _confidence.price == "high" for the values you want to trust automatically; review the unverified ones.
  6. Save as a task and schedule it weekly. Export to Google Sheets and chart price changes.

Cost for 40 pages with 5 fields: $0.0001 + 40 x $0.03 + 200 x $0.003 = $1.8001, plus the platform usage of the page fetches.


How it works

  1. Collect URLs. From urls, url and startUrls (all accepted, duplicates removed, order kept).
  2. Fetch. For each URL, the Actor runs Apify's Website Content Crawler for that single page (adaptive browser mode, so JavaScript-rendered pages work) and takes the page text. It does not follow links.
  3. Prompt. The first 12,000 characters of the page text and your field list go to the language model with strict rules: use only what is on the page, write null for missing fields, return the stated types, return only JSON, write text values in the chosen language.
  4. Parse. The JSON object is extracted from the reply (code fences are removed). A malformed reply is logged, not hidden.
  5. Validate and coerce. Each schema field is coerced to its type. The record keeps only your fields. Each field gets a status: ok, coerced, missing or required_missing.
  6. Ground. Each text or number value is searched for on the page text: high when found, unverified when not. Booleans, arrays and objects are labelled filled; empty values null.
  7. Write. One record per URL, then a summary row across all pages.

Pages are processed one after another. If one page fails, the others continue, and the failed page still gets a record (all fields null) so you can see which URL failed.


Input

FieldTypeDefaultDescription
urlsarray of stringsnonePage URLs to extract from; one record per URL. url (single string) and startUrls (Apify format [{"url": ...}]) are also accepted.
schemaobject or arrayproduct exampleThe fields you want and their types (formats below).
languagestringEnglishLanguage for text values: English, Türkçe, Deutsch, Español, Français.
modelstringclaude-haiku-4-5-20251001Advanced. The default is fast and economical; claude-sonnet-4-6 for complex pages.

urls has no default on purpose: fetching a page starts a Website Content Crawler run on your account, so nothing is fetched until you give a URL.

Schema formats

Simple map (name to type):

{"title": "string", "price": "number", "in_stock": "boolean"}

Map with required fields:

{"title": {"type": "string", "required": true}, "price": {"type": "number", "required": true}, "tags": "array"}

List of objects:

[{"name": "title", "type": "string", "required": true}, {"name": "rating", "type": "number"}]

List of strings (name or name:type):

["title", "price:number", "in_stock:boolean"]

Type aliases are accepted: str/text = string, int = integer, float/double/decimal = number, bool = boolean, list = array, dict/json = object. Unknown types fall back to string.


Output

One record per URL

KeyContent
your fieldsThe values, in your types; null when missing or not convertible
_urlThe page URL
_schema_validtrue when every required field has a value
_field_statusPer field: ok, coerced (value was converted to your type), missing, required_missing
_confidencePer field: high (found on the page), unverified (not found on the page, possible invention), filled (boolean, array or object), null (empty)

Summary row

The last row has _type: "summary" and contains:

KeyContent
page_countPages processed
schema_valid_recordsRecords with every required field filled
validity_rateschema_valid_records / page_count
field_qualityPer field: counts of ok, coerced, missing, required_missing

How grounding decides

  • Text: the first 40 characters of the value (at least 3) must appear in the page text, case-insensitive.
  • Numbers: the digits of the value (at least 2) must appear in the page's digit sequence, so 1299.9 matches 1.299,90 on the page. Single-digit numbers are always unverified, because they match almost anything.
  • Translated text (when the output language differs from the page) will often be unverified, because the translated words are not on the page. That is expected.

Grounding is a deterministic text match, not proof. Treat unverified as "check this", not as "wrong".


Pricing

This Actor charges per event. Current prices are on the Pricing tab; at the time of writing:

EventWhenPrice
actor-startOnce per run$0.0001
page-extractedPer page successfully extracted$0.03
field-validatedPer schema field, per successfully extracted page$0.003

Pages that could not be fetched or whose extraction failed are not charged (only the start fee applies to them).

Examples

RunCalculationTotal
1 page, 5 fields$0.0001 + $0.03 + 5 x $0.003$0.0451
10 pages, 5 fields$0.0001 + 10 x $0.03 + 50 x $0.003$0.4501
100 pages, 5 fields$0.0001 + 100 x $0.03 + 500 x $0.003$4.5001
100 pages, 10 fields$0.0001 + 100 x $0.03 + 1,000 x $0.003$6.0001
Empty inputnothing fetched$0

Page fetching is billed separately: each page is fetched by a Website Content Crawler run on your Apify account, and the platform usage of that run is billed to you at Apify's normal rates. It is a single-page crawl per URL.


Integrations

API

curl -X POST "https://api.apify.com/v2/acts/enezli~ai-website-to-dataset/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://example.com/product/1"],"schema":{"title":"string","price":"number"}}'

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("enezli/ai-website-to-dataset").call(run_input={
"urls": ["https://example.com/product/1", "https://example.com/product/2"],
"schema": {"title": {"type": "string", "required": True}, "price": "number", "in_stock": "boolean"},
})
for rec in client.dataset(run["defaultDatasetId"]).iterate_items():
if rec.get("_type") == "summary":
print("validity:", rec["validity_rate"])
else:
print(rec["_url"], rec["title"], rec["price"], rec["_confidence"])

Schedules

Save the input as a task and add a schedule to re-extract the same pages daily or weekly (price tracking, stock checks, content changes).

n8n, Make, Zapier

Run Actor → Get dataset items → skip the summary row → write to Google Sheets, Airtable, a database or a webhook. Use _schema_valid to route incomplete records to a review step.

AI agents (MCP)

Available through the Apify MCP server. An assistant can call it with a URL and a small schema and receive typed values with grounding labels.

Chaining with other Actors

If another Actor gives you URLs (a sitemap scraper, a search results scraper), pass its URL list into urls to turn each page into a structured record.


Tips for better results

  • Keep schemas small and specific. 5-10 well-named fields extract better than 30 vague ones.
  • Use clear field names. price_eur or founded_year tells the model more than value1.
  • Mark the fields you cannot live without as required, then filter on _schema_valid.
  • Use integer for counts and years, number for prices and ratings.
  • Use array for lists (features, tags, services) and boolean for yes/no facts.
  • Pick the stronger model for long, dense or messy pages.
  • Keep the output language equal to the page language when you care about grounding; translated values are harder to verify.
  • Check the summary row first. A low validity_rate usually means the pages do not contain the fields you asked for.

Limitations

  • One page per URL. The Actor does not crawl a site or follow links. Give it the exact pages you want.
  • First 12,000 characters of page text. Values that appear only further down a very long page may be missed.
  • Visible text only. Values that exist only in images, PDFs linked from the page, or behind a login are not available.
  • Model output size. Very large schemas or long arrays can exceed the model's reply size; the page then gets an empty record and is not charged.
  • Grounding is heuristic. A correct value phrased differently from the page can be unverified; a short common word can match by chance.
  • Sites that block crawlers may return empty pages; those pages are not charged.
  • Sequential processing. Large URL lists take proportionally longer; split them across parallel runs if speed matters.

Troubleshooting

SymptomLikely causeWhat to do
One row saying "No URL given"urls was emptyAdd page URLs
A record with every field nullThe page could not be fetched or extracted (see the log)Open the URL in a browser; try again later or check if the site blocks crawlers. The page was not charged
Many missing fieldsThe page does not contain those facts, or they are far down the pageCheck the page; use a more specific URL
_schema_valid: falseA required field is missingExpected when the page lacks it; make the field optional if that is acceptable
Numbers labelled unverifiedThe value was computed or reformatted by the modelCheck a few manually; the value may still be correct
Text labelled unverifiedOutput language differs from the page, or the model paraphrasedUse the page language, or treat as "review"
Run failed: "No valid schema provided"The schema was empty or unreadableUse one of the formats in Schema formats

FAQ

Do I need to write CSS selectors or code? No. You write field names and types. The model finds the values on the page.

Can it crawl a whole website? No. It extracts from the exact URLs you give it, one page each. For site-wide crawling, collect the URLs first with a crawler or sitemap tool, then pass them here.

Does it work on JavaScript-heavy pages? Pages are fetched by Website Content Crawler in adaptive browser mode, which renders JavaScript when needed.

Will it invent values? The model is instructed to use only page content and write null otherwise, and it runs at temperature 0. Because no model is perfect, every value is also checked against the page text and labelled high or unverified.

What happens with extra fields the model returns? They are dropped. The record contains only your schema fields, always in the same order and types.

Can it translate? Yes. Text values are written in the output language you choose, even when the page is in another language. Field names are never translated.

Am I charged for pages that failed? No. Pages that could not be fetched or extracted are not charged. The page fetch itself uses platform resources on your account.

How many URLs can I send? There is no fixed limit in the Actor. Pages are processed one after another within the run timeout; for large lists, split them across several runs.

Which number formats are understood? US and European separators (1,299.90, 1.299,90), thousands only (20,000, 1.234.567), decimal commas (12,5), and currency or unit symbols around the number.

Is my data stored? Only in your own Apify storage. Page text is sent to the language model provider solely to extract your fields.

Can I use it from an AI agent? Yes, through the Apify MCP server or the API.


What's new

  • Fair billing: pages that cannot be fetched or extracted are no longer charged.
  • Better number parsing: thousands separators such as 20,000 and 1.234.567 are read correctly.
  • Empty input is free: a run without URLs fetches nothing and explains what to fill in.

Trademarks

This Actor is not affiliated with, endorsed by, or connected to any website you extract data from. All product and company names are trademarks of their respective owners. Make sure your use of extracted data complies with the terms of the websites you process and with applicable law.