AI Web Data Extractor (URL to JSON) avatar

AI Web Data Extractor (URL to JSON)

Pricing

from $8.00 / 1,000 extracted pages

Go to Apify Store
AI Web Data Extractor (URL to JSON)

AI Web Data Extractor (URL to JSON)

Give it URLs and the fields you want; AI reads each page and returns clean JSON. No selectors or code. Works on any site, blocked pages are free. $10 per 1,000 pages.

Pricing

from $8.00 / 1,000 extracted pages

Rating

0.0

(0)

Developer

Yukai Lin

Yukai Lin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

What does AI Web Data Extractor do?

Give it a list of URLs and the fields you want. AI reads each page and returns clean, structured JSON: no CSS selectors, no XPath, no code, and it keeps working when the site's layout changes.

{ "product_name": "string", "price": "number", "in_stock": "boolean" }

becomes, for every page:

{ "product_name": "Aurora Desk Lamp", "price": 49.9, "in_stock": true }
  • 🧠 Describe, don't code: list field names and types, add plain-language instructions if needed
  • 🌐 Any website: product pages, job posts, real estate listings, company "About" pages, articles, events, docs
  • 🧾 Honest output: values that are not on the page come back as null; the AI is told never to invent data
  • 🧩 Nested data: supply a full JSON Schema for arrays and objects
  • 🛡️ Pay only for success: blocked pages (403, bot checks), errors and empty pages are free
  • ⚡ Auto mode: pages are read over plain HTTP and rendered in a real browser only when needed

How much does it cost?

EventPrice
Page extracted$10 per 1,000 pages ($0.01 per page)

That's about 3× cheaper than Apify's own AI Web Scraper ($0.03 per page), with no extra charge for browser rendering or compute. Your maximum charge limit is always respected.

How to use it

  1. Paste your URLs.
  2. In Fields to extract, list what you want, e.g. {"company_name": "string", "founded_year": "integer", "headquarters": "string"}.
  3. Optional: add Instructions, e.g. "Prices in USD as numbers, without currency symbols".
  4. Click Start. Each URL produces one row with a data object.

Input example

{
"startUrls": [{ "url": "https://github.com/apify/crawlee" }],
"fields": { "name": "string", "description": "string", "license": "string", "primary_language": "string" }
}

Output example

A real result for the input above:

{
"url": "https://github.com/apify/crawlee",
"title": "GitHub - apify/crawlee: Crawlee—A web scraping and browser automation library...",
"success": true,
"data": {
"name": "crawlee",
"description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers...",
"license": "Apache License 2.0",
"primary_language": "JavaScript"
},
"mode": "fast",
"contentTruncated": false,
"extractedAt": "2026-09-29T04:00:00.000Z"
}

Advanced: JSON Schema

For nested results, use JSON Schema instead of Fields:

{
"type": "object",
"properties": {
"jobs": {
"type": "array",
"items": {
"type": "object",
"properties": { "title": { "type": "string" }, "location": { "type": "string" }, "remote": { "type": "boolean" } }
}
}
},
"required": ["jobs"]
}

Tips for accurate results

  • Be specific with field names: newest_stable_version works better than version.
  • Use Instructions for units, formats and which item to pick when a page lists several.
  • Long pages: set a Content selector (e.g. main) so the AI reads the relevant part; very long pages are truncated (contentTruncated: true).
  • Many similar pages: combine with our Website to Markdown Crawler to collect URLs first.

Use it from AI agents and code

Works as a tool for AI agents (Apify MCP server, LangChain, LlamaIndex) and via API:

curl -X POST "https://api.apify.com/v2/acts/tidytools~ai-web-data-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls":[{"url":"https://example.com"}],"fields":{"title":"string","summary":"string"}}'

Limitations

  • AI extraction is very good but not perfect; verify critical values.
  • Only public http/https pages; no logins.
  • Up to about 8,000 tokens of each page are read.

You are responsible for how you use extracted data. Respect the target site's terms of service, copyright and privacy laws, and don't collect personal data without a legal basis.

Support

Open an issue in the Issues tab with the URL and your input.