AI Web Data Extractor (URL to JSON)
Pricing
from $8.00 / 1,000 extracted pages
AI Web Data Extractor (URL to JSON)
Give it URLs and the fields you want; AI reads each page and returns clean JSON. No selectors or code. Works on any site, blocked pages are free. $10 per 1,000 pages.
Pricing
from $8.00 / 1,000 extracted pages
Rating
0.0
(0)
Developer
Yukai Lin
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
What does AI Web Data Extractor do?
Give it a list of URLs and the fields you want. AI reads each page and returns clean, structured JSON: no CSS selectors, no XPath, no code, and it keeps working when the site's layout changes.
{ "product_name": "string", "price": "number", "in_stock": "boolean" }
becomes, for every page:
{ "product_name": "Aurora Desk Lamp", "price": 49.9, "in_stock": true }
- 🧠 Describe, don't code: list field names and types, add plain-language instructions if needed
- 🌐 Any website: product pages, job posts, real estate listings, company "About" pages, articles, events, docs
- 🧾 Honest output: values that are not on the page come back as
null; the AI is told never to invent data - 🧩 Nested data: supply a full JSON Schema for arrays and objects
- 🛡️ Pay only for success: blocked pages (403, bot checks), errors and empty pages are free
- ⚡ Auto mode: pages are read over plain HTTP and rendered in a real browser only when needed
How much does it cost?
| Event | Price |
|---|---|
| Page extracted | $10 per 1,000 pages ($0.01 per page) |
That's about 3× cheaper than Apify's own AI Web Scraper ($0.03 per page), with no extra charge for browser rendering or compute. Your maximum charge limit is always respected.
How to use it
- Paste your URLs.
- In Fields to extract, list what you want, e.g.
{"company_name": "string", "founded_year": "integer", "headquarters": "string"}. - Optional: add Instructions, e.g. "Prices in USD as numbers, without currency symbols".
- Click Start. Each URL produces one row with a
dataobject.
Input example
{"startUrls": [{ "url": "https://github.com/apify/crawlee" }],"fields": { "name": "string", "description": "string", "license": "string", "primary_language": "string" }}
Output example
A real result for the input above:
{"url": "https://github.com/apify/crawlee","title": "GitHub - apify/crawlee: Crawlee—A web scraping and browser automation library...","success": true,"data": {"name": "crawlee","description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers...","license": "Apache License 2.0","primary_language": "JavaScript"},"mode": "fast","contentTruncated": false,"extractedAt": "2026-09-29T04:00:00.000Z"}
Advanced: JSON Schema
For nested results, use JSON Schema instead of Fields:
{"type": "object","properties": {"jobs": {"type": "array","items": {"type": "object","properties": { "title": { "type": "string" }, "location": { "type": "string" }, "remote": { "type": "boolean" } }}}},"required": ["jobs"]}
Tips for accurate results
- Be specific with field names:
newest_stable_versionworks better thanversion. - Use Instructions for units, formats and which item to pick when a page lists several.
- Long pages: set a Content selector (e.g.
main) so the AI reads the relevant part; very long pages are truncated (contentTruncated: true). - Many similar pages: combine with our Website to Markdown Crawler to collect URLs first.
Use it from AI agents and code
Works as a tool for AI agents (Apify MCP server, LangChain, LlamaIndex) and via API:
curl -X POST "https://api.apify.com/v2/acts/tidytools~ai-web-data-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"startUrls":[{"url":"https://example.com"}],"fields":{"title":"string","summary":"string"}}'
Limitations
- AI extraction is very good but not perfect; verify critical values.
- Only public
http/httpspages; no logins. - Up to about 8,000 tokens of each page are read.
Is it legal?
You are responsible for how you use extracted data. Respect the target site's terms of service, copyright and privacy laws, and don't collect personal data without a legal basis.
Support
Open an issue in the Issues tab with the URL and your input.