AI Web Scraper with Your Own OpenAI or Claude Key
Pricing
Pay per event
AI Web Scraper with Your Own OpenAI or Claude Key
Extract structured data from any web page with AI. List the fields you want in plain English or paste a JSON schema, and get clean JSON rows back. Uses your own OpenAI, Anthropic, Gemini or OpenRouter key: no token markup, $4 per 1,000 pages, failed pages free.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Joshua White
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Scrape structured data from any website with AI, without writing selectors. List the fields you want in plain English (or paste a JSON schema), give it URLs, and get clean, validated JSON rows: one per page, or one per item on a listing page. It runs on your own OpenAI, Anthropic (Claude), Google Gemini or OpenRouter key, so tokens cost what your provider charges, with no markup.
Try it free, no key needed: press Start on the prefilled input and you get a free preview of the cleaned page text the model would read. Add your key to extract the fields. Google Gemini's API has a free tier.
Status: the connection to each provider is tested against that provider's real API error responses (a wrong key or unknown model stops the run in seconds). The first live extraction run on a real provider key is still pending.
How to scrape a website with AI in 3 steps
- Paste Start URLs and list Fields to extract, one per line, for example .price (number): current price, without the currency sign
- Choose one row per page (product, article, profile) or one row per item (listings, search results).
- Pick your AI provider, paste your API key and click Start. Download JSON, CSV or Excel, or use the API.
Field types: text (default), number, integer, boolean, url, date, list. Or paste a JSON schema
instead (nested objects and arrays work). Leave Model empty to use a cheap, fast model your key can use, or name
any model you like.
How much does AI web scraping cost?
$4 per 1,000 pages extracted, plus $2 per 1,000 pages that need a browser, so Apify's $5 monthly free credit covers about 1,250 pages. Failed pages, blocked pages and replies that fail validation are free. AI tokens are billed by your provider on your own key; the run summary shows the exact token count. Compared with all-in AI scrapers at around $30 per 1,000 pages, you pay $4 plus your tokens. Set a maximum cost per run in the run options: the actor stops cleanly before starting a page it couldn't charge for, so it never spends your AI tokens on pages beyond your budget.
What you can use an AI web scraper for
- Product pages → price lists. Name, price, stock, SKU and image from any shop, no matter how it's built.
- Listings → rows. Every product in a category page, every job on a careers page, every event in a calendar.
- Directories and profiles. Company name, address, phone, opening hours from business pages.
- Articles. Headline, author, date, summary and tags from news or blog posts.
- Messy one-off sites. The long tail of sites nobody has written a dedicated scraper for.
Input example
{"startUrls": [{ "url": "https://books.toscrape.com/catalogue/category/books/poetry_23/index.html" }],"fields": ["title", "price (number)", "url (url): link to the book page"],"extractionMode": "list","llmProvider": "openai","apiKey": "sk-..."}
Output example
One row per item (list mode). Your fields come first; meta tells you where each row came from and how many tokens
the page used on your key. This example is illustrative, not from a live run: the model name and token counts
are placeholders (this page is about 3,100 tokens of text).
{"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html","title": "A Light in the Attic","price": 51.77,"error": null,"errorCode": null,"meta": {"pageUrl": "https://books.toscrape.com/catalogue/category/books/poetry_23/index.html","pageTitle": "Poetry | Books to Scrape - Sandbox","fetchedWith": "http","model": "gpt-6-luna","chunks": "1/1","inputTokens": 3600,"outputTokens": 1200,"itemIndex": 0,"itemsOnPage": 19,"scrapedAt": "2026-10-05T16:40:12.511Z"}}
urlis the page URL, unless you asked for aurlfield, in which case it's the value the model found (the page URL is always inmeta.pageUrl).error,errorCode,metaandmarkdownare the actor's own columns, so don't use those as field names.- Pages that fail get a row with
erroranderrorCode, and you are not charged for them:DNS,TIMEOUT,HTTP_BLOCKED,BOT_CHALLENGE,NOT_FOUND,ROBOTS_DISALLOWED,SELECTOR_NOT_FOUND,EMPTY_PAGE,NO_ITEMS,SCHEMA_MISMATCH,OUTPUT_TRUNCATED,RATE_LIMITED,CONTEXT_TOO_LONG,CONTENT_FILTERED. - The run's
OUTPUTrecord has a summary: pages extracted and failed, items, and the total tokens used on your key, so you can work out your AI cost.
What makes it reliable
- It checks your key before it starts. A wrong key, an unknown model or an empty account stops the run in seconds, with a plain message, before any page is loaded or any token is spent. If your credit runs out mid-run it stops too, instead of failing page after page.
- Clean input for the model. Menus, footers, cookie banners and scripts are stripped; tables, lists, links and image URLs are kept; and the page's own structured data (JSON-LD, which often holds the exact price or date) is passed along. Less noise means fewer tokens and better answers.
- Validated output. Every reply is checked against your schema.
"12.99"becomes12.99for a number field, and a reply that doesn't fit gets one automatic repair before the page is marked as failed. You never get a half-broken row without being told. - Long pages are split into chunks within a token budget. With one row per page it stops reading as soon as every field is filled; with one row per item it reads every chunk and removes duplicate items.
- JavaScript pages. Pages that come back empty or blocked over plain HTTP are opened in a real headless Chrome (browser mode Auto). Each page that needs the browser adds a small browser-render charge.
- Rate limits from your provider are retried with backoff, honouring its
Retry-After.
Which provider and model?
| Provider | Key from | Default model (if you leave Model empty) |
|---|---|---|
| OpenAI | platform.openai.com → API keys | newest luna / nano tier your key can use |
| Anthropic | console.anthropic.com → API keys | newest Claude Haiku |
| Google Gemini | aistudio.google.com → Get API key (has a free tier) | newest Gemini Flash-Lite |
| OpenRouter | openrouter.ai → Keys (one key, hundreds of models) | newest OpenAI luna tier |
The defaults are auto-selected from the models your key can list; the names above describe the intended tier and haven't yet been seen live with a real key. Small models handle most extraction well. For tricky pages (lots of similar numbers, reasoning about what counts as an item), name a bigger model in Model. Any OpenAI-compatible service (Groq, Together, DeepSeek, Mistral, a self-hosted server) works too: choose OpenAI and set Base URL.
Tips for cheaper, better results
- Use a CSS selector (e.g.
#productor.results) to send only the part of the page you need. It's often 5–10× fewer tokens. - Write field descriptions like you'd brief a person: "price (number): the sale price if there is one, otherwise the regular price".
- Run Preview mode first on a few URLs to see exactly what the model will read and how many tokens it is.
- For listing pages with many items, raise Max output tokens if you see
OUTPUT_TRUNCATED.
Your API key and data
- The key field is a secret input: Apify stores it encrypted, and it never appears in logs, the dataset or the run summary. The actor sends it only to the provider you chose, over HTTPS with certificate checks, and never follows redirects with it.
- Pages are sent to your AI provider for extraction under your account and its terms.
How it behaves on the web
- It visits only the pages you give it (and, if you turn on link following, links on the same website that match your patterns, up to your limits).
- It respects robots.txt: a disallowed page is skipped and reported as
ROBOTS_DISALLOWED. - It identifies itself with the user agent
AIWebScraperBot, visits each website at most 2 requests at a time, and backs off when a site asks it to. - No logins, nothing behind a paywall. Use it for public data, and only collect personal data if you have a lawful reason to.
Limitations
- AI extraction is very good but not perfect: models can occasionally misread a page. Check a sample before you rely on a large run, and prefer a bigger model for hard pages.
- In one-row-per-page mode the actor stops reading a long page once every field has a value, so "count everything on the page" fields should use one-row-per-item mode or a larger Max tokens per chunk.
- Values that are only visible as images or CSS (for example star ratings drawn with icons) can't be read.
- Sites that block all automated visitors can't be scraped (reported as
HTTP_BLOCKED/BOT_CHALLENGE, not charged).
More tools from oldjard
- Tech Stack Detector: what any list of websites is built with.
- Sitemap URL Extractor: every URL on a website, for RAG and SEO.
- Shopify Products Scraper & Price Monitor: catalogs and price changes from any Shopify store.
- Workday, Greenhouse, Lever & Ashby Jobs Scraper: every open job from company career sites.
- Bulk Website Screenshot & URL to PDF: screenshots and PDFs of any list of pages.
- Website Change Monitor: a before/after diff by webhook, Slack or Discord when a page changes.
- Company Registry Lookup: UK Companies House, Spain, France, Finland and Norway in one schema.
- UK & EU Public Tenders: Find a Tender and TED notices in one table, with daily only-new alerts.
FAQ
Is this Apify's own AI scraper? No. It's an independent actor. The difference: you use your own AI key, so tokens cost what your provider charges, and the per-page fee here is small.
Do I need to write code or selectors? No. Describe the fields in words. A CSS selector is optional, to save tokens.
Can I scrape many pages of a site? Yes. Turn on link following (depth 1–5) with an include pattern such as
https://shop.example.com/products/*, or paste a list of URLs (for example from a sitemap).
Can I run it on a schedule? Yes. Save your input as a task and schedule it. The Monitoring field can make a run fail loudly if expected text stops appearing.
Feedback
Something extracted wrong or a site that doesn't work? Open an issue on the Issues tab with the URL, your fields and what you expected.