AI Web Scraper with Your Own OpenAI or Claude Key avatar

AI Web Scraper with Your Own OpenAI or Claude Key

Pricing

Pay per event

Go to Apify Store
AI Web Scraper with Your Own OpenAI or Claude Key

AI Web Scraper with Your Own OpenAI or Claude Key

Extract structured data from any web page with AI. List the fields you want in plain English or paste a JSON schema, and get clean JSON rows back. Uses your own OpenAI, Anthropic, Gemini or OpenRouter key: no token markup, $4 per 1,000 pages, failed pages free.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Joshua White

Joshua White

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Scrape structured data from any website with AI, without writing selectors. List the fields you want in plain English (or paste a JSON schema), give it URLs, and get clean, validated JSON rows: one per page, or one per item on a listing page. It runs on your own OpenAI, Anthropic (Claude), Google Gemini or OpenRouter key, so tokens cost what your provider charges, with no markup.

Try it free, no key needed: press Start on the prefilled input and you get a free preview of the cleaned page text the model would read. Add your key to extract the fields. Google Gemini's API has a free tier.

Status: the connection to each provider is tested against that provider's real API error responses (a wrong key or unknown model stops the run in seconds). The first live extraction run on a real provider key is still pending.

How to scrape a website with AI in 3 steps

  1. Paste Start URLs and list Fields to extract, one per line, for example
    price (number): current price, without the currency sign
    .
  2. Choose one row per page (product, article, profile) or one row per item (listings, search results).
  3. Pick your AI provider, paste your API key and click Start. Download JSON, CSV or Excel, or use the API.

Field types: text (default), number, integer, boolean, url, date, list. Or paste a JSON schema instead (nested objects and arrays work). Leave Model empty to use a cheap, fast model your key can use, or name any model you like.

How much does AI web scraping cost?

$4 per 1,000 pages extracted, plus $2 per 1,000 pages that need a browser, so Apify's $5 monthly free credit covers about 1,250 pages. Failed pages, blocked pages and replies that fail validation are free. AI tokens are billed by your provider on your own key; the run summary shows the exact token count. Compared with all-in AI scrapers at around $30 per 1,000 pages, you pay $4 plus your tokens. Set a maximum cost per run in the run options: the actor stops cleanly before starting a page it couldn't charge for, so it never spends your AI tokens on pages beyond your budget.

What you can use an AI web scraper for

  • Product pages → price lists. Name, price, stock, SKU and image from any shop, no matter how it's built.
  • Listings → rows. Every product in a category page, every job on a careers page, every event in a calendar.
  • Directories and profiles. Company name, address, phone, opening hours from business pages.
  • Articles. Headline, author, date, summary and tags from news or blog posts.
  • Messy one-off sites. The long tail of sites nobody has written a dedicated scraper for.

Input example

{
"startUrls": [{ "url": "https://books.toscrape.com/catalogue/category/books/poetry_23/index.html" }],
"fields": ["title", "price (number)", "url (url): link to the book page"],
"extractionMode": "list",
"llmProvider": "openai",
"apiKey": "sk-..."
}

Output example

One row per item (list mode). Your fields come first; meta tells you where each row came from and how many tokens the page used on your key. This example is illustrative, not from a live run: the model name and token counts are placeholders (this page is about 3,100 tokens of text).

{
"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
"title": "A Light in the Attic",
"price": 51.77,
"error": null,
"errorCode": null,
"meta": {
"pageUrl": "https://books.toscrape.com/catalogue/category/books/poetry_23/index.html",
"pageTitle": "Poetry | Books to Scrape - Sandbox",
"fetchedWith": "http",
"model": "gpt-6-luna",
"chunks": "1/1",
"inputTokens": 3600,
"outputTokens": 1200,
"itemIndex": 0,
"itemsOnPage": 19,
"scrapedAt": "2026-10-05T16:40:12.511Z"
}
}
  • url is the page URL, unless you asked for a url field, in which case it's the value the model found (the page URL is always in meta.pageUrl).
  • error, errorCode, meta and markdown are the actor's own columns, so don't use those as field names.
  • Pages that fail get a row with error and errorCode, and you are not charged for them: DNS, TIMEOUT, HTTP_BLOCKED, BOT_CHALLENGE, NOT_FOUND, ROBOTS_DISALLOWED, SELECTOR_NOT_FOUND, EMPTY_PAGE, NO_ITEMS, SCHEMA_MISMATCH, OUTPUT_TRUNCATED, RATE_LIMITED, CONTEXT_TOO_LONG, CONTENT_FILTERED.
  • The run's OUTPUT record has a summary: pages extracted and failed, items, and the total tokens used on your key, so you can work out your AI cost.

What makes it reliable

  • It checks your key before it starts. A wrong key, an unknown model or an empty account stops the run in seconds, with a plain message, before any page is loaded or any token is spent. If your credit runs out mid-run it stops too, instead of failing page after page.
  • Clean input for the model. Menus, footers, cookie banners and scripts are stripped; tables, lists, links and image URLs are kept; and the page's own structured data (JSON-LD, which often holds the exact price or date) is passed along. Less noise means fewer tokens and better answers.
  • Validated output. Every reply is checked against your schema. "12.99" becomes 12.99 for a number field, and a reply that doesn't fit gets one automatic repair before the page is marked as failed. You never get a half-broken row without being told.
  • Long pages are split into chunks within a token budget. With one row per page it stops reading as soon as every field is filled; with one row per item it reads every chunk and removes duplicate items.
  • JavaScript pages. Pages that come back empty or blocked over plain HTTP are opened in a real headless Chrome (browser mode Auto). Each page that needs the browser adds a small browser-render charge.
  • Rate limits from your provider are retried with backoff, honouring its Retry-After.

Which provider and model?

ProviderKey fromDefault model (if you leave Model empty)
OpenAIplatform.openai.com → API keysnewest luna / nano tier your key can use
Anthropicconsole.anthropic.com → API keysnewest Claude Haiku
Google Geminiaistudio.google.com → Get API key (has a free tier)newest Gemini Flash-Lite
OpenRouteropenrouter.ai → Keys (one key, hundreds of models)newest OpenAI luna tier

The defaults are auto-selected from the models your key can list; the names above describe the intended tier and haven't yet been seen live with a real key. Small models handle most extraction well. For tricky pages (lots of similar numbers, reasoning about what counts as an item), name a bigger model in Model. Any OpenAI-compatible service (Groq, Together, DeepSeek, Mistral, a self-hosted server) works too: choose OpenAI and set Base URL.

Tips for cheaper, better results

  • Use a CSS selector (e.g. #product or .results) to send only the part of the page you need. It's often 5–10× fewer tokens.
  • Write field descriptions like you'd brief a person: "price (number): the sale price if there is one, otherwise the regular price".
  • Run Preview mode first on a few URLs to see exactly what the model will read and how many tokens it is.
  • For listing pages with many items, raise Max output tokens if you see OUTPUT_TRUNCATED.

Your API key and data

  • The key field is a secret input: Apify stores it encrypted, and it never appears in logs, the dataset or the run summary. The actor sends it only to the provider you chose, over HTTPS with certificate checks, and never follows redirects with it.
  • Pages are sent to your AI provider for extraction under your account and its terms.

How it behaves on the web

  • It visits only the pages you give it (and, if you turn on link following, links on the same website that match your patterns, up to your limits).
  • It respects robots.txt: a disallowed page is skipped and reported as ROBOTS_DISALLOWED.
  • It identifies itself with the user agent AIWebScraperBot, visits each website at most 2 requests at a time, and backs off when a site asks it to.
  • No logins, nothing behind a paywall. Use it for public data, and only collect personal data if you have a lawful reason to.

Limitations

  • AI extraction is very good but not perfect: models can occasionally misread a page. Check a sample before you rely on a large run, and prefer a bigger model for hard pages.
  • In one-row-per-page mode the actor stops reading a long page once every field has a value, so "count everything on the page" fields should use one-row-per-item mode or a larger Max tokens per chunk.
  • Values that are only visible as images or CSS (for example star ratings drawn with icons) can't be read.
  • Sites that block all automated visitors can't be scraped (reported as HTTP_BLOCKED / BOT_CHALLENGE, not charged).

More tools from oldjard

FAQ

Is this Apify's own AI scraper? No. It's an independent actor. The difference: you use your own AI key, so tokens cost what your provider charges, and the per-page fee here is small.

Do I need to write code or selectors? No. Describe the fields in words. A CSS selector is optional, to save tokens.

Can I scrape many pages of a site? Yes. Turn on link following (depth 1–5) with an include pattern such as https://shop.example.com/products/*, or paste a list of URLs (for example from a sitemap).

Can I run it on a schedule? Yes. Save your input as a task and schedule it. The Monitoring field can make a run fail loudly if expected text stops appearing.

Feedback

Something extracted wrong or a site that doesn't work? Open an issue on the Issues tab with the URL, your fields and what you expected.