# AI Web Scraper with Your Own OpenAI or Claude Key (`oldjard/ai-web-scraper`) Actor

Extract structured data from any web page with AI. List the fields you want in plain English or paste a JSON schema, and get clean JSON rows back. Uses your own OpenAI, Anthropic, Gemini or OpenRouter key: no token markup, $4 per 1,000 pages, failed pages free.

- **URL**: https://apify.com/oldjard/ai-web-scraper.md
- **Developed by:** [Joshua White](https://apify.com/oldjard) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## AI Web Scraper with Your Own OpenAI or Claude Key

**Scrape structured data from any website with AI**, without writing selectors. List the fields you want in plain
English (or paste a JSON schema), give it URLs, and get clean, validated JSON rows: one per page, or one per item on a
listing page. It runs on **your own OpenAI, Anthropic (Claude), Google Gemini or OpenRouter key**, so tokens cost
what your provider charges, with **no markup**.

**Try it free, no key needed:** press **Start** on the prefilled input and you get a free preview of the cleaned page
text the model would read. Add your key to extract the fields. Google Gemini's API has a free tier.

**Status:** the connection to each provider is tested against that provider's real API error responses (a wrong key
or unknown model stops the run in seconds). The first live extraction run on a real provider key is still pending.

### How to scrape a website with AI in 3 steps

1. Paste **Start URLs** and list **Fields to extract**, one per line, for example `price (number): current price,
   without the currency sign`.
2. Choose **one row per page** (product, article, profile) or **one row per item** (listings, search results).
3. Pick your **AI provider**, paste your **API key** and click **Start**. Download JSON, CSV or Excel, or use the API.

Field types: text (default), `number`, `integer`, `boolean`, `url`, `date`, `list`. Or paste a **JSON schema**
instead (nested objects and arrays work). Leave **Model** empty to use a cheap, fast model your key can use, or name
any model you like.

### How much does AI web scraping cost?

**$4 per 1,000 pages** extracted, plus **$2 per 1,000** pages that need a browser, so Apify's $5 monthly free credit covers about 1,250 pages. Failed pages, blocked pages and
replies that fail validation are free. AI tokens are billed by your provider on your own key; the run summary shows
the exact token count. Compared with all-in AI scrapers at around $30 per 1,000 pages, you pay $4 plus your tokens.
Set a **maximum cost per run** in the run options: the actor stops cleanly before starting a page it couldn't charge
for, so it never spends your AI tokens on pages beyond your budget.

### What you can use an AI web scraper for

- **Product pages → price lists.** Name, price, stock, SKU and image from any shop, no matter how it's built.
- **Listings → rows.** Every product in a category page, every job on a careers page, every event in a calendar.
- **Directories and profiles.** Company name, address, phone, opening hours from business pages.
- **Articles.** Headline, author, date, summary and tags from news or blog posts.
- **Messy one-off sites.** The long tail of sites nobody has written a dedicated scraper for.

### Input example

```json
{
  "startUrls": [{ "url": "https://books.toscrape.com/catalogue/category/books/poetry_23/index.html" }],
  "fields": ["title", "price (number)", "url (url): link to the book page"],
  "extractionMode": "list",
  "llmProvider": "openai",
  "apiKey": "sk-..."
}
```

### Output example

One row per item (list mode). Your fields come first; `meta` tells you where each row came from and how many tokens
the page used on your key. **This example is illustrative, not from a live run:** the model name and token counts
are placeholders (this page is about 3,100 tokens of text).

```json
{
  "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
  "title": "A Light in the Attic",
  "price": 51.77,
  "error": null,
  "errorCode": null,
  "meta": {
    "pageUrl": "https://books.toscrape.com/catalogue/category/books/poetry_23/index.html",
    "pageTitle": "Poetry | Books to Scrape - Sandbox",
    "fetchedWith": "http",
    "model": "gpt-6-luna",
    "chunks": "1/1",
    "inputTokens": 3600,
    "outputTokens": 1200,
    "itemIndex": 0,
    "itemsOnPage": 19,
    "scrapedAt": "2026-10-05T16:40:12.511Z"
  }
}
```

- `url` is the page URL, unless you asked for a `url` field, in which case it's the value the model found (the page
  URL is always in `meta.pageUrl`).
- `error`, `errorCode`, `meta` and `markdown` are the actor's own columns, so don't use those as field names.
- **Pages that fail get a row with `error` and `errorCode`**, and you are **not charged** for them: `DNS`,
  `TIMEOUT`, `HTTP_BLOCKED`, `BOT_CHALLENGE`, `NOT_FOUND`, `ROBOTS_DISALLOWED`, `SELECTOR_NOT_FOUND`, `EMPTY_PAGE`,
  `NO_ITEMS`, `SCHEMA_MISMATCH`, `OUTPUT_TRUNCATED`, `RATE_LIMITED`, `CONTEXT_TOO_LONG`, `CONTENT_FILTERED`.
- The run's `OUTPUT` record has a summary: pages extracted and failed, items, and the **total tokens used on your
  key**, so you can work out your AI cost.

### What makes it reliable

- **It checks your key before it starts.** A wrong key, an unknown model or an empty account stops the run in seconds,
  with a plain message, before any page is loaded or any token is spent. If your credit runs out mid-run it stops
  too, instead of failing page after page.
- **Clean input for the model.** Menus, footers, cookie banners and scripts are stripped; tables, lists, links and
  image URLs are kept; and the page's own structured data (JSON-LD, which often holds the exact price or date) is
  passed along. Less noise means fewer tokens and better answers.
- **Validated output.** Every reply is checked against your schema. `"12.99"` becomes `12.99` for a number field, and
  a reply that doesn't fit gets **one automatic repair** before the page is marked as failed. You never get a
  half-broken row without being told.
- **Long pages are split** into chunks within a token budget. With one row per page it stops reading as soon as every
  field is filled; with one row per item it reads every chunk and removes duplicate items.
- **JavaScript pages.** Pages that come back empty or blocked over plain HTTP are opened in a real headless Chrome
  (browser mode *Auto*). Each page that needs the browser adds a small *browser-render* charge.
- **Rate limits** from your provider are retried with backoff, honouring its `Retry-After`.

### Which provider and model?

| Provider | Key from | Default model (if you leave Model empty) |
|---|---|---|
| OpenAI | platform.openai.com → API keys | newest *luna* / *nano* tier your key can use |
| Anthropic | console.anthropic.com → API keys | newest Claude Haiku |
| Google Gemini | aistudio.google.com → Get API key (has a free tier) | newest Gemini Flash-Lite |
| OpenRouter | openrouter.ai → Keys (one key, hundreds of models) | newest OpenAI luna tier |

The defaults are auto-selected from the models your key can list; the names above describe the intended tier and
haven't yet been seen live with a real key. Small models handle most extraction well. For tricky pages (lots of similar numbers, reasoning about what counts as
an item), name a bigger model in **Model**. Any **OpenAI-compatible** service (Groq, Together, DeepSeek, Mistral, a
self-hosted server) works too: choose OpenAI and set **Base URL**.

### Tips for cheaper, better results

- Use a **CSS selector** (e.g. `#product` or `.results`) to send only the part of the page you need. It's often 5–10×
  fewer tokens.
- Write **field descriptions** like you'd brief a person: "price (number): the sale price if there is one, otherwise
  the regular price".
- Run **Preview mode** first on a few URLs to see exactly what the model will read and how many tokens it is.
- For listing pages with many items, raise **Max output tokens** if you see `OUTPUT_TRUNCATED`.

### Your API key and data

- The key field is a **secret input**: Apify stores it encrypted, and it never appears in logs, the dataset or the run
  summary. The actor sends it only to the provider you chose, over HTTPS with certificate checks, and never follows
  redirects with it.
- Pages are sent to your AI provider for extraction under your account and its terms.

### How it behaves on the web

- It visits only the pages you give it (and, if you turn on link following, links on the **same website** that match
  your patterns, up to your limits).
- It **respects robots.txt**: a disallowed page is skipped and reported as `ROBOTS_DISALLOWED`.
- It identifies itself with the user agent `AIWebScraperBot`, visits each website at most 2 requests at a time, and
  backs off when a site asks it to.
- No logins, nothing behind a paywall. Use it for public data, and only collect personal data if you have a lawful
  reason to.

### Limitations

- AI extraction is very good but not perfect: models can occasionally misread a page. Check a sample before you rely
  on a large run, and prefer a bigger model for hard pages.
- In one-row-per-page mode the actor stops reading a long page once every field has a value, so "count everything on
  the page" fields should use one-row-per-item mode or a larger **Max tokens per chunk**.
- Values that are only visible as images or CSS (for example star ratings drawn with icons) can't be read.
- Sites that block all automated visitors can't be scraped (reported as `HTTP_BLOCKED` / `BOT_CHALLENGE`, not charged).

### More tools from oldjard

- [Tech Stack Detector](https://apify.com/oldjard/tech-stack-detector): what any list of websites is built with.
- [Sitemap URL Extractor](https://apify.com/oldjard/sitemap-url-extractor): every URL on a website, for RAG and SEO.
- [Shopify Products Scraper & Price Monitor](https://apify.com/oldjard/shopify-products-price-monitor): catalogs and price changes from any Shopify store.
- [Workday, Greenhouse, Lever & Ashby Jobs Scraper](https://apify.com/oldjard/ats-career-site-jobs): every open job from company career sites.
- [Bulk Website Screenshot & URL to PDF](https://apify.com/oldjard/screenshot-pdf): screenshots and PDFs of any list of pages.
- [Website Change Monitor](https://apify.com/oldjard/website-change-monitor): a before/after diff by webhook, Slack or Discord when a page changes.
- [Company Registry Lookup](https://apify.com/oldjard/company-registry-lookup): UK Companies House, Spain, France, Finland and Norway in one schema.
- [UK & EU Public Tenders](https://apify.com/oldjard/uk-eu-public-tenders): Find a Tender and TED notices in one table, with daily only-new alerts.

### FAQ

**Is this Apify's own AI scraper?** No. It's an independent actor. The difference: you use your own AI key, so tokens
cost what your provider charges, and the per-page fee here is small.

**Do I need to write code or selectors?** No. Describe the fields in words. A CSS selector is optional, to save tokens.

**Can I scrape many pages of a site?** Yes. Turn on link following (depth 1–5) with an include pattern such as
`https://shop.example.com/products/*`, or paste a list of URLs (for example from a sitemap).

**Can I run it on a schedule?** Yes. Save your input as a task and schedule it. The *Monitoring* field can make a run
fail loudly if expected text stops appearing.

### Feedback

Something extracted wrong or a site that doesn't work? Open an issue on the Issues tab with the URL, your fields and
what you expected.

# Actor input Schema

## `startUrls` (type: `array`):

The web pages to extract data from, one per line. Paste as many as you like; duplicates are removed. Turn on link following below to also visit pages linked from these.

## `fields` (type: `array`):

One field per line, in plain English: <code>name</code>, <code>name: what it is</code> or <code>name (type): what it is</code>. Types: text (default), number, integer, boolean, url, date, list. Example: <code>price (number): the current price, without the currency sign</code>. Each field becomes a column. Leave empty if you paste a JSON schema below.

## `jsonSchema` (type: `object`):

Instead of the field list, paste a JSON Schema for one record (<code>{"type": "object", "properties": {...}}</code>). A schema with <code>"type": "array"</code> switches on list mode. Nested objects and arrays are supported. Overrides the field list.

## `instructions` (type: `string`):

Extra guidance for the model, e.g. "Prices in USD. Skip sponsored listings." With no fields and no schema, describe what you want here and the model picks the field names (less consistent across pages).

## `extractionMode` (type: `string`):

<b>One row per page</b> fills your fields once per page. <b>One row per item</b> finds every matching item on the page (each product in a category page, each job in a list) and returns one row per item.

## `llmProvider` (type: `string`):

Whose model reads the pages. Any OpenAI-compatible service (Groq, Together, DeepSeek, Mistral, a self-hosted server) works too: pick OpenAI and set the Base URL below.

## `apiKey` (type: `string`):

Your API key for the provider above. Without a key the actor runs a free preview of the first 3 pages (the cleaned text the model would read, no extracted fields).

## `model` (type: `string`):

Leave empty to use the newest cheap, fast model your key can use (for example OpenAI's luna or nano tier, Claude Haiku, Gemini Flash-Lite). Or name any model your key can use, e.g. <code>gpt-6-sol</code>, <code>claude-sonnet-5-5</code>, <code>gemini-3.8-flash</code>, or for OpenRouter <code>anthropic/claude-sonnet-5.5</code>.

## `baseUrl` (type: `string`):

Only for OpenAI-compatible services or proxies, e.g. <code>https://api.groq.com/openai/v1</code>. Must be https. Leave empty for the provider's own API.

## `renderMode` (type: `string`):

<b>Auto</b> (default) loads each page with a fast HTTP request and opens it in headless Chrome only if it is blocked or is an empty JavaScript app. Each page opened in Chrome adds one <i>browser-render</i> charge. For runs with many browser pages, give the run 2 GB of memory or more.

## `contentScope` (type: `string`):

What the model reads. <b>Main content</b> saves tokens and keeps the model focused. Choose <b>Whole page</b> if a field lives in the header, footer or menu.

## `cssSelector` (type: `string`):

Only read the part of the page that matches this CSS selector, e.g. <code>#product</code> or <code>.job-listing</code>. Cuts tokens a lot on big pages. Pages where it matches nothing are reported and not charged.

## `includeLinks` (type: `boolean`):

Give the model the URL of every link, so it can fill URL fields. Turn off to save tokens if you need no URLs.

## `includeImages` (type: `boolean`):

Give the model image URLs and alt text, so it can fill image fields.

## `maxCrawlDepth` (type: `integer`):

0 = only the start URLs. 1 = also the pages they link to (same website only), and so on. Use the include pattern below to stay on the pages you want, e.g. product pages.

## `maxPagesPerStartUrl` (type: `integer`):

Stop following links from a start URL after this many pages (the start page counts).

## `linkIncludeGlobs` (type: `array`):

Glob patterns over the full URL; <code>*</code> matches within one path segment and <code>\*\*</code> matches anything. Example: <code>https://shop.example.com/products/*</code>. Empty = every same-site link.

## `linkExcludeGlobs` (type: `array`):

Glob patterns for links to skip, e.g. <code>**/login\*</code> or <code>**?sort=\*</code>.

## `maxPages` (type: `integer`):

Stop after this many pages across the whole run (0 = no limit). You can also set a maximum cost in the run options.

## `maxTokensPerChunk` (type: `integer`):

Pages longer than this (estimated) are split into chunks, each sent to the model separately and the results merged. Lower it to cut cost per call; raise it for models with big context windows.

## `maxChunksPerPage` (type: `integer`):

Read at most this many chunks of a very long page. With one row per page the actor also stops early once every field is filled.

## `maxOutputTokens` (type: `integer`):

The most tokens the model may write per call (reasoning models count their thinking here too). Raise it for listing pages with many items.

## `maxConcurrency` (type: `integer`):

How many pages are processed at once. Lower it if your provider answers with rate-limit errors. Each website is still visited politely, at most 2 requests at a time.

## `requestTimeoutSecs` (type: `integer`):

Give up on a page that hasn't answered within this many seconds. Reported as a timeout and not charged.

## `includeMarkdown` (type: `boolean`):

Add a <code>#markdown</code> column with the cleaned page text the model read. Handy for checking results.

## `dryRun` (type: `boolean`):

Convert the pages to clean markdown and estimate their tokens without calling the model or needing a key. Use it to tune the CSS selector and token settings before a big run. Charged per page like extraction.

## `canaryExpectations` (type: `object`):

Optional, for scheduled health checks: a map of URL to text that must appear in that page's result, e.g. <code>{"https://example.com/": \["Example Domain"]}</code>. The run fails if more than 20% of the checks miss.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
    }
  ],
  "fields": [
    "title: the book title",
    "price (number): price in GBP, without the £ sign",
    "inStock (boolean): whether it is in stock",
    "copiesAvailable (integer): how many copies are available",
    "upc: the UPC code from the product information table"
  ],
  "extractionMode": "single",
  "llmProvider": "openai",
  "renderMode": "auto",
  "contentScope": "auto",
  "includeLinks": true,
  "includeImages": true,
  "maxCrawlDepth": 0,
  "maxPagesPerStartUrl": 10,
  "maxPages": 0,
  "maxTokensPerChunk": 8000,
  "maxChunksPerPage": 3,
  "maxOutputTokens": 8000,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 30,
  "includeMarkdown": false,
  "dryRun": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
        }
    ],
    "fields": [
        "title: the book title",
        "price (number): price in GBP, without the £ sign",
        "inStock (boolean): whether it is in stock",
        "copiesAvailable (integer): how many copies are available",
        "upc: the UPC code from the product information table"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("oldjard/ai-web-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html" }],
    "fields": [
        "title: the book title",
        "price (number): price in GBP, without the £ sign",
        "inStock (boolean): whether it is in stock",
        "copiesAvailable (integer): how many copies are available",
        "upc: the UPC code from the product information table",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("oldjard/ai-web-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
    }
  ],
  "fields": [
    "title: the book title",
    "price (number): price in GBP, without the £ sign",
    "inStock (boolean): whether it is in stock",
    "copiesAvailable (integer): how many copies are available",
    "upc: the UPC code from the product information table"
  ]
}' |
apify call oldjard/ai-web-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,oldjard/ai-web-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/YmEIabGqbxx7OcIA6/builds/vE8UwxgG7vsbEhSwn/openapi.json
