# Web Scraping Extractor Data API - Import.io Crawl Runs, Reports (`nabeelbaghoor/web-scraping-extractor-data-api`) Actor

Export the output of your own Import.io web data extractors: every extracted row of the latest finished or any crawl run, crawl run history and stats, data and change report rows, extractors, inputs, plan and usage. Read only: nothing is started or changed. Bring your own key.

- **URL**: https://apify.com/nabeelbaghoor/web-scraping-extractor-data-api.md
- **Developed by:** [Nabeel Hassan](https://apify.com/nabeelbaghoor) (community)
- **Categories:** Developer tools, E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 extractor data record returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Web Scraping Extractor Data API - Import.io Crawl Runs, Reports

Export what your own Import.io web data extractors have already collected: every extracted row of an extractor's latest finished crawl run or of any past crawl run, crawl run history with URL and row counts, data and change report rows, and your extractors and their inputs, as clean rows ready for a warehouse, a spreadsheet or the next step of a pipeline.

### What it collects

- **Extracted rows**: the newest finished crawl run of an extractor, or any crawl run by id, one row per extracted row with every column under `data`, keyed by the extractor's own field names (product name, price, availability, rating, URL, whatever you trained it to capture).
- **Crawl run history**: recent crawl runs across the account or per extractor, with state (finished, started, pending, cancelled, failed), start and stop time, total, successful and failed URLs, rows, queries used, screen captures and proxy traffic. Filter by state.
- **Report rows**: the CSV rows of any data report run, or the JSON rows of a change report run, which is where price and availability changes between crawl runs show up.
- **Reports and report runs**: every data and change report, and each run's status, summary and inputs.
- **Extractors and inputs**: every extractor with its fields, tags, chaining and report settings, and each extractor's current URLs or inputs. Webhook headers, credentials and any field that looks like a secret are left out.
- **Plan and usage**: queries used and allowed in the current period, plan, renewal and expiry.
- Read only, pay per result, bring your own key.

### Input

| Field | What it does |
| --- | --- |
| What to read | Extractors (default, needs nothing else), latest result, crawl runs, result of a crawl run, crawl run details, extractor details, extractor inputs, reports, report runs, rows of a report run, report run details, or plan and usage. |
| Ids | One per line: extractor ids, crawl run ids, report ids or report run ids, depending on the service. |
| Crawl run state | Crawl runs: keep only finished, running, pending, cancelled or failed runs. |
| Order | Lists: newest first (default) or oldest first. |
| Report file | Rows of a report run: CSV (any report) or JSON (change reports only). |
| Maximum results | Row cap for the run. |
| Requests per minute | Pacing for calls to the provider. |
| API key | Your own API key, as a secret input. |

A typical first run is the default extractors service. Copy an extractor id from it, then run latest result with that id to get the data.

### FAQ

#### What is this web scraping extractor data API used for?

Getting data out of extractors that already run on a schedule and into the systems that use it. A pricing team pulls the latest competitor product and price crawl every morning into BigQuery or Snowflake. An ecommerce analyst exports a change report run to see which products changed price or went out of stock. An operations team reviews failed crawl runs and failed URL counts per extractor. A finance team checks queries used against the plan limit.

#### Which data source does this actor read?

The Import.io extractor API (version 2.0) at api.import.io, through the read routes documented at docs.import.io: `GET /extractors/`, `GET /extractors/{extractorId}`, `GET /extractors/{extractorId}/inputs`, `GET /extractors/{extractorId}/crawlruns`, `GET /crawlruns/`, `GET /crawlruns/{crawlrunId}`, `GET /crawlruns/{crawlrunId}/json`, `GET /reports/`, `GET /reports/{reportId}/reportruns`, `GET /reportruns/`, `GET /reportruns/{reportRunId}`, `GET /reportruns/{reportRunId}/{csv|json}` and `GET /users/current/subscription`.

#### Do I need an API key?

Yes. This actor is bring-your-own-key and never ships one. Your API key is under account settings in your Import.io dashboard. Paste it into the input, or set it once as the `DATA_API_KEY` environment secret. A missing or refused key ends the run cleanly with a message saying which it was.

#### Does reading data use up my Import.io queries?

The provider's documentation says that endpoints returning data an extractor has already collected do not count as queries toward the plan total. This actor only reads what exists and never runs an extractor. The plan and usage service shows the count, so you can compare it before and after a run.

#### How do I find extractor ids and crawl run ids?

Run the extractors service first; each row carries `extractorId`. The latest result service needs only an extractor id and finds the newest finished crawl run itself. For an older crawl, run the crawl runs service with the extractor id and pick a `crawlRunId`.

#### Can this actor start an extractor or change anything?

No. It never starts or stops an extractor or a report, and no create, update, duplicate or delete route is wired.

#### What happens when there is nothing for an id?

It becomes its own row with `found: false` and a note: an extractor with no finished crawl run yet, a crawl run whose file holds no rows, an id the account cannot see, or a JSON file asked for on a data report. Those rows are never charged.

#### How is it priced?

Pay per result: one flat price per record returned, whether an extracted row, a report row, a crawl run, a report run, a report, an extractor, an input or the plan record. Rows for ids that produced nothing are free. Your Import.io subscription applies separately.

### Example output

```json
{
  "service": "latestResult",
  "serviceLabel": "Extracted rows of the latest finished crawl run",
  "requested": "0c5e2b7a-4f1d-4b8e-9a63-2d7f1e8c9b40",
  "found": true,
  "extractorId": "0c5e2b7a-4f1d-4b8e-9a63-2d7f1e8c9b40",
  "crawlRunId": "7b1d9e22-3c4a-4f5b-8e6d-1a2b3c4d5e6f",
  "state": "FINISHED",
  "startedAt": "2026-09-28T05:00:04.112Z",
  "stoppedAt": "2026-09-28T05:14:37.905Z",
  "rowIndex": 0,
  "data": {
    "Product name": "Trail Running Shoe 1042",
    "Price": "89.99",
    "Availability": "In stock",
    "Rating": "4.6",
    "Product link": "https://shop.example.com/products/trail-running-shoe-1042"
  },
  "retrievedAt": "2026-09-28T09:14:52.118Z",
  "note": null
}
```

Values are illustrative; the columns under `data` are whatever your extractor captures.

### Keyword map

web scraping API, extractor data export, web data extraction API, ecommerce product data export, competitor price scraping data, price change report export, crawl run history, scraping usage and quota, scraper output to database, Import.io API, import.io extractors, import.io crawl runs.

# Actor input Schema

## `service` (type: `string`):

Extractors lists every extractor in the account and needs nothing else, so it is the default; it is where extractor ids come from. Latest result returns every extracted row of an extractor's newest finished crawl run. Crawl runs lists run history, which is where crawl run ids come from. Reports and report runs list data and change reports, which is where report run ids come from.

## `identifiers` (type: `array`):

One per line. Extractor ids for latest result, extractor details and inputs, and optionally to narrow crawl runs. Crawl run ids for result and details of a crawl run. Report run ids for rows and details of a report run, and optionally report ids to narrow report runs. Extractors, reports and plan and usage read none.

## `crawlRunState` (type: `string`):

Crawl runs only: keep only runs in this state. Any keeps every run.

## `sortDirection` (type: `string`):

Lists only (extractors, crawl runs, reports, report runs): sort by creation time, newest first by default.

## `reportFileType` (type: `string`):

Rows of a report run only: the results file read. CSV works for every report; JSON exists for change reports only.

## `maxResults` (type: `integer`):

Stop after this many rows. One crawl run can hold many thousands of extracted rows.

## `requestsPerMinute` (type: `integer`):

Pacing ceiling for calls to the provider. Rate limited answers are retried after a pause.

## `apiKey` (type: `string`):

Your own API key, copied from account settings in your Import.io dashboard. This actor is bring-your-own-key and never ships one. Leave blank to use the DATA\_API\_KEY environment secret. The key is never written to a row or logged.

## `baseUrl` (type: `string`):

Override the host the actor calls. Only useful for testing against a different environment.

## Actor input object example

```json
{
  "service": "extractors",
  "crawlRunState": "any",
  "sortDirection": "DESC",
  "reportFileType": "csv",
  "maxResults": 1000,
  "requestsPerMinute": 60
}
```

# Actor output Schema

## `records` (type: `string`):

One row per extracted row, report row, crawl run, report run, report, extractor, input or plan.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("nabeelbaghoor/web-scraping-extractor-data-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("nabeelbaghoor/web-scraping-extractor-data-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call nabeelbaghoor/web-scraping-extractor-data-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nabeelbaghoor/web-scraping-extractor-data-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7uXXXlzKzX9aTtChI/builds/7QxtCK76PgRCkulGq/openapi.json
