Web Scraping Extractor Data API - Import.io Crawl Runs, Reports avatar

Web Scraping Extractor Data API - Import.io Crawl Runs, Reports

Pricing

$5.00 / 1,000 extractor data record returneds

Go to Apify Store
Web Scraping Extractor Data API - Import.io Crawl Runs, Reports

Web Scraping Extractor Data API - Import.io Crawl Runs, Reports

Export the output of your own Import.io web data extractors: every extracted row of the latest finished or any crawl run, crawl run history and stats, data and change report rows, extractors, inputs, plan and usage. Read only: nothing is started or changed. Bring your own key.

Pricing

$5.00 / 1,000 extractor data record returneds

Rating

0.0

(0)

Developer

Nabeel Hassan

Nabeel Hassan

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Export what your own Import.io web data extractors have already collected: every extracted row of an extractor's latest finished crawl run or of any past crawl run, crawl run history with URL and row counts, data and change report rows, and your extractors and their inputs, as clean rows ready for a warehouse, a spreadsheet or the next step of a pipeline.

What it collects

  • Extracted rows: the newest finished crawl run of an extractor, or any crawl run by id, one row per extracted row with every column under data, keyed by the extractor's own field names (product name, price, availability, rating, URL, whatever you trained it to capture).
  • Crawl run history: recent crawl runs across the account or per extractor, with state (finished, started, pending, cancelled, failed), start and stop time, total, successful and failed URLs, rows, queries used, screen captures and proxy traffic. Filter by state.
  • Report rows: the CSV rows of any data report run, or the JSON rows of a change report run, which is where price and availability changes between crawl runs show up.
  • Reports and report runs: every data and change report, and each run's status, summary and inputs.
  • Extractors and inputs: every extractor with its fields, tags, chaining and report settings, and each extractor's current URLs or inputs. Webhook headers, credentials and any field that looks like a secret are left out.
  • Plan and usage: queries used and allowed in the current period, plan, renewal and expiry.
  • Read only, pay per result, bring your own key.

Input

FieldWhat it does
What to readExtractors (default, needs nothing else), latest result, crawl runs, result of a crawl run, crawl run details, extractor details, extractor inputs, reports, report runs, rows of a report run, report run details, or plan and usage.
IdsOne per line: extractor ids, crawl run ids, report ids or report run ids, depending on the service.
Crawl run stateCrawl runs: keep only finished, running, pending, cancelled or failed runs.
OrderLists: newest first (default) or oldest first.
Report fileRows of a report run: CSV (any report) or JSON (change reports only).
Maximum resultsRow cap for the run.
Requests per minutePacing for calls to the provider.
API keyYour own API key, as a secret input.

A typical first run is the default extractors service. Copy an extractor id from it, then run latest result with that id to get the data.

FAQ

What is this web scraping extractor data API used for?

Getting data out of extractors that already run on a schedule and into the systems that use it. A pricing team pulls the latest competitor product and price crawl every morning into BigQuery or Snowflake. An ecommerce analyst exports a change report run to see which products changed price or went out of stock. An operations team reviews failed crawl runs and failed URL counts per extractor. A finance team checks queries used against the plan limit.

Which data source does this actor read?

The Import.io extractor API (version 2.0) at api.import.io, through the read routes documented at docs.import.io: GET /extractors/, GET /extractors/{extractorId}, GET /extractors/{extractorId}/inputs, GET /extractors/{extractorId}/crawlruns, GET /crawlruns/, GET /crawlruns/{crawlrunId}, GET /crawlruns/{crawlrunId}/json, GET /reports/, GET /reports/{reportId}/reportruns, GET /reportruns/, GET /reportruns/{reportRunId}, GET /reportruns/{reportRunId}/{csv|json} and GET /users/current/subscription.

Do I need an API key?

Yes. This actor is bring-your-own-key and never ships one. Your API key is under account settings in your Import.io dashboard. Paste it into the input, or set it once as the DATA_API_KEY environment secret. A missing or refused key ends the run cleanly with a message saying which it was.

Does reading data use up my Import.io queries?

The provider's documentation says that endpoints returning data an extractor has already collected do not count as queries toward the plan total. This actor only reads what exists and never runs an extractor. The plan and usage service shows the count, so you can compare it before and after a run.

How do I find extractor ids and crawl run ids?

Run the extractors service first; each row carries extractorId. The latest result service needs only an extractor id and finds the newest finished crawl run itself. For an older crawl, run the crawl runs service with the extractor id and pick a crawlRunId.

Can this actor start an extractor or change anything?

No. It never starts or stops an extractor or a report, and no create, update, duplicate or delete route is wired.

What happens when there is nothing for an id?

It becomes its own row with found: false and a note: an extractor with no finished crawl run yet, a crawl run whose file holds no rows, an id the account cannot see, or a JSON file asked for on a data report. Those rows are never charged.

How is it priced?

Pay per result: one flat price per record returned, whether an extracted row, a report row, a crawl run, a report run, a report, an extractor, an input or the plan record. Rows for ids that produced nothing are free. Your Import.io subscription applies separately.

Example output

{
"service": "latestResult",
"serviceLabel": "Extracted rows of the latest finished crawl run",
"requested": "0c5e2b7a-4f1d-4b8e-9a63-2d7f1e8c9b40",
"found": true,
"extractorId": "0c5e2b7a-4f1d-4b8e-9a63-2d7f1e8c9b40",
"crawlRunId": "7b1d9e22-3c4a-4f5b-8e6d-1a2b3c4d5e6f",
"state": "FINISHED",
"startedAt": "2026-09-28T05:00:04.112Z",
"stoppedAt": "2026-09-28T05:14:37.905Z",
"rowIndex": 0,
"data": {
"Product name": "Trail Running Shoe 1042",
"Price": "89.99",
"Availability": "In stock",
"Rating": "4.6",
"Product link": "https://shop.example.com/products/trail-running-shoe-1042"
},
"retrievedAt": "2026-09-28T09:14:52.118Z",
"note": null
}

Values are illustrative; the columns under data are whatever your extractor captures.

Keyword map

web scraping API, extractor data export, web data extraction API, ecommerce product data export, competitor price scraping data, price change report export, crawl run history, scraping usage and quota, scraper output to database, Import.io API, import.io extractors, import.io crawl runs.