# Web Scraping Results API - Dexi Robots, Runs and Datasets (`nabeelbaghoor/web-scraping-robot-results-api`) Actor

Export the output of your own Dexi web scraping robots: every scraped row of the latest or any execution, execution history and usage statistics, dataset rows with filters, plus your robots and runs. Read only: nothing is started or changed. Bring your own key.

- **URL**: https://apify.com/nabeelbaghoor/web-scraping-robot-results-api.md
- **Developed by:** [Nabeel Hassan](https://apify.com/nabeelbaghoor) (community)
- **Categories:** Developer tools, E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 scraping result record returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Web Scraping Results API - Dexi Robots, Runs and Datasets

Export what your own Dexi web scraping robots have already collected: every scraped row of a run's latest result or of any past execution, the execution history of a run, usage statistics, and dataset rows, as clean rows ready for a warehouse, a spreadsheet or the next step of a pipeline.

### What it collects

- **Scraped rows**: the latest result of a run (from its latest successful execution by default) or the result of any execution by id, one row per scraped row with every field under `data`, keyed by the robot's own column names.
- **Execution history**: every execution of a run, newest first, with its state (OK, failed, stopped, running, pending, queued), start and finish time, optionally within a date range.
- **Execution statistics**: page visits, requests, time and traffic used, current, failed and total results, concurrency and who started it.
- **Dataset rows**: the rows collected into a dataset, with the provider's own query filters (equals, in, between, greater than and more) and nested fields by dot path.
- **Robots and runs**: every robot in the account and every run (saved configuration), with ids and names, which is where the ids for everything else come from. Nested robot definitions and any field that looks like a credential are left out.
- Read only, pay per result, bring your own key.

### Input

| Field | What it does |
| --- | --- |
| What to read | Robots (default, needs nothing else), runs, executions of a run, latest result of a run, result of an execution, execution statistics or dataset rows. |
| Ids | One per line: run ids, execution ids or dataset ids, depending on the service. Robot ids optionally narrow runs. |
| Latest result from executions in state | Latest result: OK by default, or any state. |
| From date / To date | Executions: the date range. |
| Dataset query | Dataset rows: the provider's JSON query. |
| Maximum results | Row cap for the run. |
| Requests per minute | Pacing for calls to the provider. |
| Account id | Your account id, as a secret input. |
| API key | Your own API key, as a secret input. |

### FAQ

#### What is this web scraping results API used for?

Getting data out of scraping robots that already run on a schedule and into the systems that use it. A pricing team pulls the latest competitor price scrape every morning into BigQuery or Snowflake. An analyst exports last month's executions of one run to see which ones failed. An operations team checks page visits and traffic per execution to control scraping costs. A product team reads a dataset of collected listings filtered by category and price.

#### Which data source does this actor read?

The Dexi API at api.dexi.io, through the read routes documented on its public developer portal: `GET /robots`, `GET /runs`, `GET /runs/{runId}/executions`, `GET /runs/{runId}/latest/result`, `GET /executions/{executionId}/result`, `GET /executions/{executionId}/stats` and `POST /datasets/{dataSetId}/rows`, which is a query sent as a POST and writes nothing.

#### Do I need an API key?

Yes. This actor is bring-your-own-key and never ships one. Your account id and API key are both shown on the API page of your Dexi account. Paste both into the input, or paste `accountId:apiKey` into the API key field, or set them once as the `DATA_API_ACCOUNT_ID` and `DATA_API_KEY` environment secrets. The actor computes the documented access key (MD5 of account id and API key) itself, so the key is never sent. Missing or refused credentials end the run cleanly with a message saying which it was.

#### How do I find run ids and execution ids?

Run the robots service first, then the runs service (optionally with robot ids) to get run ids. The executions service lists a run's executions with their ids and states. The latest result service needs only a run id, which is usually all you need.

#### Can this actor start a robot or change anything?

No. It never starts, stops or retries an execution, and no create, update or delete route is wired. It reads what your robots have already collected.

#### What happens when there is nothing for an id?

It becomes its own row with `found: false` and a note: a run with no successful execution yet, an execution id the account cannot see, or a dataset query that matches nothing. Those rows are never charged.

#### How is it priced?

Pay per result: one flat price per record returned, whether a scraped row, a dataset row, an execution, a statistics row, a run or a robot. Rows for ids that produced nothing are free. Your Dexi subscription applies separately.

### Example output

```json
{
  "service": "latestResult",
  "serviceLabel": "Latest result of a run",
  "requested": "b64a8397-2159-4466-a907-cc615d6530de",
  "found": true,
  "runId": "b64a8397-2159-4466-a907-cc615d6530de",
  "executionId": null,
  "rowIndex": 0,
  "data": {
    "product_name": "Trail Running Shoe 1042",
    "price": "89.99",
    "currency": "EUR",
    "in_stock": "true",
    "url": "https://shop.example.com/products/trail-running-shoe-1042"
  },
  "retrievedAt": "2026-09-28T09:14:52.118Z",
  "note": null
}
```

Values are illustrative; the columns under `data` are whatever your robot extracts.

### Keyword map

web scraping API, scraping results export, web data extraction API, scraper output to database, scheduled scraping export, competitor price scraping data, scraping execution history, scraping usage statistics, dataset export API, Dexi API, dexi.io robots, web extraction robots.

# Actor input Schema

## `service` (type: `string`):

Robots lists every robot in the account and needs nothing else, so it is the default. Runs lists each robot's saved configurations, which is where run ids come from. Executions lists the times a run ran, which is where execution ids come from. Latest result and result of an execution return every scraped row. Execution statistics give usage and result counts. Dataset rows reads rows collected into a dataset.

## `identifiers` (type: `array`):

One per line. Run ids for executions and latest result, execution ids for result of an execution and execution statistics, dataset ids for dataset rows, and optionally robot ids to narrow runs. Robots reads none.

## `resultState` (type: `string`):

Latest result only: the state of the execution whose result is read. OK, the default, reads the latest successful execution; any sends no state and reads the latest execution whatever its state.

## `fromDate` (type: `string`):

Executions only: list executions from this date, as YYYY-MM-DD or YYYY-MM-DD HH:mm:ss.

## `toDate` (type: `string`):

Executions only: list executions up to this date, as YYYY-MM-DD (the whole day) or YYYY-MM-DD HH:mm:ss.

## `datasetQuery` (type: `object`):

Dataset rows only: the provider's query, a JSON object of field name to query options, such as {"price": {"type": "GT", "value": 10}} or {"category": {"type": "IN", "values": \["Shoes", "Bags"]}}. Nested fields use dots. Left empty, every row is read.

## `maxResults` (type: `integer`):

Stop after this many rows. One execution result can hold thousands of scraped rows.

## `requestsPerMinute` (type: `integer`):

Pacing ceiling for calls to the provider. Rate limited answers are retried after a pause.

## `accountId` (type: `string`):

Your account id, shown on the API page of your Dexi account. Leave blank to use the DATA\_API\_ACCOUNT\_ID environment secret, or paste accountId:apiKey into the API key field instead.

## `apiKey` (type: `string`):

Your own API key, generated on the API page of your Dexi account (not the access key, which this actor computes from the two). This actor is bring-your-own-key and never ships one. Leave blank to use the DATA\_API\_KEY environment secret. The key itself is never sent, written to a row or logged.

## `baseUrl` (type: `string`):

Override the host the actor calls. Only useful for testing against a different environment.

## Actor input object example

```json
{
  "service": "robots",
  "resultState": "OK",
  "maxResults": 1000,
  "requestsPerMinute": 60
}
```

# Actor output Schema

## `records` (type: `string`):

One row per scraped result row, dataset row, execution, run or robot.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("nabeelbaghoor/web-scraping-robot-results-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("nabeelbaghoor/web-scraping-robot-results-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call nabeelbaghoor/web-scraping-robot-results-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nabeelbaghoor/web-scraping-robot-results-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kHKAiSfssx7MEpp30/builds/aqIaglDZMl0AMyN89/openapi.json
