# Actor Regression Watch — Dataset Schema, Count & Run Checks (`mouadapi/actor-regression-watch`) Actor

Checks your own Apify datasets and runs for regressions: item count drops, fields added or removed, type changes, empty fields, failed rows and run time; only counts and field names, never row contents; failed and unchanged rows are never charged. Not affiliated with or endorsed by Apify.

- **URL**: https://apify.com/mouadapi/actor-regression-watch.md
- **Developed by:** [COMPASS DEV](https://apify.com/mouadapi) (community)
- **Categories:** Developer tools, Integrations
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 dataset checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Checks your own Apify datasets and runs for regressions (item count drops, fields added or removed, type changes, empty
fields, failed rows, run time) and returns one row per dataset with counts and field names only; failed and unchanged rows
are never charged.

Pick your datasets, or connect it as an integration of your Actor or task, and every run's output is compared with the
last good ones: you see at once when a scraper starts returning fewer items, loses a field, gets prices as text or fills
half its rows with errors. It runs with Apify's limited permissions and needs no API token: it can read only the datasets
you give it. *Not affiliated with or endorsed by Apify.*

### What it does

- **Profiles each dataset:** its item count, every field with its most common type and how often it is empty, the share of
  failed rows (a `status` of failed or error, or a filled `error` field) and of no-data rows, and a short schema
  fingerprint.
- **Compares it with a baseline:** the median of the last checks (5 by default) of the same Actor, task or watch list, and
  says each regression in one plain sentence:
  - the item count fell by `maxItemDropPercent` or more (30% by default), or the dataset is empty;
  - a field is gone, or changed type (`price: number → string`);
  - a field is empty in many more rows (`maxEmptyIncreasePoints`, 20 points by default);
  - failed or no-data rows rose (`maxFailedIncreasePoints`, 10 points by default);
  - a required field (`requiredFields`) is empty in some rows;
  - as an integration: the run did not succeed, or took much longer (`maxRunTimeIncreasePercent`, 100% and at least 60 s
    more by default).
- **Export mode** (no watch list): the datasets of the run are compared with each other, oldest first.
- **Watch mode** (give a watch list name in `stateName`): each new dataset is compared with the list's last checks, and a
  dataset already checked is skipped. A watch run with nothing new returns **exactly one free `no_data` row** that says
  so.
- **Never row contents:** rows are read in memory only to count. The output holds counts, percentages, field names and type
  names, never a value from your data. A field name that could be a person's name, holds an @ or looks like data is counted,
  never listed.
- **Never a false alarm from a dataset still being written:** Apify updates a dataset's item count a few seconds after the
  writes. A dataset written in the last 2 minutes that looks empty, or whose count fell, is read again up to 3 times,
  10 seconds apart, before the finding is reported. If the count never settles (or the dataset shows writes but no row yet),
  the row is a free `no_data` row with `noDataReason: "pending_dataset_write"`: check it again in a minute.
- **You are never charged for failed results:** failed, no_data and example rows are free.

### Quick start

Compare two datasets (the second against the first):

```json
{ "datasetIds": ["WkzbQMuFYuamGv3YF", "Hk3QxqZmV8B7sYp2a"] }
```

Check every run of your scraper automatically: open your Actor or task, go to **Integrations**, add **Actor Regression
Watch**, and give it this input (Apify fills in the dataset of each run that finishes):

```json
{ "datasetIds": ["{{resource.defaultDatasetId}}"], "stateName": "my-scraper" }
```

The first check is the baseline; each later run is compared with the last checks of the same Actor or task, with its run
status and run time. Leave the input empty to see a built-in example (free).

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `datasetIds` | list of dataset IDs | — (empty: the built-in example) | Your datasets, oldest first; the picker gives the run read access to them only. As an integration: `["{{resource.defaultDatasetId}}"]` |
| `mode` | `watch` or `export` | — | Empty: watch when `stateName` is given, otherwise export. An explicit mode wins |
| `stateName` | string | — | Watch list name. Watch mode without a name uses the list `default` |
| `requiredFields` | list of strings | `[]` | Fields every row must have, filled |
| `maxItemDropPercent` | integer 1–100 | `30` | An item count this many percent under the baseline is a regression |
| `maxEmptyIncreasePoints` | integer 1–100 | `20` | A field empty in this many more percentage points of rows is a regression |
| `maxFailedIncreasePoints` | integer 1–100 | `10` | Failed or no-data rows rising by this many points is a regression |
| `maxRunTimeIncreasePercent` | integer 1–10,000 | `100` | Integration runs: a run time this many percent over the baseline (and 60 s more) is a regression |
| `baselineChecks` | integer 1–20 | `5` | How many earlier checks form the baseline |
| `sampleSize` | integer 100–10,000 | `1000` | Rows read per dataset for the field checks (the first ones) |
| `maxItems` | integer 1–1,000 | `100` | Most charged checks per run; in watch mode the rest come in the next run |

Only `datasetIds` gives the run read access to datasets, so no other field name is read for them.

### Output

```json
{
    "status": "ok",
    "attempts": 1,
    "error": null,
    "input": "WkzbQMuFYuamGv3YF",
    "mode": "watch",
    "watchList": "my-scraper",
    "example": false,
    "datasetId": "WkzbQMuFYuamGv3YF",
    "datasetName": null,
    "datasetCreatedAt": "2026-10-04T06:30:12.000Z",
    "baselineKey": "task:Hk3QxqZmV8B7sYp2a",
    "actorId": "aYG0l9s7dL7xXKFh1",
    "actorTaskId": "Hk3QxqZmV8B7sYp2a",
    "runId": "Fj3kd8sLq0ZpX7mNc",
    "runStatus": "SUCCEEDED",
    "runStartedAt": "2026-10-04T06:30:10.000Z",
    "runFinishedAt": "2026-10-04T06:36:40.000Z",
    "runTimeSecs": 390,
    "runTimeBaselineSecs": 120,
    "runTimeChangePercent": 225,
    "verdict": "regression",
    "regressionCount": 4,
    "regressions": [
        "Item count fell 40% (100 → 60; the baseline is the median of 5 checks)",
        "Field rating is gone (it was in the baseline)",
        "Field price changed type: number → string",
        "Run time rose 225% (120 s → 390 s)"
    ],
    "itemCount": 60,
    "itemCountBaseline": 100,
    "itemCountChangePercent": -40,
    "sampledRows": 60,
    "fieldCount": 6,
    "fieldsHidden": 0,
    "schemaFingerprint": "3f9a1c0b7d2e",
    "fieldsAdded": [],
    "fieldsRemoved": ["rating"],
    "typeChanges": ["price: number → string"],
    "emptyRateChanges": [],
    "requiredFieldsMissing": [],
    "failedRowsPercent": 0,
    "failedRowsBaselinePercent": 0,
    "noDataRowsPercent": 0,
    "fieldProfile": ["currency: string", "inStock: boolean", "price: string", "status: string", "title: string", "url: string"],
    "baselineChecks": 5,
    "url": "https://console.apify.com/storage/datasets/WkzbQMuFYuamGv3YF",
    "source": "Your own Apify dataset, read with this run's limited-permission token (only the access you gave the run). Only counts and field names are returned.",
    "scrapedAt": "2026-10-04T06:37:02.000Z"
}
```

| Status | Meaning | Charged? |
|---|---|---|
| `ok` | A dataset checked: `verdict` is `regression`, `pass`, or `baseline` (the first check of its Actor, task or list) | Yes |
| `no_data` | No such dataset, nothing new since the last run of the watch list, the built-in example, or a dataset still being written (`noDataReason` says which) | No |
| `failed` | Invalid entry, a dataset this run may not read, or no answer after retries (`error` says why) | No |

The key-value store holds `RUN_REPORT` (counts, charged and free rows, requests, stop reason) and, when something fails,
the API's answer (`SNAPSHOT_*`; never a page of your rows, only its size). The watch list is a named key-value store in your
account (`actor-regression-watch-state-<name>`): the last profiles of each key, with counts and field names only.

### Use it from AI agents

One clear main input, `datasetIds`; every row has `status`, `verdict`, `regressions`, `error`, `url` and `scrapedAt`. Call
it through the Apify API, the Apify MCP server (`mouadapi/actor-regression-watch`) or x402 agentic payments. The datasets
must be yours (the run may read only those). Copy-paste call (your Apify token in `APIFY_TOKEN`):

```bash
curl -s -X POST "https://api.apify.com/v2/acts/mouadapi~actor-regression-watch/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"datasetIds": ["WkzbQMuFYuamGv3YF", "Hk3QxqZmV8B7sYp2a"]}'
```

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('mouadapi/actor-regression-watch').call({ datasetIds: ['WkzbQMuFYuamGv3YF', 'Hk3QxqZmV8B7sYp2a'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => [r.datasetId, r.verdict, r.regressions]));
```

### Output fields

| Field | Type | Description |
|---|---|---|
| `status` | string (or null) | ok = a dataset checked (charged); no_data = nothing to check: no such dataset, nothing new since the last run of the watch list, the built-in example, or a dataset still being written (noDataReason says which) (free); failed = invalid input, no access to the dataset, or no answer (free) |
| `attempts` | integer (or null) | Requests made for this entry (retries included) |
| `error` | string (or null) | Why a row is no_data or failed; null on ok rows |
| `input` | string (or null) | The entry that produced this row: a dataset ID, the watch list for the nothing-new row, or the example dataset |
| `mode` | string (or null) | watch (compared with the watch list's last checks) or export (compared with this run's earlier datasets) |
| `watchList` | string (or null) | The watch list name in watch mode; null in export mode |
| `example` | boolean (or null) | true only on the built-in example's free rows (nothing of yours was read) |
| `noDataReason` | string (or null) | Why a row is no_data: not_found (no such dataset), nothing_new (the watch list checked every dataset before), example (the built-in example), pending_dataset_write (written in the last 2 minutes and its item count had not settled: free, check it again in a minute); null on ok and failed rows |
| `datasetId` | string (or null) | The checked dataset's ID |
| `datasetName` | string (or null) | The dataset's name, or null for an unnamed dataset |
| `datasetCreatedAt` | string (or null) | When the dataset was created (ISO 8601) |
| `baselineKey` | string (or null) | What the dataset is compared with: task:<id> or actor:<id> for an integration's runs, list:<name> for a watch list, list for this run's own datasets |
| `actorId` | string (or null) | The Actor whose run made the dataset (integration runs only) |
| `actorTaskId` | string (or null) | The task whose run made the dataset (integration runs only) |
| `runId` | string (or null) | The run that made the dataset (integration runs only) |
| `runStatus` | string (or null) | How that run ended (integration runs only); anything but SUCCEEDED is a regression |
| `runStartedAt` | string (or null) | When that run started (ISO 8601; integration runs only) |
| `runFinishedAt` | string (or null) | When that run finished (ISO 8601; integration runs only) |
| `runTimeSecs` | number (or null) | How long that run took, in seconds (integration runs only) |
| `runTimeBaselineSecs` | number (or null) | The median run time of the baseline's checks, in seconds |
| `runTimeChangePercent` | number (or null) | The run time's change against the baseline, in percent |
| `verdict` | string (or null) | regression = at least one finding below; pass = none; baseline = the first check of its key (no earlier check to compare with) |
| `regressionCount` | integer (or null) | How many findings the check has |
| `regressions` | array (or null) | Each regression in one plain sentence (counts and field names only, never a value from a row) |
| `itemCount` | integer (or null) | The dataset's item count (its own total, not the sample) |
| `itemCountBaseline` | number (or null) | The median item count of the baseline's checks |
| `itemCountChangePercent` | number (or null) | The item count's change against the baseline, in percent |
| `sampledRows` | integer (or null) | How many rows were read for the field checks (the first ones, up to sampleSize) |
| `fieldCount` | integer (or null) | Fields in the sample (a key in at least 2 rows and 1% of them) |
| `fieldsHidden` | integer (or null) | Fields counted but never listed: a name that could be a person's, holds an @, or is over 80 characters |
| `schemaFingerprint` | string (or null) | A short hash of the listed fields' names and types: the same value means the same schema |
| `fieldsAdded` | array (or null) | Fields that are new against the baseline (not a regression on their own) |
| `fieldsRemoved` | array (or null) | Fields of the baseline that are gone |
| `typeChanges` | array (or null) | Fields whose most common type changed: "field: before → now" |
| `emptyRateChanges` | array (or null) | Fields empty in more rows than in the baseline (by at least maxEmptyIncreasePoints) |
| `requiredFieldsMissing` | array (or null) | Required fields (requiredFields) that are empty in some sampled rows, with the count |
| `failedRowsPercent` | number (or null) | Sampled rows with a status of failed or error, or a filled error field, in percent |
| `failedRowsBaselinePercent` | number (or null) | The baseline's median share of failed rows, in percent |
| `noDataRowsPercent` | number (or null) | Sampled rows with a status of no_data, not_found or empty, in percent |
| `fieldProfile` | array (or null) | Each listed field with its most common type and, when above 0, its empty share (at most 100 fields, by name) |
| `baselineChecks` | integer (or null) | How many earlier checks the comparison used |
| `url` | string (or null) | The dataset's page in Apify Console (the Store page for the built-in example) |
| `source` | string (or null) | Where the data comes from |
| `scrapedAt` | string (or null) | When the row was made (ISO 8601) |

### Pricing

Pay per event: one `dataset-check` event per checked dataset (an `ok` row, whatever its verdict). `no_data`, `failed` and
example rows are free.

| Plan | Price per 1,000 checks |
|---|---|
| Free | $1.50 |
| Bronze | $1.30 |
| Silver | $1.15 |
| Gold (and higher) | $1.00 |

Apify also charges its small per-run start event. There are no usage fees on top.

### Limits

- One request at a time to the Apify API, at least 250 ms apart. If the API answers HTTP 429, the run pauses 60, 120 and
  240 seconds on the same connection, then stops with free `failed` rows. No other token, IP or proxy is ever tried.
- Each dataset: its item count, then its first `sampleSize` rows (1,000 by default) in pages of up to 200 rows, fewer when
  rows are large (a page stays near 8 MB), at most 60 pages. Apify updates a dataset's item count a few seconds after the
  writes, so the count used is the larger of that count and the rows read.
- At most 100 datasets per run.

### Known issues

- The field checks read the first rows of a dataset: a change that shows only in later rows is not seen unless
  `sampleSize` covers it. The item count is always the dataset's own total.
- Fields are top-level keys: a nested object is one field of type `object`.
- Run status and run time come only with an integration (a dataset alone does not say how its run went).
- Datasets are given by ID: a dataset's name (or username~name) is refused, free, because the picker grants read access by
  ID only. As an integration, the input must name the run's dataset (`["{{resource.defaultDatasetId}}"]`): a dataset that
  is only in the run's details cannot be read, and that row says so (free).

### FAQ

**Does it need my API token?** No. It runs with Apify's limited permissions and reads only the datasets you pick: the
picker grants it read access to those, including the run's dataset an integration names in `datasetIds`.

**Can it see my other data?** No. Limited permissions let it read only its own storages, its watch list and the datasets you
give it.

**Are my rows returned or kept?** No. Rows are read in memory to count fields, types and empty or failed rows. The output
and the watch list hold counts, percentages, field names and type names only.

**Why did a watch run return one `no_data` row?** Every dataset it was given was checked before by that watch list; the
row says how many. It is free.

**What is a `pending_dataset_write` row?** The dataset was written in the last 2 minutes and its item count did not settle
within 30 seconds (Apify updates the count a little after the writes), so it was not checked. It is free, and it is not
recorded: run the check again with that dataset in a minute, and it is checked then.

**Why is a field missing from `fieldProfile`?** A key counts as a field when it is in at least 2 sampled rows and 1% of
them; one that could be a person's name, holds an @ or is over 80 characters is counted in `fieldsHidden`, never listed.

### Data and licence

- Source: your own datasets, through the Apify API v2 (documented at https://docs.apify.com/api/v2), read with the run's own
  limited-permission token.
- Your data stays yours (Apify General Terms 5.8). This Actor returns only what it counted.
- This Actor is not affiliated with, endorsed by or provided by Apify.

# Actor input Schema

## `datasetIds` (type: `array`):

Your datasets, oldest first. Without a watch list, each is compared with the ones before it; with one, with that list's last checks. As an integration of your Actor or task, use \["{{resource.defaultDatasetId}}"]: the dataset of the run that just finished. The picker gives this run read access to these datasets only. Leave empty to see a built-in example (free).

## `mode` (type: `string`):

Empty: watch when a watch list name is given, otherwise export. "watch" compares each dataset with the watch list's last checks and skips datasets it already checked; "export" compares the datasets of this run with each other. An explicit mode wins.

## `stateName` (type: `string`):

Name of your watch list (kept in your own storage between runs). Each Actor or task an integration names keeps its own baseline in it; without an integration, the list itself is the baseline. Giving a name turns on watch mode; watch mode without a name uses the list "default".

## `requiredFields` (type: `array`):

Fields every row must have, filled. A sampled row without one is a regression, with the count. One field name per line.

## `maxItemDropPercent` (type: `integer`):

An item count this many percent under the baseline (the median of the earlier checks) is a regression. An empty dataset always is.

## `maxEmptyIncreasePoints` (type: `integer`):

A field empty in this many more percentage points of the sampled rows than in the baseline is a regression.

## `maxFailedIncreasePoints` (type: `integer`):

Failed rows (a status of failed or error, or a filled error field) or no-data rows rising by this many percentage points is a regression.

## `maxRunTimeIncreasePercent` (type: `integer`):

Integration runs only: a run time this many percent over the baseline, and at least 60 seconds more, is a regression.

## `baselineChecks` (type: `integer`):

How many earlier checks of the same Actor, task or list form the baseline.

## `sampleSize` (type: `integer`):

Rows read from each dataset for the field checks (the first ones). The item count is always the dataset's own total.

## `maxItems` (type: `integer`):

Most datasets checked and charged per run. In watch mode the rest come in the next run.

## Actor input object example

```json
{
  "datasetIds": [
    "WkzbQMuFYuamGv3YF"
  ],
  "requiredFields": [],
  "maxItemDropPercent": 30,
  "maxEmptyIncreasePoints": 20,
  "maxFailedIncreasePoints": 10,
  "maxRunTimeIncreasePercent": 100,
  "baselineChecks": 5,
  "sampleSize": 1000,
  "maxItems": 100
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset with one row per checked dataset (verdict, findings, counts and field names), or per entry when nothing is checked

## `runReport` (type: `string`):

Summary of the run (counts, charged and free rows, requests, stop reason)

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("mouadapi/actor-regression-watch").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("mouadapi/actor-regression-watch").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call mouadapi/actor-regression-watch --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mouadapi/actor-regression-watch"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HDr3pZpX5Hq8cj5JE/builds/95GsiGWzRG6kSxN2b/openapi.json
