Actor Regression Watch — Dataset Schema, Count & Run Checks
Pricing
from $1.00 / 1,000 dataset checkeds
Actor Regression Watch — Dataset Schema, Count & Run Checks
Checks your own Apify datasets and runs for regressions: item count drops, fields added or removed, type changes, empty fields, failed rows and run time; only counts and field names, never row contents; failed and unchanged rows are never charged. Not affiliated with or endorsed by Apify.
Pricing
from $1.00 / 1,000 dataset checkeds
Rating
0.0
(0)
Developer
COMPASS DEV
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Checks your own Apify datasets and runs for regressions (item count drops, fields added or removed, type changes, empty fields, failed rows, run time) and returns one row per dataset with counts and field names only; failed and unchanged rows are never charged.
Pick your datasets, or connect it as an integration of your Actor or task, and every run's output is compared with the last good ones: you see at once when a scraper starts returning fewer items, loses a field, gets prices as text or fills half its rows with errors. It runs with Apify's limited permissions and needs no API token: it can read only the datasets you give it. Not affiliated with or endorsed by Apify.
What it does
- Profiles each dataset: its item count, every field with its most common type and how often it is empty, the share of
failed rows (a
statusof failed or error, or a fillederrorfield) and of no-data rows, and a short schema fingerprint. - Compares it with a baseline: the median of the last checks (5 by default) of the same Actor, task or watch list, and
says each regression in one plain sentence:
- the item count fell by
maxItemDropPercentor more (30% by default), or the dataset is empty; - a field is gone, or changed type (
price: number → string); - a field is empty in many more rows (
maxEmptyIncreasePoints, 20 points by default); - failed or no-data rows rose (
maxFailedIncreasePoints, 10 points by default); - a required field (
requiredFields) is empty in some rows; - as an integration: the run did not succeed, or took much longer (
maxRunTimeIncreasePercent, 100% and at least 60 s more by default).
- the item count fell by
- Export mode (no watch list): the datasets of the run are compared with each other, oldest first.
- Watch mode (give a watch list name in
stateName): each new dataset is compared with the list's last checks, and a dataset already checked is skipped. A watch run with nothing new returns exactly one freeno_datarow that says so. - Never row contents: rows are read in memory only to count. The output holds counts, percentages, field names and type names, never a value from your data. A field name that could be a person's name, holds an @ or looks like data is counted, never listed.
- Never a false alarm from a dataset still being written: Apify updates a dataset's item count a few seconds after the
writes. A dataset written in the last 2 minutes that looks empty, or whose count fell, is read again up to 3 times,
10 seconds apart, before the finding is reported. If the count never settles (or the dataset shows writes but no row yet),
the row is a free
no_datarow withnoDataReason: "pending_dataset_write": check it again in a minute. - You are never charged for failed results: failed, no_data and example rows are free.
Quick start
Compare two datasets (the second against the first):
{ "datasetIds": ["WkzbQMuFYuamGv3YF", "Hk3QxqZmV8B7sYp2a"] }
Check every run of your scraper automatically: open your Actor or task, go to Integrations, add Actor Regression Watch, and give it this input (Apify fills in the dataset of each run that finishes):
{ "datasetIds": ["{{resource.defaultDatasetId}}"], "stateName": "my-scraper" }
The first check is the baseline; each later run is compared with the last checks of the same Actor or task, with its run status and run time. Leave the input empty to see a built-in example (free).
Input
| Field | Type | Default | Description |
|---|---|---|---|
datasetIds | list of dataset IDs | — (empty: the built-in example) | Your datasets, oldest first; the picker gives the run read access to them only. As an integration: ["{{resource.defaultDatasetId}}"] |
mode | watch or export | — | Empty: watch when stateName is given, otherwise export. An explicit mode wins |
stateName | string | — | Watch list name. Watch mode without a name uses the list default |
requiredFields | list of strings | [] | Fields every row must have, filled |
maxItemDropPercent | integer 1–100 | 30 | An item count this many percent under the baseline is a regression |
maxEmptyIncreasePoints | integer 1–100 | 20 | A field empty in this many more percentage points of rows is a regression |
maxFailedIncreasePoints | integer 1–100 | 10 | Failed or no-data rows rising by this many points is a regression |
maxRunTimeIncreasePercent | integer 1–10,000 | 100 | Integration runs: a run time this many percent over the baseline (and 60 s more) is a regression |
baselineChecks | integer 1–20 | 5 | How many earlier checks form the baseline |
sampleSize | integer 100–10,000 | 1000 | Rows read per dataset for the field checks (the first ones) |
maxItems | integer 1–1,000 | 100 | Most charged checks per run; in watch mode the rest come in the next run |
Only datasetIds gives the run read access to datasets, so no other field name is read for them.
Output
{"status": "ok","attempts": 1,"error": null,"input": "WkzbQMuFYuamGv3YF","mode": "watch","watchList": "my-scraper","example": false,"datasetId": "WkzbQMuFYuamGv3YF","datasetName": null,"datasetCreatedAt": "2026-10-04T06:30:12.000Z","baselineKey": "task:Hk3QxqZmV8B7sYp2a","actorId": "aYG0l9s7dL7xXKFh1","actorTaskId": "Hk3QxqZmV8B7sYp2a","runId": "Fj3kd8sLq0ZpX7mNc","runStatus": "SUCCEEDED","runStartedAt": "2026-10-04T06:30:10.000Z","runFinishedAt": "2026-10-04T06:36:40.000Z","runTimeSecs": 390,"runTimeBaselineSecs": 120,"runTimeChangePercent": 225,"verdict": "regression","regressionCount": 4,"regressions": ["Item count fell 40% (100 → 60; the baseline is the median of 5 checks)","Field rating is gone (it was in the baseline)","Field price changed type: number → string","Run time rose 225% (120 s → 390 s)"],"itemCount": 60,"itemCountBaseline": 100,"itemCountChangePercent": -40,"sampledRows": 60,"fieldCount": 6,"fieldsHidden": 0,"schemaFingerprint": "3f9a1c0b7d2e","fieldsAdded": [],"fieldsRemoved": ["rating"],"typeChanges": ["price: number → string"],"emptyRateChanges": [],"requiredFieldsMissing": [],"failedRowsPercent": 0,"failedRowsBaselinePercent": 0,"noDataRowsPercent": 0,"fieldProfile": ["currency: string", "inStock: boolean", "price: string", "status: string", "title: string", "url: string"],"baselineChecks": 5,"url": "https://console.apify.com/storage/datasets/WkzbQMuFYuamGv3YF","source": "Your own Apify dataset, read with this run's limited-permission token (only the access you gave the run). Only counts and field names are returned.","scrapedAt": "2026-10-04T06:37:02.000Z"}
| Status | Meaning | Charged? |
|---|---|---|
ok | A dataset checked: verdict is regression, pass, or baseline (the first check of its Actor, task or list) | Yes |
no_data | No such dataset, nothing new since the last run of the watch list, the built-in example, or a dataset still being written (noDataReason says which) | No |
failed | Invalid entry, a dataset this run may not read, or no answer after retries (error says why) | No |
The key-value store holds RUN_REPORT (counts, charged and free rows, requests, stop reason) and, when something fails,
the API's answer (SNAPSHOT_*; never a page of your rows, only its size). The watch list is a named key-value store in your
account (actor-regression-watch-state-<name>): the last profiles of each key, with counts and field names only.
Use it from AI agents
One clear main input, datasetIds; every row has status, verdict, regressions, error, url and scrapedAt. Call
it through the Apify API, the Apify MCP server (mouadapi/actor-regression-watch) or x402 agentic payments. The datasets
must be yours (the run may read only those). Copy-paste call (your Apify token in APIFY_TOKEN):
curl -s -X POST "https://api.apify.com/v2/acts/mouadapi~actor-regression-watch/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"datasetIds": ["WkzbQMuFYuamGv3YF", "Hk3QxqZmV8B7sYp2a"]}'
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('mouadapi/actor-regression-watch').call({ datasetIds: ['WkzbQMuFYuamGv3YF', 'Hk3QxqZmV8B7sYp2a'] });const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items.map((r) => [r.datasetId, r.verdict, r.regressions]));
Output fields
| Field | Type | Description |
|---|---|---|
status | string (or null) | ok = a dataset checked (charged); no_data = nothing to check: no such dataset, nothing new since the last run of the watch list, the built-in example, or a dataset still being written (noDataReason says which) (free); failed = invalid input, no access to the dataset, or no answer (free) |
attempts | integer (or null) | Requests made for this entry (retries included) |
error | string (or null) | Why a row is no_data or failed; null on ok rows |
input | string (or null) | The entry that produced this row: a dataset ID, the watch list for the nothing-new row, or the example dataset |
mode | string (or null) | watch (compared with the watch list's last checks) or export (compared with this run's earlier datasets) |
watchList | string (or null) | The watch list name in watch mode; null in export mode |
example | boolean (or null) | true only on the built-in example's free rows (nothing of yours was read) |
noDataReason | string (or null) | Why a row is no_data: not_found (no such dataset), nothing_new (the watch list checked every dataset before), example (the built-in example), pending_dataset_write (written in the last 2 minutes and its item count had not settled: free, check it again in a minute); null on ok and failed rows |
datasetId | string (or null) | The checked dataset's ID |
datasetName | string (or null) | The dataset's name, or null for an unnamed dataset |
datasetCreatedAt | string (or null) | When the dataset was created (ISO 8601) |
baselineKey | string (or null) | What the dataset is compared with: task: |
actorId | string (or null) | The Actor whose run made the dataset (integration runs only) |
actorTaskId | string (or null) | The task whose run made the dataset (integration runs only) |
runId | string (or null) | The run that made the dataset (integration runs only) |
runStatus | string (or null) | How that run ended (integration runs only); anything but SUCCEEDED is a regression |
runStartedAt | string (or null) | When that run started (ISO 8601; integration runs only) |
runFinishedAt | string (or null) | When that run finished (ISO 8601; integration runs only) |
runTimeSecs | number (or null) | How long that run took, in seconds (integration runs only) |
runTimeBaselineSecs | number (or null) | The median run time of the baseline's checks, in seconds |
runTimeChangePercent | number (or null) | The run time's change against the baseline, in percent |
verdict | string (or null) | regression = at least one finding below; pass = none; baseline = the first check of its key (no earlier check to compare with) |
regressionCount | integer (or null) | How many findings the check has |
regressions | array (or null) | Each regression in one plain sentence (counts and field names only, never a value from a row) |
itemCount | integer (or null) | The dataset's item count (its own total, not the sample) |
itemCountBaseline | number (or null) | The median item count of the baseline's checks |
itemCountChangePercent | number (or null) | The item count's change against the baseline, in percent |
sampledRows | integer (or null) | How many rows were read for the field checks (the first ones, up to sampleSize) |
fieldCount | integer (or null) | Fields in the sample (a key in at least 2 rows and 1% of them) |
fieldsHidden | integer (or null) | Fields counted but never listed: a name that could be a person's, holds an @, or is over 80 characters |
schemaFingerprint | string (or null) | A short hash of the listed fields' names and types: the same value means the same schema |
fieldsAdded | array (or null) | Fields that are new against the baseline (not a regression on their own) |
fieldsRemoved | array (or null) | Fields of the baseline that are gone |
typeChanges | array (or null) | Fields whose most common type changed: "field: before → now" |
emptyRateChanges | array (or null) | Fields empty in more rows than in the baseline (by at least maxEmptyIncreasePoints) |
requiredFieldsMissing | array (or null) | Required fields (requiredFields) that are empty in some sampled rows, with the count |
failedRowsPercent | number (or null) | Sampled rows with a status of failed or error, or a filled error field, in percent |
failedRowsBaselinePercent | number (or null) | The baseline's median share of failed rows, in percent |
noDataRowsPercent | number (or null) | Sampled rows with a status of no_data, not_found or empty, in percent |
fieldProfile | array (or null) | Each listed field with its most common type and, when above 0, its empty share (at most 100 fields, by name) |
baselineChecks | integer (or null) | How many earlier checks the comparison used |
url | string (or null) | The dataset's page in Apify Console (the Store page for the built-in example) |
source | string (or null) | Where the data comes from |
scrapedAt | string (or null) | When the row was made (ISO 8601) |
Pricing
Pay per event: one dataset-check event per checked dataset (an ok row, whatever its verdict). no_data, failed and
example rows are free.
| Plan | Price per 1,000 checks |
|---|---|
| Free | $1.50 |
| Bronze | $1.30 |
| Silver | $1.15 |
| Gold (and higher) | $1.00 |
Apify also charges its small per-run start event. There are no usage fees on top.
Limits
- One request at a time to the Apify API, at least 250 ms apart. If the API answers HTTP 429, the run pauses 60, 120 and
240 seconds on the same connection, then stops with free
failedrows. No other token, IP or proxy is ever tried. - Each dataset: its item count, then its first
sampleSizerows (1,000 by default) in pages of up to 200 rows, fewer when rows are large (a page stays near 8 MB), at most 60 pages. Apify updates a dataset's item count a few seconds after the writes, so the count used is the larger of that count and the rows read. - At most 100 datasets per run.
Known issues
- The field checks read the first rows of a dataset: a change that shows only in later rows is not seen unless
sampleSizecovers it. The item count is always the dataset's own total. - Fields are top-level keys: a nested object is one field of type
object. - Run status and run time come only with an integration (a dataset alone does not say how its run went).
- Datasets are given by ID: a dataset's name (or username~name) is refused, free, because the picker grants read access by
ID only. As an integration, the input must name the run's dataset (
["{{resource.defaultDatasetId}}"]): a dataset that is only in the run's details cannot be read, and that row says so (free).
FAQ
Does it need my API token? No. It runs with Apify's limited permissions and reads only the datasets you pick: the
picker grants it read access to those, including the run's dataset an integration names in datasetIds.
Can it see my other data? No. Limited permissions let it read only its own storages, its watch list and the datasets you give it.
Are my rows returned or kept? No. Rows are read in memory to count fields, types and empty or failed rows. The output and the watch list hold counts, percentages, field names and type names only.
Why did a watch run return one no_data row? Every dataset it was given was checked before by that watch list; the
row says how many. It is free.
What is a pending_dataset_write row? The dataset was written in the last 2 minutes and its item count did not settle
within 30 seconds (Apify updates the count a little after the writes), so it was not checked. It is free, and it is not
recorded: run the check again with that dataset in a minute, and it is checked then.
Why is a field missing from fieldProfile? A key counts as a field when it is in at least 2 sampled rows and 1% of
them; one that could be a person's name, holds an @ or is over 80 characters is counted in fieldsHidden, never listed.
Data and licence
- Source: your own datasets, through the Apify API v2 (documented at https://docs.apify.com/api/v2), read with the run's own limited-permission token.
- Your data stays yours (Apify General Terms 5.8). This Actor returns only what it counted.
- This Actor is not affiliated with, endorsed by or provided by Apify.