# Dataset Change Monitor (`egra_van/dataset-change-monitor`) Actor

Compare every scraper run with the previous one and get only new, changed and removed items, with field-level before/after diffs. Alerts via Telegram, Slack, webhook or email. Works with any Apify Actor, dataset, JSON or CSV.

- **URL**: https://apify.com/egra\_van/dataset-change-monitor.md
- **Developed by:** [Argentin Vazdautan](https://apify.com/egra_van) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.70 / 1,000 change detected (new, changed or removed item)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Dataset Change Monitor

Run your scraper on a schedule and **get only what changed since the last run**. This Actor compares each new run's results with the previous one and outputs **new items, changed items (with before → after values per field) and removed items**. You can also get an **alert on Telegram, Slack, email or any webhook**.

It works with **any Apify Actor or dataset**, and also with JSON or CSV files. Typical uses: price monitoring, stock and availability tracking, new job postings, new real-estate listings, competitor catalogs, new reviews, and SEO or content change detection.

### Why use it

- 🆕 **New items**: new products, listings, jobs or posts since the last run
- ✏️ **Changed items** with a field-level diff: `price: 19.99 → 17.49`, `stock: true → false`
- ❌ **Removed items**: sold out, delisted or deleted
- 🔔 **Alerts** only when something changed: Telegram, Slack, webhook (Make, Zapier, n8n) and email
- 🎯 **You choose what counts as a change**: watch only `price` and `stock`, or everything except `scrapedAt`
- 🔌 **No code**: add it as an integration to your scraper and it runs after every scrape
- 📦 **Handles large datasets**: reads page by page and keeps a compact, compressed snapshot between runs
- 💸 **Pay only for results**: a small fee per 1,000 items compared plus a fee per reported change

### Quick start: run it after every scrape (recommended)

1. Open your scraper (or saved task) in Apify Console → **Integrations** → **Add integration** → **Dataset Change Monitor**.
2. Set **Unique key fields** (e.g. `url`), **Fields to ignore** (e.g. `scrapedAt`) and, if you want, your Telegram or Slack details.
3. Schedule the scraper as usual. After each run, this Actor compares the new results with the previous ones.

The first run saves a **baseline** and reports nothing (unless you enable *Report all items as new on the first run*). From the second run on, you only get the differences.

When used as an integration, the dataset of the finished run is detected automatically, and the monitor name defaults to the scraper's task or Actor ID, so every scraper keeps its own history.

### Other ways to provide data

| Input | When to use |
|---|---|
| `datasetId` | Compare a specific dataset (ID or name). |
| `actorRunId` | Compare the default dataset of a given run. |
| `datasetUrl` | A public JSON, JSON Lines or CSV link, e.g. an Apify dataset export URL or a Google Sheet published as CSV. |
| `items` | Pass an array directly (testing, Make, Zapier, n8n). |

### Input example

```json
{
  "datasetId": "aBcD1234efGh5678",
  "monitorName": "amazon-laptops",
  "idFields": ["asin"],
  "compareFields": ["price", "stock", "title"],
  "trackRemoved": true,
  "telegramBotToken": "123456:ABC-your-bot-token",
  "telegramChatId": "123456789",
  "notifyMaxItems": 10
}
```

#### Matching options

- **Monitor name**: every name keeps its own memory of the previous run. Use one name per scraper, search or list you track. Runs with the same name are compared with each other.
- **Unique key fields** (`idFields`): identify the same item across runs, e.g. `url`, `id`, `asin`, `sku`. Several fields are combined, and nested fields use dots (`offer.id`). If you leave it empty, the whole item is the key, so items can only be *new* or *removed*, never *changed*.
- **Fields to watch** (`compareFields`): only these fields decide whether an item changed. Leave empty to watch all fields.
- **Fields to ignore** (`ignoreFields`): fields that change on every run, such as `scrapedAt`, `timestamp` or `position`.

If you change the key, watched or ignored fields, the next run starts a new baseline automatically, because old and new snapshots are no longer comparable.

### Output

Each change is one dataset item:

```json
{
  "changeType": "changed",
  "key": "https://shop.example/p/123",
  "changedFields": ["price", "stock"],
  "changes": [
    { "field": "price", "before": 19.99, "after": 17.49 },
    { "field": "stock", "before": true, "after": false }
  ],
  "totalChangedFields": 2,
  "item": { "url": "https://shop.example/p/123", "title": "…", "price": 17.49, "stock": false },
  "monitorName": "shop-example",
  "detectedAt": "2026-09-26T08:00:00.000Z"
}
```

- `changeType`: `new`, `changed` or `removed`
- `item`: the current item. For `removed`, it holds the last known values of the compared fields.
- `changes`: only for `changed`. Long values are shortened.

The `OUTPUT` record in the run's key-value store contains a summary: counts of new, changed, removed and unchanged items, items compared, notification results, and whether the snapshot was updated.

### Notifications

Notifications are optional and are sent only when something changed (or after every run if you enable *Also notify when nothing changed*). Messages look like this:

```
🔔 amazon-laptops: 3 new, 2 changed, 1 removed (1,240 items checked)
🆕 NEW: B0CX23V2ZK
✏️ CHANGED: B0BSHF7WHW
    price: 999 → 899
❌ REMOVED: B0C1234567
…and 2 more
Full results: https://console.apify.com/storage/datasets/…
```

- **Telegram**: create a bot with [@BotFather](https://t.me/BotFather), paste the token, and set your chat ID. For groups and channels, add the bot first.
- **Slack**: create an [incoming webhook](https://api.slack.com/messaging/webhooks) and paste its URL.
- **Webhook**: receives a `POST` with JSON: `event`, `monitorName`, `summary`, `message`, `resultsUrl` and the first changes (`webhookMaxItems`, default 100). Great for Make, Zapier, n8n, Google Apps Script or your own backend.
- **Email**: sent through the official `apify/send-mail` Actor. That call runs on your Apify account as a separate, very small run (usually a fraction of a cent). Separate several addresses with commas.

The bot token, Slack URL and webhook URL are stored as **secret inputs** and are encrypted by Apify.

A failed notification never fails the run: the result is shown in the log and in `OUTPUT.notifications`.

### Pricing

Pay per event:

- **Items compared**: charged per started 1,000 items read from your data.
- **Change detected**: charged per reported new, changed or removed item.

A run where nothing changed only costs the comparison. Change types you switch off are not reported and not charged. If a run hits your **maximum cost per run**, it stops cleanly, keeps what it already reported and **does not update the snapshot**, so no change is lost. The next run compares against the same previous data.

### Safety features

- **Empty result protection**: if a scrape returns 0 items (for example because it was blocked), nothing is reported as removed and the previous snapshot is kept. Enable *Allow an empty source* if empty results are real.
- **Crash-safe state**: the new snapshot is written completely before it replaces the old one.
- **Duplicates**: when the same key appears twice in one run, only the first item is used, and the log tells you so.
- **Items without a key** are skipped and counted.

### Where the history is stored

Snapshots live in a named key-value store in your account, `change-monitor-state` by default, one entry per monitor name. Named storages are kept until you delete them. To start over, enable **Start over** for one run, or delete the `state-<monitor-name>-…` records.

### Limits

- Only fields up to 4 levels deep are listed individually in `changes`. Deeper differences and arrays are shown as one changed field.
- Very long values are stored as a fingerprint plus a 200-character preview. Changes are still detected exactly, but `before` shows only the preview.
- The snapshot is held in memory while comparing. About 100,000 typical items fit in 1 GB of memory. For millions of items, give the run more memory, or turn off *Show which fields changed* to keep only fingerprints.
- `datasetUrl` files are downloaded whole, so use `datasetId` for very large data.
- With *Max items to compare*, removals are not reported for that run, because unseen items might still exist.

### FAQ

**Why did the first run report nothing?** The first run saves the baseline. Enable *Report all items as new on the first run* if you want everything reported.

**Every item shows as changed.** A field changes on every run, for example a timestamp, rank or session ID. Add it to *Fields to ignore*, or list only the fields you care about in *Fields to watch*.

**Every item shows as new.** The key is not stable. Pick a stable field such as a product ID or canonical URL as *Unique key fields*.

**Can I monitor several scrapers?** Yes. Use a different *Monitor name* for each one. As an integration this happens automatically.

**Can I use it outside Apify?** Yes. Send your items with `items` or `datasetUrl` through the Apify API, Make, Zapier or n8n, and use the same monitor name each time.

# Actor input Schema

## `datasetId` (type: `string`):

ID or name of the Apify dataset to check, e.g. the default dataset of your scraper's latest run. Leave empty when this Actor runs as an integration after another Actor: the finished run's dataset is used automatically.

## `actorRunId` (type: `string`):

ID of a finished Actor run. Its default dataset is compared.

## `datasetUrl` (type: `string`):

Public link to a JSON array, JSON Lines or CSV file, e.g. an Apify dataset export link (https://api.apify.com/v2/datasets/…/items?format=json).

## `items` (type: `array`):

Array of objects to compare. Handy for testing or when calling the Actor from Make, Zapier or n8n.

## `monitorName` (type: `string`):

Each name keeps its own memory of the previous run. Use a different name for every scraper or search you track, e.g. "amazon-laptops". If empty, the name is taken from the source Actor or task (integration) or "default".

## `idFields` (type: `array`):

Field(s) that identify the same item between runs, e.g. url, id or asin. Nested fields use dots (offer.id). If empty, the whole item is the key: items can then only be new or removed, never changed.

## `compareFields` (type: `array`):

Only these fields decide whether an item changed, e.g. price, stock, title. Leave empty to watch all fields except the ignored ones.

## `ignoreFields` (type: `array`):

Fields that change on every run and are not real changes, e.g. scrapedAt, timestamp, position. Used when "Fields to watch" is empty.

## `reportNew` (type: `boolean`):

Items whose key did not exist in the previous run.

## `reportChanged` (type: `boolean`):

Items whose watched fields differ from the previous run.

## `trackRemoved` (type: `boolean`):

Items that were in the previous run but are missing now.

## `reportAllOnFirstRun` (type: `boolean`):

By default the first run only saves a baseline and reports nothing.

## `detailedDiff` (type: `boolean`):

Keeps a compact copy of each item between runs. Turn off for very large datasets to use less memory; changed items are still detected.

## `maxChangesPerItem` (type: `integer`):

Limits the size of the "changes" list for items with many modified fields.

## `maxItems` (type: `integer`):

0 = no limit. When the limit is hit, removed items are not reported for that run.

## `notifyMaxItems` (type: `integer`):

How many changes Telegram, Slack and email messages list before "…and N more".

## `notifyOnNoChanges` (type: `boolean`):

Sends a short "no changes" message after every run, useful to know the monitor is alive.

## `telegramBotToken` (type: `string`):

Create a bot with @BotFather and paste its token.

## `telegramChatId` (type: `string`):

Your user, group or channel ID (e.g. 123456789 or -1001234567890). Add the bot to the group or channel first.

## `slackWebhookUrl` (type: `string`):

https://hooks.slack.com/services/…

## `webhookUrl` (type: `string`):

Receives a POST with a JSON summary and the first changes. Works with Make, Zapier, n8n or your own server.

## `webhookMaxItems` (type: `integer`):

Maximum number of change records included in the webhook JSON. The full list is always in the dataset.

## `emailTo` (type: `string`):

Sends a summary email through the apify/send-mail Actor (billed as a separate small run on your account). Several addresses: separate with commas.

## `emailSubject` (type: `string`):

Default: "<monitor>: 3 new, 1 changed, 0 removed".

## `resetState` (type: `boolean`):

Forget the previous snapshot of this monitor and create a new baseline.

## `allowEmptySource` (type: `boolean`):

By default an empty dataset is treated as a failed scrape: nothing is reported as removed and the snapshot is kept. Enable if an empty result can be real.

## `stateStoreName` (type: `string`):

Named key-value store that keeps the snapshots between runs. Letters, digits and "-" only.

## `payload` (type: `object`):

Filled automatically when this Actor runs as an integration after another Actor. Leave empty.

## Actor input object example

```json
{
  "items": [
    {
      "url": "https://shop.example.com/p/1",
      "name": "Blue mug",
      "price": 12.5
    },
    {
      "url": "https://shop.example.com/p/2",
      "name": "Red mug",
      "price": 9.9
    }
  ],
  "monitorName": "demo-monitor",
  "idFields": [
    "url"
  ],
  "ignoreFields": [
    "scrapedAt",
    "crawledAt",
    "timestamp"
  ],
  "reportNew": true,
  "reportChanged": true,
  "trackRemoved": true,
  "reportAllOnFirstRun": false,
  "detailedDiff": true,
  "maxChangesPerItem": 20,
  "maxItems": 0,
  "notifyMaxItems": 10,
  "notifyOnNoChanges": false,
  "webhookMaxItems": 100,
  "resetState": false,
  "allowEmptySource": false,
  "stateStoreName": "change-monitor-state"
}
```

# Actor output Schema

## `results` (type: `string`):

All result items of this run.

## `summary` (type: `string`):

Summary of the run with counts and download links.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "items": [
        {
            "url": "https://shop.example.com/p/1",
            "name": "Blue mug",
            "price": 12.5
        },
        {
            "url": "https://shop.example.com/p/2",
            "name": "Red mug",
            "price": 9.9
        }
    ],
    "monitorName": "demo-monitor",
    "idFields": [
        "url"
    ],
    "ignoreFields": [
        "scrapedAt",
        "crawledAt",
        "timestamp"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("egra_van/dataset-change-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "items": [
        {
            "url": "https://shop.example.com/p/1",
            "name": "Blue mug",
            "price": 12.5,
        },
        {
            "url": "https://shop.example.com/p/2",
            "name": "Red mug",
            "price": 9.9,
        },
    ],
    "monitorName": "demo-monitor",
    "idFields": ["url"],
    "ignoreFields": [
        "scrapedAt",
        "crawledAt",
        "timestamp",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("egra_van/dataset-change-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "items": [
    {
      "url": "https://shop.example.com/p/1",
      "name": "Blue mug",
      "price": 12.5
    },
    {
      "url": "https://shop.example.com/p/2",
      "name": "Red mug",
      "price": 9.9
    }
  ],
  "monitorName": "demo-monitor",
  "idFields": [
    "url"
  ],
  "ignoreFields": [
    "scrapedAt",
    "crawledAt",
    "timestamp"
  ]
}' |
apify call egra_van/dataset-change-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,egra_van/dataset-change-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gxu7wWDHYEppTFCV9/builds/86fGgUcesbCfyK3BS/openapi.json
