Dataset Diff: Compare Datasets, Get New & Changed Rows avatar

Dataset Diff: Compare Datasets, Get New & Changed Rows

Pricing

from $0.70 / 1,000 changes

Go to Apify Store
Dataset Diff: Compare Datasets, Get New & Changed Rows

Dataset Diff: Compare Datasets, Get New & Changed Rows

Compare two Apify datasets or CSV, JSON or Excel files, or each run of a scraper with its last run: get only the new, changed and removed rows, with which fields changed and their old values. Chain it after any actor to monitor prices, stock, listings or jobs. Pay per change.

Pricing

from $0.70 / 1,000 changes

Rating

0.0

(0)

Developer

Michael Costa

Michael Costa

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

15 hours ago

Last modified

Share

What does Dataset Diff do?

Dataset Diff compares two Apify datasets or CSV, JSON or Excel files, or each run of a scraper with its last run, and returns only the new, changed and removed rows, with which fields changed and what they were before. Chain it after any actor to watch prices, stock, listings, jobs or any other data for changes.

It's the "what's different since last time?" step between a scheduled scraper and your alerts, spreadsheet or database: only the 12 products whose price moved, not all 5,000 again.

Input fields · API · Use it from Claude, ChatGPT or Cursor (MCP)

Jump to: Monitor a scraper · Fields · Price · How to use · Input · Output · AI agents (MCP) · Limits · FAQ

Try it in one click: the input comes pre-filled with the US Treasury's official exchange rates at two dates (2025-09-30 and 2025-12-31, public CSV files), compared by currency on the rate: 120 rates changed, 1 currency is new and 1 is gone, so 122 rows, about $0.12 (122 × $0.001, plus 167 rows compared × $0.00001 and $0.00005 for the run start). Then pick your own datasets or files.

Monitor any scraper: get only what changed since its last run

Leave Compare with empty and each run compares with what the same comparison saw on its previous run, remembered in your own Apify account. A run where nothing changed writes no rows and costs only the rows compared ($0.01 per 1,000) and the $0.00005 start fee.

  1. Open the scraper (or its saved task), go to the Integrations tab and add Apify actor → Dataset Diff, triggered when a run succeeds.
  2. Set the input to
    {"datasetId": "{{resource.defaultDatasetId}}", "keyFields": ["url"], "ignoreFields": ["scrapedAt"]}
    : Apify fills in the finished run's dataset, keyFields says what identifies a row (a URL, an id, or name, city), and ignoreFields lists what the scraper stamps on every run, so it doesn't count as a change.
  3. Every successful scrape now produces a dataset with only the new, changed and removed rows. The first run reports every row as new (set First run with no earlier data to "Only remember the rows" to skip that).
  4. To get them in Slack, email or a webhook, open the Integrations tab of Dataset Diff (or of its saved task) and add one, triggered when a run succeeds: Slack with {{resource.statusMessage}} and a link to https://console.apify.com/storage/datasets/{{resource.defaultDatasetId}}, Gmail with the dataset attached, or an HTTP webhook whose resource.defaultDatasetId is the dataset of changes.

The memory follows the actor that wrote the dataset (so each scraper run's new dataset id doesn't matter), or the saved task running Dataset Diff, or Comparison name when you set one. Changes cut by Max rows per run or your maximum cost per run aren't lost: the next run reports them.

What data does Dataset Diff return?

Your own rows, with three fields added. From the example:

FieldExampleNotes
(your fields)"country": "Afghanistan", "exchange_rate": "65.96"The new row as it is; for a removed row, the old one.
changeTypechangednew, changed, removed (or unchanged, if you ask for those too).
changedFields["exchange_rate"]The compared fields that differ, for a changed row; empty otherwise.
previousValues{"exchange_rate": "67.33"}What the changed fields held before; null for new and removed rows.

The run's RUN_STATS record counts the rows read, compared, new, changed, removed and unchanged, rows without a key and repeated keys. More under Output.

How much does it cost to compare datasets?

You pay for each row written (a change) and for the rows compared: $1.00 per 1,000 changes and $0.01 per 1,000 rows compared, plus $0.00005 each time a run starts.

Changes are cheaper on paid Apify plans: $0.90 per 1,000 on Starter, $0.80 on Scale and $0.70 on Business. The prices on this page are the Free-plan price, so on a paid plan you pay less than the examples show.

  • The example: 122 changes × $0.001 + 167 rows compared × $0.00001 = about $0.12, plus the start fee.
  • A month, for example: a daily product scraper of 5,000 items where 50 change a day: 30 × (5,000 × $0.00001 + 50 × $0.001) = $3.00, plus 30 starts ($0.0015). The same with nothing changing: $1.50.
  • Caps: Max rows per run in the input, and Maximum cost per run in the run options. The run stops cleanly at whichever comes first, and what it didn't write is reported next run.

Never charged: unchanged rows (unless you ask for them), rows without a key, repeated keys, and the older data you compare with.

How to compare two datasets or files

  1. Open Dataset Diff and click Try for free (or Start if you're signed in).
  2. Pick the newer data in New data: Apify dataset (or paste a link in Or: new data file URL).
  3. Pick the older data in Compare with (a dataset or a file link), or leave it empty to compare with the last run.
  4. Set Key fields (what identifies a row, e.g. url), and if you like Fields to compare (e.g. price, stock) or Fields to ignore (e.g. scrapedAt).
  5. Click Start, then open the Output tab and export as JSON, CSV or Excel.

Example: exchange rates that changed in a quarter

The pre-filled input (two public CSV files, compared by currency on the rate):

{"fileUrl": "https://api.fiscaldata.treasury.gov/services/api/fiscal_service/v1/accounting/od/rates_of_exchange?format=csv&filter=record_date:eq:2025-12-31&page[size]=300",
"previousFileUrl": "https://api.fiscaldata.treasury.gov/services/api/fiscal_service/v1/accounting/od/rates_of_exchange?format=csv&filter=record_date:eq:2025-09-30&page[size]=300",
"keyFields": ["country_currency_desc"], "compareFields": ["exchange_rate"]}

It compared 167 currencies (7 listed twice were counted once) and wrote 122 rows. One of each kind (real output from a local run on 2026-10-03; the Treasury's bookkeeping columns left out):

[
{"country": "Afghanistan", "currency": "Afghani", "exchange_rate": "65.96", "record_date": "2025-12-31",
"changeType": "changed", "changedFields": ["exchange_rate"], "previousValues": {"exchange_rate": "67.33"}},
{"country": "Cyprus", "currency": "Euro", "exchange_rate": "0.851", "record_date": "2025-12-31",
"changeType": "new", "changedFields": [], "previousValues": null},
{"country": "Belarus", "currency": "New Ruble", "exchange_rate": "3.025", "record_date": "2025-09-30",
"changeType": "removed", "changedFields": [], "previousValues": null}
]

Ready-to-run examples

Each example opens with the input already filled in. Run it as it is, or change the input first.

Input

FieldWhat it does
New data: Apify datasetThe latest data: one of your datasets, picked from the list (or its id or name). Read-only.
Or: new data file URLA direct link to a CSV/TSV, JSON, JSON Lines or Excel (.xlsx, first sheet) file, up to 200 MB.
Compare with (dataset or file URL)The older data. Leave both empty to compare with the last run.
Key fieldsWhat identifies a row: url, id, or several (name, city). Dots reach nested fields (product.sku). Empty: whole rows.
Fields to compareOnly these make a row changed (e.g. price, stock). Empty: every top-level field.
Fields to ignoreNever make a row changed (e.g. scrapedAt).
Report new / changed / removed / unchanged rowsWhich kinds to write (new, changed and removed by default).
First run with no earlier dataReport every row as new (default), or only remember them.
Comparison nameKeeps a comparison's memory apart from others (or starts afresh).
Max rows per runCap the rows written; the rest is reported next run.

Values compare as they read: 12 and "12" are equal (so a CSV export and a dataset compare fine), and so are null, "", [], {} and a missing field; text that differs only in how an accent is encoded is the same. Objects and lists compare as a whole.

A full example, for a product scraper chained into Dataset Diff:

{
"datasetId": "{{resource.defaultDatasetId}}",
"keyFields": ["url"],
"compareFields": ["price", "inStock"],
"includeRemoved": true,
"firstRun": "baseline"
}

Output

One dataset row per change: the row's own fields plus changeType, changedFields and previousValues. A field of yours with one of those three names is overwritten. The RUN_STATS record has the counts. The dataset can be downloaded as CSV, JSON, Excel, XML or HTML at any time, or passed on to the next actor.

Run it on a schedule, or from your own code

Chained after a scraper (above) is the usual set-up. Otherwise, save your input as a task and add it to a schedule, for example daily after a CSV feed you watch is updated. Collect the changes from the API (GET https://api.apify.com/v2/actor-tasks/<task id>/runs/last/dataset/items?status=SUCCEEDED, with your API token), a webhook, or Make, Zapier and n8n through Apify's integrations.

Can I use Dataset Diff from an AI agent (MCP)?

Yes, through Apify's MCP server, from Claude, ChatGPT, Cursor or any other MCP client. Add this to your client's MCP configuration (or let the agent find it with the server's actor search); your client signs you in to Apify:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=humble-echidna/dataset-diff"
}
}
}

To use an Apify API token instead of signing in, add "headers": {"Authorization": "Bearer <APIFY_TOKEN>"} next to url.

An agent that has run a scraper twice can pass both datasets, e.g.

{"datasetId": "<new>", "previousDatasetId": "<old>", "keyFields": ["url"]}
, and get back only what changed.

Who it's for

Anyone who runs a scraper on a schedule and only cares about the difference: new listings or jobs, price and stock changes, products that disappeared, companies that changed a field. Also for comparing two exports (CSV, Excel) of the same table.

Why this one?

  • Works with any actor's data. Your own datasets or CSV/TSV, JSON, JSON Lines and Excel files, with any key, nested fields included. No scraper of its own, so nothing on a website can break it.
  • Says what changed. Each changed row lists the fields that differ and their old values; removed rows come back as they were.
  • Never loses a change. The memory only moves on for rows you actually got, so a run cut by your limits reports the rest next time.
  • Cheap when nothing happens. A quiet run costs $0.01 per 1,000 rows compared; changes $1.00 per 1,000 ($0.90 on Starter), with no compute, proxy or storage bill on top.
  • Tells you what happened. The status counts new, changed, removed and unchanged rows, and warns about field names no row has (a typo in a key field is the most common mistake).
  • Polite and safe with files. It identifies itself honestly (User-Agent HumbleEchidnaApify), follows each site's robots.txt, and only requests public web addresses on the standard ports (80 and 443).

Limits

  • Up to 100,000 rows on each side per run; the status says so if the new data has more (filter it first with Dataset Transformer). Very wide or long rows can stop a run earlier at the default 1 GB memory; the status says to give the run more memory.
  • Files: up to 200 MB (after decompression), public web addresses on ports 80 and 443; Excel's first sheet only.
  • A key appearing twice in the same data counts once (the first row); rows with every key field empty are skipped.
  • Removed rows are only checked when every row of the new data was read.
  • The last run's snapshot is kept in a key-value store in your account (dataset-diff-memory); rows over 20 KB keep only their key, so their old values aren't shown.
  • A single output row can be up to 5 MB; larger ones are skipped and counted.

FAQ

What if my rows have no single id?

Use several key fields (name, city), or leave Key fields empty to compare whole rows: then a row that changed shows up as one removed and one new row.

Why is every row "new"?

It's the comparison's first run, or its memory belongs to another set-up: the key, compared and ignored fields, and the comparison's name are part of what's remembered. Changing them starts a new comparison.

Why are rows "changed" when nothing important moved?

The scraper writes something new on every run (a timestamp, a position, a session id). Add it to Fields to ignore, or list only the fields you care about in Fields to compare.

Can it compare datasets from different actors?

Yes: put one in New data and the other in Compare with. Both are read with your own account's access.

It only processes data you choose: your own Apify datasets, and files at addresses you supply, fetched politely (robots.txt honoured, honest User-Agent). Make sure you're allowed to use the files you point it at. The example data is the US Treasury's Fiscal Data API, which is "offered free, without restriction" for commercial and non-commercial use.

Something that used to work now fails. Why?

The run log and the status say what went wrong. Please open an issue with the input you used.

ActorUse it when
Dataset TransformerYou want to filter, dedupe or reshape the changes (or the data before comparing it), or save them as CSV.
Join DatasetsYou want to add fields from another dataset to each row, matched by a key (VLOOKUP for datasets).
Website Change MonitorYou want to watch web pages themselves for changes, with a diff of the text.

Feedback and support

Found a bug, or need a comparison that isn't here? Open an issue on the Issues tab with the input you used.

Versions

Current version: 0.1. See the Changelog tab for what changed in each version.