# Changelog of Dataset Diff: Compare Datasets, Get New & Changed Rows (`humble-echidna/dataset-diff`) Actor

- **URL**: https://apify.com/humble-echidna/dataset-diff/changelog.md
- **Full Actor documentation**: https://apify.com/humble-echidna/dataset-diff.md

## Changelog

Versions follow MAJOR.MINOR.PATCH (`src/version.py`); Apify shows MAJOR.MINOR from `.actor/actor.json`.
Every run logs its version and records it in the `RUN_STATS` key-value record.

### 0.1.2 (2026-10-04)

- README: the ready-to-run Store examples.

### 0.1.1 (2026-10-04)

- The status counts the older rows before matching (it said "compared with 0 rows" in last-run mode).

### 0.1.0 (2026-10-03)

First release.

- Compares new data (an Apify dataset, or a CSV/TSV, JSON, JSON Lines or Excel file at a URL) with older data (the
  same kinds), or with what the same comparison saw on its last run, and writes the new, changed and removed rows
  with `changeType`, `changedFields` and `previousValues`.
- Rows are matched by key fields (nested paths allowed) or whole rows; compared on all fields or the ones named,
  minus ignored ones; values compare as they read (12 = "12"; null = "" = \[] = {} = missing; accents normalized).
- Last-run memory in the user's own `dataset-diff-memory` store, named after the comparison name, the saved task, the
  actor that wrote the dataset, or the file URL, plus the key/compare/ignore set-up. Snapshots are written in gzipped
  parts under a new generation before the old one is deleted; a quiet run doesn't rewrite it.
- Memory only moves on for what the user got: changes cut by max rows, max cost or a source failing midway are
  reported next run. Removed rows are only checked when every new row was read.
- Charged per change (`apify-default-dataset-item`) and per row compared (`row-compared`, charged before each 1,000
  rows are compared, so the maximum cost per run stops the work, not just the writing).
- Limits: 100,000 rows per side, a memory budget of 40% of the run's memory for the snapshots, 200 MB files, 5 MB
  output rows.
