# Changelog of Car data scraper: specs, 0-60 times and power (`mrbridge/car-data-scraper`) Actor

- **URL**: https://apify.com/mrbridge/car-data-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/mrbridge/car-data-scraper.md

## Changelog

### 0.1.41 (2026-09-25)

*A run that finds nothing now says so.*

- **An unknown make or model is reported instead of ending in silence.** Such a run used to finish SUCCEEDED with zero rows and no message. It now writes one error row with the code `TARGET_NOT_FOUND`, ends with `stopReason: "badInput"` and a `stopMessage` naming what did not match, and the run header reads "Stopped (nothing matched)". The error row is not charged.
- **A listing page served as an anti-bot page now counts as a refused request.** It is retried first. A run blocked there ends FAILED instead of succeeding with an empty dataset.
- **A listing page skipped after its retries is no longer dropped silently.** It adds an error row carrying its URL, not charged.
- The dataset schema now lists the four error codes the Actor emits, `TARGET_BLOCKED`, `EXTRACTION_FAILED`, `TARGET_NOT_FOUND` and `BAD_INPUT`, instead of two codes it no longer uses.
- Billing is unchanged.

### 2026-08-30

*Restored: auto-data changed its HTML and the Actor read it as a block.*

- **The Actor stopped returning data between 2026-08-26 and 2026-08-30, and it was not a
  block.** auto-data.net rebuilt its spec sheet, moving every label and value out of
  `<table class="cardetailsout"><tr>` and into `<div class="cardetailsout">` with
  `<div class="row"><div class="par">label</div><div class="val">value</div></div>`. Not
  one `<tr>` is left on the page. The parser's selector matched nothing, no field was
  extracted, and the "zero fields" guard reported `TARGET_BLOCKED`. The site had been
  answering HTTP 200 with complete pages the whole time.
  **The parser now reads both shapes**, the current one and the legacy one, using the
  same label mappings as before. Nothing about the extracted values changes: the real
  captured page yields 62 fields and a complete identity.
- **A run that delivers nothing because the source refused it now FAILS.** It used to
  finish SUCCEEDED with exit code 0 over a dataset containing only error rows, which told
  every integration downstream that all was well. A run that delivers some rows and loses
  others stays successful and reports itself as partial.
- **A blocked run now stops in seconds instead of minutes.** A run capped at 10 records
  produced 357 error rows over 216 seconds, because the cap counts delivered rows and a
  blocked run delivers none. The run now stops after a few consecutive refusals from the
  backbone. The cap keeps its meaning: it counts records, not failures.
- **New `stopReason: "targetBlocked"` and a new `blockedBackboneRequests` count in the
  `OUTPUT` record**, so a refusal is distinguishable from a timeout and from an internal
  error, which need different responses.
- **`totalErrorRows` was reported as 0 while error rows were being written.** It was only
  counted on the invalid-input path. It now counts every error row, after the write
  confirms. Error rows are still never charged.
- **The error rows view was half empty.** `make`, `model`, `generation` and `modification`
  are never present on an error row, because the failure happens before a variant is
  identified, and `url`, the address that failed, was missing. The view is now `error`,
  `errorMessage`, `source`, `url`, `scrapedAt`: five columns that are always filled.
- **`lastFullyCompletedModel` claimed a model on runs that produced no rows.** It stays
  null unless at least one record was delivered.
- **A page that fails to parse is no longer cached.** A blocked or unreadable response
  used to be stored for thirty days, so re-running could not recover even after the site
  came back. Only a usable page is cached now, and a stored page that no longer validates
  is dropped.
- Every response is now classified as it arrives, and the reason appears in the log and in
  the error row: an anti-bot page, a markup change, an explicit refusal and a truncated
  response are told apart rather than all reported as a block.

### 2026-08-28

*The demo video.*

- **A video walkthrough at the top of this README.** A clickable thumbnail sits directly under
  the title, and a second link in the how-to section for readers who scroll past it. The
  walkthrough runs six minutes and covers the input form, a live run, the four result views and
  the export. Documentation only: no code, schema, pricing or Store field changed.

### 2026-08-24

*Electrified powertrains, honest copy, and observability.*

#### Changed

- **Electric and hybrid variants now carry their power, and one existing column changes
  value on hybrids.** auto-data publishes the total output of an electrified variant under
  `System power` and `System torque`, rows this Actor never read, while the `Power` row it
  did read describes the combustion engine alone. A battery-electric variant therefore
  returned no power at all, and a hybrid returned the wrong one. `powerHp` and `torqueNm`
  now carry the system figure whenever the page prints one.
  **This changes a value that was already present, on hybrids only.** Measured on the
  fixtures: a Chevrolet Volt II 1.5 Plug-in Hybrid published `powerHp` 101 and now
  publishes 150, which is the figure its own variant name quotes and the figure auto-data
  answers to "what is the power output". A Ford Mondeo IV Hybrid published `torqueNm` 173
  and now publishes 300. **Combustion and mild-hybrid rows are unchanged**, because those
  pages print no `System power` row: an Audi TT RS stays at 394 hp and 480 Nm, a Mercedes
  E 220d Mild Hybrid at 197 hp and 440 Nm.
  **Nothing published is overwritten and no value is computed.** The nested `engine`
  section keeps `power_hp` and `torque_nm` exactly as the `Power` and `Torque` rows printed
  them, beside the new system keys, so the combustion figure of a hybrid is still there.
  The `canonicalVariantId` reads the nested figure and therefore does not move.
  The matcher also reads the nested figure, so cross-source matching behaviour is
  unchanged by this release.

- The progress counter is now monotonic. A heartbeat could report `collected 0 variant(s)`
  after `Collected 8 variant(s)` had already been shown, because it read a counter the
  pipeline only filled at the end, and because two status writes could settle out of order.
  Both are closed: one shared counter that never decreases, one serialized write path, and
  each message rendered from the counter at the moment it is written.

- The terminal status of a run stopped by its timeout now reads
  `Stopped (run timeout reached)` rather than `Stopped (deadline)`, and states that
  delivered rows are kept. `stopReason` in `OUTPUT` is still `deadline`, so integrations
  reading it are unaffected.

- The **Input mode** field now defaults to `Model names list`, the same value the form is
  prefilled with. An API call that omitted the field and supplied both a brand and a models
  list used to expand the whole brand. A call that supplies only a brand, or only models,
  behaves exactly as before, so existing tasks and integrations are unaffected.

- The **Max items** description now states that run time follows the number of variants a
  model carries rather than the cap alone.

- Power figures from auto-data are now compared against carperformancedata using two
  anchors, the published figure and its metric-to-SAE conversion, whichever is nearer. The
  published `power_hp` is untouched: the rule applies inside matching only, and only to that
  one source pair, whose two possible conventions are documented.

- **`dedupeConfidence` can change as a result, even when the matched variant and its
  confidence are identical.** On the evaluation corpus one row moves, from 0.96 to 1.00,
  with the same winning variant and the same `exact` confidence. No row changes which
  variant it matched, and no `power_hp` value changes.

#### Added

- The run summary now reports enrichment coverage per source, and separately where each
  source's candidates stopped. The two are counted in different units, delivered rows and
  candidate records, and are never combined.
- A source that delivered rows but enriched none of them is named in the run's terminal
  status message. It is informative and never changes the run's outcome.
- **Four new flat columns, all from auto-data.** `batteryCapacityKwh` is the gross battery
  capacity, `electricRangeKm` is the all-electric range as published under `Electric range`
  or `Electric range (WLTP)`, and `powerBasis` and `torqueBasis` name which auto-data row
  filled `powerHp` and `torqueNm`, with the value `system` or `engine`. All four are `null`
  where the source publishes nothing, and none of them is estimated or converted.
- **Seven new nested keys** in the `engine` and `performance` sections: `system_power_hp`,
  `system_power_rpm`, `system_torque_nm`, `system_torque_rpm`, `electric_motor_power_hp`,
  `electric_motor_torque_nm`, `battery_capacity_kwh` and `electric_range_km`.
- `OUTPUT` now carries **`sourcesMatched`**, the number of delivered rows carrying each
  enabled source. `sourcesEnabled` says what was configured, `sourcesMatched` says what
  contributed, and a run with four sources enabled and three zeros is a single-source run.
  It is derived from the coverage ledger rather than counted separately, so the two can
  never disagree.

### 2026-08-11

*More of the spec sheet, and provenance you can audit.*

- **CO2 is no longer a single number of unknown provenance.** auto-data publishes NEDC and WLTP figures under labels that differed only by a parenthetical, and the parser collapsed them. They are now separate: `fuelEconomy.co2_wltp_gkm` and `fuelEconomy.co2_nedc_gkm` are **new nested keys**, and the **new flat column `co2Standard`** names which standard the reference figure follows.
  **`co2GKm` and `fuelEconomy.co2_gkm` keep their exact previous values.** They remain the reference figure, WLTP when the source publishes one and NEDC otherwise. This release does not change either of them for any row.
- **`measurementKind` is a new provenance property** on a traced value, present only when the source states how a figure was obtained. It never defaults to `measured`, because a source that says nothing has not told you it measured anything. Its single flat mirror is the **new column `zeroToSixtyMphMethod`**, which is `null` when the source did not say.
- **`upstreamSourceName` and `upstreamSourceUrl` are new provenance properties** on the `sourcesDetail` entries. carperformancedata.com republishes figures measured by someone else and credits that publication on some of its records. That credit is now preserved. `source` stays the technical id and `sourceUrl` stays the page actually read, so the two describe the upstream publication and nothing else. Both are absent when no credit is published, and the other three sources publish none.
- **Four fields that already existed but were always empty now carry values.** `compression_ratio`, `bore_mm`, `stroke_mm` and `trunk_l` (boot volume, seats in place) were declared before this release and nothing populated them, so they never appeared in a record. They are now read from auto-data.
  **This changes what you see in an existing column**: the flat column **`compression`** was `null` on every row and is now filled wherever auto-data publishes a compression ratio, as `"10:1"`. On a whole-BMW run that is 970 of 1000 rows. **No value that was already present has been altered**: the change is empty to filled, never one value to another.
- **Eighteen genuinely new nested keys from auto-data**, none promoted to a flat column: engine internals (`injection_system`, `valvetrain`, `engine_layout`, `engine_configuration`, `powertrain_architecture`), servicing capacities (`oil_capacity_l`, `coolant_capacity_l`), payload and towing geometry (`gross_weight_kg`, `payload_kg`, `roof_load_kg`, `width_mirrors_mm`, `front_overhang_mm`, `rear_overhang_mm`, `turning_circle_m`), the boot maximum `trunk_max_l` (kept separate from the minimum so the two can never be mixed), transmission prose (`drivetrain_architecture`) and chassis text (`power_steering`, `wheel_rims`).
- **Three kinds of addition, kept distinct.** New nested keys live in the six traced sections. New flat columns are only `co2Standard` and `zeroToSixtyMphMethod`. New provenance properties are `measurementKind`, `upstreamSourceName` and `upstreamSourceUrl`, and they hang off values and source entries rather than being fields of their own.
- **No column was removed and no populated value was altered.** The only change to an existing field is the empty-to-filled case described above. The four dataset views keep the same columns in the same order.

### 2026-08-08

*A run that crashes now says so.*

- **An unexpected internal error now ends the run as FAILED**, where it previously reported SUCCEEDED with an error recorded inside the run summary. A crash that reports success is invisible to alerts and to any integration reading the run status, which is the opposite of what a failure should be.
- **The run summary is still written first.** `OUTPUT` is attempted before the run fails, so a failed run is still explainable: it carries `stopReason: "error"` and the original message. If writing the summary also fails, that is logged as a secondary problem and never replaces the original error.
- **Nothing else changes status.** A stop on the time budget, a stop on the charge limit, invalid input, and a run whose only output was error rows all still finish SUCCEEDED. Those are outcomes you asked for or can act on, not infrastructure failures.

### 2026-08-07

*Three missing brands, and three descriptions that did not match the product.*

- **Evolute, HWA and UMO are selectable in the Brand dropdown.** The source catalogue lists 392 makes, the dropdown offered 389. No entry was removed.
- **The dataset view formerly called "Errors" is now "Error columns"**, and says what it does: it projects the error columns across every row and does not filter, because an Apify dataset view can only choose fields. A row with an empty error cell is a normal successful row. The view id is unchanged, so saved links and `view=errors` API calls keep working.
- **The overview now presents auto-data.net as the backbone catalogue**, with the other three sources described as conditional enrichment, two of them opt-in. The previous wording led with four sources, which overstated a default run.
- **Removed a claim the Actor does not yet keep**: carperformancedata.com figures were described as carrying original-source attribution. They do not, today. The claim returns when the feature ships.

### 2026-08-06

*Runs that stop on their own terms, and a billing contract that matches reality.*

- **A run now stops itself before the platform timeout** instead of being killed by it. It keeps the records already produced, finishes as a successful run with a partial dataset, and says so. Previously a run that ran out of time was terminated mid-flight and a whole-brand run could return nothing at all.
- **Records are written as each model completes.** A brand run used to hold every row until the entire brand had been walked, so its dataset stayed empty for most of the run and then filled in one burst. It now fills progressively, and stopping early keeps every model already finished.
- **A run summary is written to the key-value store under `OUTPUT`** on every ending the Actor controls: normal completion, invalid input, time budget reached, spend limit reached, or an unexpected error. It reports how the run ended, how many records were delivered and billed, and the last model fully completed. It cannot be written when the platform kills the container outright.
- **The billed record is described as it really is.** The old wording promised each charged record was cross-checked across all enabled sources. In practice a record carries whatever the enabled sources actually matched, which for many models is auto-data.net alone. Error rows do not trigger the **Car data** event, and the standard **Actor Start** event is charged separately for every run.
- **`canonicalVariantId` is documented as a grouping key, not a unique one.** Variants that differ only in fuel, gearbox or door count share an id. For a unique key, use the auto-data source URL in `sourcesDetail`. The id format itself is unchanged in this release.
- **`excludeModels` is documented as an exact, case-insensitive match** on the model or make and model. It always behaved this way, the description said "contains".
- **Every run must carry at least one limit.** `maxItems: 0` (no record cap) is now rejected on a run that also has no timeout, since nothing would stop it. **Behaviour change for API callers:** `maxItems: null` now means the default cap of 1000, where it previously meant no cap.
- Progress is reported by a heartbeat scheduled every 15 seconds, which never overlaps itself, in both the run status and the log, so a long walk is visibly working.

### 2026-07-06

*Faster, cheaper default runs.*

- The default source set is now **auto-data.net + carperformancedata.com**, the fast and reliable core. ultimatespecs.com and zeperfs.com are opt-in via a new **Extra sources** field (or the API `sources` option): ultimatespecs.com is slower and nests performance variants under their base model.
- The default input example is now a single model (`Audi RS3`) with a low **Max items**, so a first run finishes in seconds for a few cents instead of walking a whole brand. Use the Brand mode and raise Max items for larger pulls.
- Consolidated rows are written to the dataset **incrementally** as they are produced, so results appear during the run and a long run is not lost if it is stopped early.
  - *Annotated 2026-08-07, scope correction:* this was true **per target**, and a whole brand was a single target, so a brand run still filled its dataset in one burst at the end. Per-model flushing arrived in the 2026-08-06 entry above. The original wording is kept as published.

### 2026-06-25

*Input UX and matching improvements.*

- Reworked the input around an **Input mode** menu: choose **Brand** (pick a make from a dropdown to scrape all its models) or **Model names list** (list specific models such as "Audi TT"). Choosing from menus instead of typing free text prevents typos.
- Improved cross-source matching so carperformancedata figures consolidate onto more variants.

### 2026-06-24

*Initial release (v0.1.0).*

- Consolidates car technical specifications and performance figures from four reference sources: auto-data.net (backbone catalog), ultimatespecs.com, carperformancedata.com and zeperfs.com.
- Resolves a brand or a list of models into all engine variants, then matches records across sources by fingerprint (make, model, displacement, power, production years) into one consolidated row per variant.
- Per-field source traceability: each value carries its source and alternative values, with a `hasConflict` flag and the maximum deviation when sources disagree beyond a built-in tolerance.
- Input filters: brand or model list, year range, exclude models, per-source toggles, and a `maxItems` safety cap.
- Pay-per-event billing: one charge per car data record emitted; error rows are returned for visibility and never charged.
