# Changelog of UK Contract Expiry Radar - Recompete Leads (`datagrit/uk-contract-expiry-radar`) Actor

- **URL**: https://apify.com/datagrit/uk-contract-expiry-radar/changelog.md
- **Full Actor documentation**: https://apify.com/datagrit/uk-contract-expiry-radar.md

## Changelog

### 0.1

- Initial release: contracts ending in a chosen window from Contracts Finder award notices, with incumbent supplier, buyer, value, CPV, extension and call-off flags.
- Source read sequentially by day; proactive limiter of 12 requests per 121 s (server limit: 12 per 120 s, otherwise HTTP 429 with Retry-After 120); Retry-After on 429 is still honoured. State (count, emitted ids) persists across platform migrations.
- Default memory 256 MB. Prefill/default run: no keywords, 30-day lookback, 12-month window, 20 results.
- Manual checks: typical (30 days, no keywords, 20 results), edge (4 days, ends in 24-60 months, CPV 72/48, min value 50000, no call-offs, contact email on: 3 results, all constraints respected), no results (nonsense keyword: one unbilled status row), invalid input (expiresAfterMonths > expiresBeforeMonths: clear error, run fails).

### 0.2 (rework after review 4)

- Hard run-time budget: new input `maxRunSeconds` (default 240, 0 = unlimited). The limiter refuses to wait past the deadline and backoff sleeps are checked too; on the cut the results found so far are kept and the run status says how many lookback days were read. Daily Apify test (default input) can no longer exceed 5 minutes.
- Default and prefill `publishedLookbackDays` 30 -> 7. Measured on the default input with a keyword that matches nothing (full 7-day scan, 575 award records): 137 s including one 121 s limiter pause; with matching contracts the 20-result prefill run is much shorter (see 0.3 for the current measurement).
- Differentiator rewritten (no claim about competitors; expiry-window dataset for Contracts Finder only). seoTitle names Contracts Finder; README states the coverage (no Find a Tender) and drops the "under a minute" claim.
- Price 0.01 -> 0.005 USD per result (narrower coverage than both-portal competitors); cost re-measured (superseded by the 0.3 measurement below).
- Manual checks: typical (prefill, 20 results, 8 s), edge (keyword matching nothing, 7 days: full scan 137 s, one unbilled status row), budget (30-day window with `maxRunSeconds` 30: stops with a status message naming the days read, exit 0).
- Run status now reports source coverage of the key field: "Contract end date present in X of Y valid records" (partial drift of `contractPeriod` is visible; total loss still fails the run).

### 0.3 (rework after review 6)

- One row per award, also for corrected notices: records with the same award id are resolved per publication day (latest award date wins, on a tie the later one in the feed; a newer publication day beats an older one) before filters run. The run status now says how many duplicates were skipped and how many of them carried different data (live check: 2 duplicates, 1 with different data).
- Consequence: a whole publication day is read before its results are written, so the first results appear after ~2 requests instead of after the first page.
- Volume text corrected to what the source delivers: about 150 to 200 award notices per working day, about 2 requests, none at weekends (about 20 s per lookback day given the 12 requests / 120 s limit).
- Differentiator names the closest competitor (publicdata/uk-contracts-finder-find-a-tender: end-date filters, both portals, same price) and states the real differences; no price claim.
- Cost re-measured with `smoke --live --write-cost`: prefill input, 20 rows, 44 s, 0.0306 USD per 1000 results = 163x below the 0.005 USD price. Timing depends on the shared source rate limit.
- Manual checks: typical (prefill, 20 results), edge (2 days lookback, window 0 to 60 months, unlimited results: 113 contracts, conflict counter shown in the status), no results (fixture run: one unbilled status row).

### 0.4 (rework after review 7)

- Volume text re-measured day by day on the live source (2026-09-22 to 09-29): 125, 193, 158, 152, 147, 118 award notices per working day, 0 at weekends, always 2 requests per working day. README and the lookback input description now say "roughly 120 to 190" instead of "150 to 200".
- Build pinned: `package-lock.json` and `npm ci --omit=dev` in the Dockerfile; `package.json` version aligned with `actor.json`.
- No change in run behaviour.
- Cost re-measured with `smoke --live --write-cost`: prefill input, 20 rows, 6.1 s, 0.0042 USD per 1000 results = 1189x below the 0.005 USD price (time depends on the shared source rate limit; 44 s was measured in 0.3 when the limiter paused).

### 0.5 (rework after review 8)

- Duplicate awards are merged field by field instead of by position in the feed. Live scan of 7 days (594 records, 12 duplicate pairs): all 12 pairs have the same award date, so the "later award date wins" rule never decided anything; the feed position did, and it gave QBE021 an award value of 0 instead of 150,263.40 GBP. Now a missing value (null, empty, 0 award value, no supplier) loses to the other copy's value and daysToExpiry, monthsToExpiry and contractLengthMonths are recomputed from the merged dates.
- Real disagreements (award value 504,113 vs 50,411,351, contract end 2027-02-28 vs 2028-02-28, title and description typos, category) are no longer silent: new fields `recordDisputed`, `disputedFields` and `disputedOtherValues` on every row (false/null when there is no dispute). The kept copy is the one with the later award date, then the later one in the feed. Live scan after the change: 586 rows, 1 with a value filled in, 5 disputed, 15 rows with award value 0 remain, none of them has a duplicate copy that carries a value.
- README no longer presents the award-date rule as the working one; unit tests now include the 12 real duplicate pairs recorded from the live source (tests/fixtures/duplicates.json, contact points removed) next to synthetic cases.
- Run-time claims re-measured: full 7-day scan = 12 requests, 14.0 s and 15.2 s on two live runs, no limiter pause. README FAQ and the lookback field description now say that only windows beyond 12 requests wait (about 20 s per extra working day); the 137 s figure came from 0.2 and is removed.
- Added .dockerignore (tests, storage, node\_modules out of the image).
- Cost re-measured with `smoke --live --write-cost`: prefill input, 20 rows, 5.6 s, 0.0039 USD per 1000 results = 1281x below the 0.005 USD price. Manual checks: typical (prefill, 20 results, 5.6 s), edge (live 7-day full scan through the resolver: 598 records, 586 rows, 14.0 s, 12 requests, 5 disputed rows), no results (fixture run: one unbilled status row); offline smoke 33/33, live 23/23, unit tests 28/28.

### 0.6 (rework after review 11)

- `frameworkCallOff` now comes from the source's own procedure field (`procurementMethodDetails`: "Call-off from a framework agreement" or "Call-off from a dynamic purchasing system"), not from a regex over title and description. On a live sample of 417 awards (3 working days) the old regex flagged 266 and the procedure field says 246: 20 notices (4.8 percent) were stand-alone contracts that only mention call-off in their text (for example a restricted procedure for a framework of care services) and `includeFrameworkCallOffs=false` silently dropped them. After a merge of duplicate copies the flag is recomputed from the procedure of the kept copy.
- New field `callOffMentioned` for the text mention (title, description, award description); it never drives the filter. Dynamic purchasing system call-offs count as call-offs and this is stated in the field and input descriptions and in the README.
- `awardValue`: the source never sends null for the value, only 0 for "not disclosed" (9 of 417 awards, 2.2 percent, none null). The parser now emits null for 0, so the row, the dedupe merge and the `minValue` filter agree: with `minValue` above 0 an undisclosed value is skipped, with 0 it is kept. Input schema, dataset schema and README say so.
- Tests: 12 real releases recorded from the live source (tests/fixtures/call-off.json, contact points removed) cover both call-off procedures, text-only mentions under other procedures and a zero-value award; plus filter and duplicate-merge cases. Unit tests 35/35, offline smoke 35/35, live smoke 28/28.
- Cost re-measured with `smoke --live --write-cost`: prefill input, 20 rows, 5.0 s, 0.0035 USD per 1000 results = 1451x below the 0.005 USD price.
- Manual checks: typical (prefill, 20 results, live), edge (live sample of 417 awards through parser, resolver and filters, window 0-12 months: 210 rows; with call-offs off 82 rows, 10 of them mention call-off in the text and are kept; minValue 50000: 117 rows, none with a null value), no results (fixture run: one unbilled status row).

### 0.7 (rework after review 12)

- C16c: an empty feed (0 award records read, typical for a 1-day lookback on a weekend) no longer tells the client to widen working filters. The run status now says Contracts Finder published no award notices in the window, that no filter was applied, and to raise the lookback days; the "widen the expiry window, lookback days or keywords" advice is kept only for runs that read notices and matched none (`noResultsHint` in `src/filter.js`, unit-tested for both branches and the budget cut).
- C14: `AwardResolver.settled` kept the full resolved record of every award for the whole run just to count cross-day conflicts (1972 B per record, over 200 MB at the documented maximum of 1095 lookback days). It now keeps a 16-character hash of each version seen (`id -> hashes`): measured 211 B per record (35,000 records = 7 MB), about 22 MB at 105,000. Side effect fixed: the counter compared a finished row (with `recordDisputed`, recomputed days) against a fresh parse, so every cross-day copy of a disputed award counted as a conflict; it now compares parsed versions with parsed versions and counts only a version not seen before.
- E19: run-time numbers are ranges with conditions, not single measurements. Live feed, 21 days (2026-09-10 to 09-30): 97 to 193 award notices per working day (1 or 2 requests), 0 to 9 at weekends (not "none"), and all 15 sliding 7-day windows came to 11 or 12 requests and 653 to 797 award records (earlier runs 594 to 674). README FAQ and the lookback field now say: 11 to 12 requests, roughly 600 to 800 records, about 15 s with no pause; a day above 200 notices would add a 13th request that waits about two minutes, still inside the 240 s budget. New FAQ entry explains the two "no results" messages.
- Cost re-measured with `smoke --live --write-cost`: prefill input, 20 rows, 5.1 s, 0.0035 USD per 1000 results = 1414x below the 0.005 USD price.
- Manual checks: typical (prefill, 20 results, live, 5.1 s), edge (default 7-day live scan with a keyword matching nothing: 679 records, 17 s, 12 duplicates, 5 disputed, zero pushed rows besides the status row; 21-day request/volume probe above), no results (fixture run with `publishedLookbackDays: 1` on an empty feed: one unbilled status row and the new empty-feed message). Unit tests 39/39, offline smoke 35/35, live smoke 28/28.

### 0.8 (2026-10-01, rework after review 13)

- Pricing: a small fee when a run starts, then a price per result that depends on your Apify plan (lower on paid plans). Results are no longer free within each run; the Apify free plan includes monthly credit for trying the Actor. README and input schema no longer promise free results; `PRICING.json.freeItems` is 0.
- Differentiator and FAQ no longer compare prices with publicdata/uk-contracts-finder-find-a-tender; they state only the functional differences.

### 0.9 (2026-10-01, rework after review 14)

- Keywords are matched against the full award description; only the `description` field in the result stays shortened to 800 characters. Before, a keyword that appeared only after the 800th character silently dropped the contract (live feed, one publication day: 16 of the descriptions were longer than 800 characters and 5 keyword hits were lost). The full text is kept internally per award (also through duplicate merging) and is never written to the dataset. Test with a description of more than 800 characters in `tests/keyword.test.js`.
- Run status: "missing value filled in from the other copy" now counts only awards where a field was actually copied from the other copy; when the copy that was kept already had the value, nothing is counted. README describes both cases.
- Differentiator and README rewritten: the nearest Actor on the same source is dataio/uk-public-contract-awards (Contracts Finder only, end-date window in days), so the expiry view and computed days are not claimed as differences. What sets this Actor apart: call-off flag from the procurement procedure field plus a separate call-off wording flag, extension-clause flag, one row per award with corrected notices merged and conflicts marked, expiry window in months.
- Manual checks: typical (live prefill, 20 results, 5.0 s; cost re-measured 0.0035 USD per 1000 results = 1429x below the price), edge (keyword found only after the 800th character of the description: row kept, `description` shortened, `fullDescription` not in the dataset - `tests/keyword.test.js`), no results (fixture run: one unbilled status row). smoke 35/35 offline, 28/28 live, tests 41/41.

### 0.10 (2026-10-05, rework: source soft-throttle)

- Contracts Finder answers HTTP 200 with an empty `releases` list when one IP sends many requests in a row (measured 2026-10-02: the same query returned 100 notices and then 0). Before, an empty working day was read as "no notices published" and the run ended with 0 rows and a green status.
- A working day (Monday to Friday, UTC; today only after 12:00 UTC) that comes back empty is now checked against a control request for a day that returned notices earlier in the run (one notice, a single light request). If the control request is empty too, the source is throttling: the Actor waits 30 s, 60 s and 120 s, checking again each time, and re-reads the day as soon as the control request returns data. If the control request is fine, the day is a real bank holiday or quiet day and the run goes on without any wait.
- A day that stays empty after all waits is skipped and named in the final status. When nothing at all was read and at least one working day stayed unconfirmed, the run fails with a message that names the days and suggests a later run or a proxy, instead of returning 0 rows. Waits respect `maxRunSeconds`; a budget cut during a wait counts as an unconfirmed day.
- Tests: `tests/throttle.test.js` (bank holiday with a healthy control request: no wait; throttle that lifts after two waits; permanent throttle: pauses 30/60/120 s, then the day is unconfirmed; wait beyond the budget). 46 tests, smoke 28/28 live.
- Manual checks: typical (live prefill, 20 results, 6.4 s, cost re-measured 0.0045 USD per 1000 results), edge (4 inputs in smoke --live: 35, 2, 0 and 5 rows, no violations; one run met a genuine HTTP 429 with Retry-After 120 and finished in 126 s), no results (fixture run: one unbilled status row).

### 0.11 (2026-10-05, follow-ups after review 17)

- Run status counts disputed contracts in the result separately from disputed award records removed by your filters (before, it counted disputes before filters, so a user could look for `recordDisputed=true` rows that were not in the dataset). README updated.
- Control request for a throttle check now falls back to two days (7 and 14 days back) when no day has returned notices yet, so a bank holiday on the first control day no longer causes a needless 210 s wait.
- Input descriptions and README no longer give a fixed record count per week (it drifted every week); the proxy hint now says a proxy helps when the source throttles your IP.
- `PRICING.json.measuredCost.note` records the upper bound for a run that meets a rate-limit pause.
