# Changelog of Totaljobs Scraper 💰 $1.39/1K — UK’s Largest Job Board (`blackfalcondata/totaljobs-scraper`) Actor

- **URL**: https://apify.com/blackfalcondata/totaljobs-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/blackfalcondata/totaljobs-scraper.md

## Changelog

### 1.3.71 — 2026-10-04

#### Improved

- The buyer's maximum cost per run is honoured: the run collects only the rows the limit pays for and ends SUCCEEDED instead of ABORTED. In incremental mode, rows held back by the limit are not saved as seen and are delivered next run.

#### Fixed

- SERP: totaljobs A/B arm `herdwick` omits harmonisedId; retry such pages on a fresh session (bounded by maxRequestRetries, circuit-breaker), backfill only from ids the site sent for the same jobKey in-run, diag-emit residual nulls.

### Unreleased — 2026-09-30

#### Improved

- Runs that need a browser load pages faster and use less network: images, fonts, videos and third-party trackers are no longer downloaded. The job data returned is unchanged.
- When some pages need a second attempt, only those pages are retried the slower way; the rest of the run continues at normal speed. Retry counts and result limits are unchanged.

### 1.3.69 — 2026-09-27

#### Fixed

- Sorting by date or salary now applies to every results page. Before, only the first 25 results followed the chosen order and later pages came back in relevance order, so date-sorted runs could miss newer jobs.

### 1.3.68 — 2026-09-27

#### Fixed

- Sorting by salary now works on the UK and Irish job sites (high to low, or low to high). Before, it silently returned relevance order.
- On StepStone Germany, Austria, Belgium, the Netherlands and Pnet: sorting by salary is not offered by this job site. Choosing it now says so in the run log and sorts by relevance, instead of doing so silently.

### 1.3.4 — 2026-09-14

#### Fixed

- Fixed: a job carried over to the new identity is no longer reported as expired on the next full run. Re-keying a listing is not the same as the listing disappearing, and only the latter should ever be reported.
- Improved: carrying existing change-tracking data over to the new identity no longer depends on the source sending the identifier it had dropped. Jobs are now also recognised by their content, so nothing saved before this update is re-reported or re-charged.
- Fixed: in change-tracking mode the same job could be reported as new again on a later run, and charged again, even though nothing about it had changed. The source does not always send the internal identifier the tracking was keyed on, so a job's identity shifted between runs. Tracking now uses the job's own stable listing id. Existing tracking data carries over — the first run after this update recognises jobs saved under the old identifier, so nothing is re-reported or re-charged, and nothing is wrongly marked as expired.
- Searches whose keyword contains a full stop returned nothing and reported the site as blocked. "Vue.js developer", "node.js developer" and ".NET developer" were all affected: the full stop made the search address invalid, and the resulting error page looked like a block. The full stop is now treated as a word break, so these searches return results. Other punctuation (C++, C#, R\&D) was never affected and is unchanged.

### 1.3.3 — 2026-08-15

#### Fixed

- More unsupported filter values are now corrected instead of hanging or being silently ignored. Totaljobs no longer honours the contract-type, working-hours, remote-work, company or application-method filters — some are ignored, one returns an error page and one matches nothing at all — so those are skipped with a note in the run log rather than distorting the result. A salary filter is now sent with the yearly-amount qualifier it needs to take effect.
- A request that never comes back is no longer reported as a block. Timeouts and an actively refused connection were treated as the same thing, so a search that simply asked for a value the site does not answer was filed as "blocked". Runs still fail in both cases; only the reported reason changes.

### 1.3.2 — 2026-08-13

#### Fixed

- Searches that used an unsupported distance or date-filter value no longer end with no results. Totaljobs supports a fixed set of values for those two filters; a value outside that set is now adjusted to the closest supported one — never narrower than requested — and the run continues, with a note in the run log saying what was adjusted.
- The "Posted Within" filter now actually applies. It was previously accepted but had no effect, so results came back unfiltered while the run reported the filter as applied.
- The "Experience Level" filter is no longer sent, because Totaljobs no longer offers it. The run continues with a note in the log instead of failing, and the input is kept so existing task configurations keep working.

### 2026-06-16

- Added schema drift monitoring for incremental runs, with source-field baselines persisted in the configured state store.

### 2026-05-14

- Fixed: `maxPages` no longer silently caps `maxResults`. When you request more results than the default page budget allows, the actor now extends pagination automatically.

### 1.3.1 — 2026-05-04

#### Critical billing — caught by external review of v1.3.0 before ship

- **Search-results-only fast path no longer double-bills.** With `includeDetails: false`, the fast path pushed every row directly, but its usage flag evaluated to false because nothing was queued for detail enrichment. The standard path then ran the same start URLs again, producing duplicate dataset rows and charging twice. The flag now derives from "did the fast path successfully parse a search-results page", and the search-results-only push routes through the central `pushOutputForJob` helper. v1.3.0 shipped this regression but was held before going live; production stayed on v1.2.x.
- **Push contract is now wired into every engine path** (browser search results plus detail, alternate detail, browser-fetch detail, detail crawler, plus the `onDetailFailed` search-results fallback push in `main.ts`). v1.3.0 only reached routes.ts and the fast path, leaving the rest with the same C2/I1/I2 bugs the helper was meant to retire:
  - `markSeen`-on-dedup-hit: local `maybePush` returned `true` on dedup, callers then unconditionally called `incremental.markSeen` for rows that never reached the dataset, poisoning incremental state.
  - `detailsFetched: true` for parser failures: a 200 with no JSON-LD JobPosting now correctly emits `detailsFetched: false`. Was unconditionally `true` whenever the HTTP response succeeded.
  - Search-results-only rows missing `changeType` / `firstSeenAt` / `lastSeenAt`: the engine paths advertised these lifecycle fields but only one path (post-v1.3.0) populated them.

#### Diagnostics

- **`writeFailureDiagnostics` now fires on every early-exit branch.** Invalid `maxResults`, invalid `startUrl`, and state-lock conflict each write a `FAILURE_DIAGNOSTICS` KV record so consumers can distinguish "ran clean with 0 results" from "never started". Lock-conflict happens before the outer try/catch wrap, so it gets its own write rather than relying on the `thrown_error` fallback.

#### State store

- **Incremental state is pruned on save.** Two-phase: age-based prune drops entries whose `lastSeenAt` is older than 90 days, then a soft size budget (8 MB, leaving headroom under Apify's 9 MB per-record limit) triggers oldest-first pruning if the JSON payload still exceeds it. Without this, the v2 timestamp-per-id format would eventually breach the KV limit and start failing every save on long-lived state stores. `INCREMENTAL_RETENTION_MS` and `INCREMENTAL_MAX_BYTES` are exported so other actors can tune.

### 1.3.0 — 2026-05-04

#### Critical correctness — push contract

- **`pushOutputForJob` helper centralises every dataset write.** Owns changeType, firstSeenAt, lastSeenAt, transform, push, dedup, and `incremental.markSeen`. Callers now go through one function instead of 7 ad-hoc copies, so future fixes don't have to be replicated across engine paths.
- **`markSeen` is now strictly post-push.** Previously `routes.ts` search-results-only and LABEL_DETAIL paths called `incremental.markSeen` BEFORE `attemptPush`, so a cap rejection or transient push failure left the job permanently locked out of future incremental runs. Helper invariant now enforces "markSeen iff push succeeded".
- **Cap-rejected jobs are no longer marked seen.** The search-results `onDetailJob` callback used to mark cap-overflow jobs as seen ("// cap reached — mark remaining jobs seen"), causing silent data loss across runs. Removed; future runs now rediscover those jobs as NEW.

#### State model

- **`IncrementalState` v2 KV format.** Per-id timestamps `{firstSeenAt, lastSeenAt}` replace the bare seen-set. Legacy v1 `{ids}` state migrates automatically on first load by synthesising timestamps from the run-level `updatedAt`. `OutputItem.firstSeenAt` / `lastSeenAt` now carry real cross-run semantics.
- **State lock TTL bumped 30 min → 90 min.** With actor `timeoutSecs: 1800` (30 min), the previous TTL was identical to the timeout, so any long-but-legal run could trigger a `stale_override` race mid-write.
- **Lock release is compare-and-delete.** Previously `setValue(lockKey, null)` ran unconditionally, so a late-exiting run could nuke a concurrent run's freshly-acquired lock. Now reads current lock first and only clears if `runId` matches.

#### Detail engine bugs from the v1.2.0 audit pass

The "comprehensive correctness pass" of v1.1.0 missed both detail blocks (primary + deferred retry). All three engine-specific bugs are now fixed in those paths:

- **Currency: `parseDetailJsonLd` now receives `defaultCurrencyForGeo(geo)`.** UK was getting EUR salaries; now correctly GBP.
- **`contentHash` is positional null-safe** — was using `.filter(Boolean).join('|')` which dropped empty strings and shifted slot positions, so a missing field would re-hash an unrelated value.
- **`detailsFetched = Boolean(detail)`** — was unconditionally `true` when the detail response succeeded, claiming detail success even when the parser couldn't read the JSON-LD.

#### Cap math

- **Search-results collection at `main.ts:916` now compares `pushed + reserved + pendingDetailJobs.length`** against `maxResults` instead of just `pendingDetailJobs.length`. Same I7 bug from the v1.1.0 audit, different code path that had been missed.

#### Input contract

- **`geo` runtime default: `'TOTALJOBS'`.** API/CLI callers omitting `geo` now match the "Fixed to Totaljobs (UK)" promise.
- **`includeDetails` runtime default: `true`.** Console UI users and API/CLI callers now get the same default detail-enriched output.
- **Default drift audit added** — guards against future mismatches between Console defaults and API/CLI defaults.

#### Diagnostics

- **Optimised search-results failures now record block signals.** A 403/429/transport block on the optimised search-results path used to surface as "no signals + 0 pushed" in run summary, indistinguishable from an empty query.
- **Failure diagnostics on early exit / thrown errors.** Validation failures (missing query, invalid geo) and unhandled exceptions now write a `FAILURE_DIAGNOSTICS` KV record so consumers can tell a clean run from one that never started.

### 1.2.0 — 2026-05-04

#### Performance

- **Inline detail fetches now run with bounded concurrency**. Each search-results page used to fetch its 25 details one-at-a-time; now bounded by `min(maxConcurrency, 8)` per search-results page. End-to-end inline-detail throughput up ~4-8× on default settings.
- **Alternate detail crawler reuses browser contexts** across jobs (up to 10 uses per context, then auto-recycled). Previously created and immediately closed a fresh context per job (and per retry), which OOM'd on 500+-job runs and added ~1-2s of overhead per fetch. Retries still get a fresh context after a failed identity.
- **External retry backoff now jittered** (±20% uniform). Multiple actors retrying at the same tick produced thundering-herd spikes; jitter spreads the second wave.

#### Schema

- **Output view drift fixed**: `descriptionHtml`, `descriptionMarkdown`, `contentHash`, `changeType` were missing from the "all" export view. Display labels added for 11 lifecycle/repost/extraction fields that previously appeared with raw key names. A drift-audit test guards against future regressions.

### 1.1.0 — 2026-05-03

#### Critical fixes

- **Pagination cap removed**: `maxPages` was previously clamped to 1 in the browser-fetch path, so users requesting `maxPages=10` got only the first page of results. The clamp was a leftover speculative optimization; the search-results path now runs pagination end-to-end. (\[reported by @cleme_ntino])
- **Pass-2 escalation no longer double-bills**: `pendingDetailJobs` is now cleared after each detail phase. Previously, escalating after a failed pass re-pushed the entire pass-1 set against the cap, billing users twice for the same listings.
- **Detail retry no longer false-fails**: Firefox detail crawler used to retry whenever description was empty, even when the JSON-LD JobPosting block was present and complete. Retries now only fire when JSON-LD is entirely missing.
- **Incremental state isolation**: Two runs with identical `query+geo+location` but different `age`/`radius`/`contractType`/etc. used to share state, silently suppressing fresh hits in run B. Filter dimensions are now hashed into the state-key prefix.
- **Phone & URL extraction now read post-format text** (no longer broken by HTML→Markdown conversion); **email extraction reads raw HTML** (so `mailto:` anchors aren't lost). The previous code did the opposite of both.
- **`changeType: 'NEW'` now wired across all detail engines**. Was missing on multiple paths, so incremental subscribers couldn't tell new from existing items.
- **`contentHash` is now null-safe**: previously a missing field would throw inside the SHA-256 hashing call.
- **Lock-acquisition errors no longer mask root cause**: `Actor.fail()` was being thrown awkwardly during state-lock acquisition, swallowing the underlying error message. Now throws a plain `Error`.
- **State lock always released on failure**: try/catch added around the main run body so a crash mid-run still releases the lock instead of holding it for the full TTL.

#### Important fixes

- **Currency mapping is now geo-aware**: UK GBP, EU EUR, ZA ZAR. Salaries from JSON-LD without explicit currency previously defaulted to EUR for everything.
- **External detail success criterion fixed**: changed from `html.length > 5000` to JSON-LD presence check. Long block pages used to count as success; legitimate compact templates used to count as failure.
- **Telegram/WhatsApp message splits at semantic boundaries**: notifications now split at `\n\n` boundaries before falling back to hard slices, preventing job entries from being chopped mid-sentence.
- **Notification dispatch gated on success**: previously dispatched even when the run had failed mid-way.
- **Detail uniqueKey discriminated by pass**: pass-2 detail retries used to be deduped against pass-1 entries by the request queue, making escalation a no-op. UniqueKey now includes the pass label.
- **`startUrls` hostname validation**: invalid hostnames are rejected up front instead of failing mid-run with a confusing error.
- **`onDetailJob` cap math**: pass-2 escalation now correctly accounts for `pushed + reserved + pendingDetailJobs.length` against `maxResults`.
- **`stateStoreName` default**: now consistently defaults to `"totaljobs-state"` for all entry points.

#### Compact output

- `salaryMin` / `salaryMax` added to compact field set (essential for AI-agent salary filtering).

#### Operational

- Default memory bumped from 1024MB → 2048MB; default timeout 300s → 1800s. Browser detail paths previously OOM'd on larger runs and timed out on `maxPages>5`.

### 1.0.x — 2026-04-30

- Fixed: `startUrls` now processes all URLs in the array. Previously the optimized search-results path only used the first URL; subsequent URLs were silently ignored. Each URL is now its own pagination universe with shared dedup + maxResults cap.

### 0.1.x — 2026-04-14

- Added: `descriptionHtml`, `descriptionMarkdown` output fields (triple-format descriptions for RAG/LLM pipelines)
- Added: `contentHash` output field (SHA-256 hash of content-identifying fields)

### 0.1.x — 2026-04-14

- Added: cross-run repost detection (`isRepost`, `repostOfId`, `repostDetectedAt`)
- Added: `skipReposts` input to exclude detected reposts from output

### 1.0.0 — 2026-03-26

#### Added

- Initial release
- Search UK job listings on totaljobs.com by keyword, location, and filters
- Salary data, full descriptions, company profiles, and contact info
- Detail enrichment with apply URLs and employer metadata
- Incremental mode with change detection
- Compact output mode for AI-agent and MCP workflows
