# Changelog of Welcome to the Jungle Jobs Scraper — Europe Jobs & Salaries (`yasaslive/job-board-scraper`) Actor

- **URL**: https://apify.com/yasaslive/job-board-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/yasaslive/job-board-scraper.md

## Changelog

All notable changes to this actor are documented here. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and the actor uses `MAJOR.MINOR` versions as required by the Apify platform.

### \[1.0] - 2026-09-27

#### Added

- Search Welcome to the Jungle by keywords, location, site edition (fr/en/es/cz/sk), contract types, remote policy and publication age.
- Full job details: description (HTML + text), apply URL, salary, experience/education level, benefits, tools/skills tags, all office locations.
- Company data: name, slug, sector(s), size, founding year, logo, website and description.
- Start URLs: individual job pages and job-search pages (InstantSearch `query`, `aroundQuery`, `refinementList[...]` parameters).
- Pay-per-event pricing: `actor-start` ($0.01), `job-scraped` ($0.002), `job-with-description` ($0.001 add-on).
- Run statistics in the `STATS` key-value record (logged every 30 s), including categorised error counters and `budgetReached`.
- `SiteAdapter` interface (`src/sites/`) so other regional boards can be added as one adapter each.

#### Design decisions

- **Search via Algolia REST, HTTP only.** The site's own search runs on Algolia. The public, search-only app ID and key are read at run start from the site's runtime config (`/api/env`) instead of being hard-coded, and are re-read once if they rotate mid-run. The key is referer-restricted, so requests send the site's `Origin`/`Referer`, exactly as the browser does. `analytics: false` keeps scraper traffic out of the site's search analytics.
- **Details from the public JSON API, not `__NEXT_DATA__`.** The spec assumed a Next.js `__NEXT_DATA__` payload. The current site has none: job pages embed a React Query cache in `window.__INITIAL_DATA__` (a double-encoded JSON string). The primary detail source is the site's public, unauthenticated JSON endpoint `api.welcometothejungle.com/api/v1/organizations/{org}/jobs/{slug}`, which returns the same object. The HTML `__INITIAL_DATA__` parser is kept as a fallback. Both are unit-tested against saved fixtures.
- **Anti-bot challenges are detected, never solved.** Job pages sit behind AWS WAF, which after bursts answers `202` with `x-amzn-waf-action: challenge` and a JavaScript proof-of-work page. Those responses (and 200-status challenge pages) are classified as `blocked` and retried with a new session/IP and exponential backoff. If details still fail, the job is emitted with search data only (`detailsScraped: false`) and no description add-on is charged.
- **Got Scraping directly instead of CheerioCrawler.** All targets are JSON APIs, so no HTML crawling framework is needed. Crawlee's `SessionPool` and Apify's `ProxyConfiguration` still provide session and IP rotation. No browser (Playwright) is used.
- **`maxJobs` is the "max results" field.** The actor spec calls it `maxJobs` with default 100. It takes the place of the generic `maxResults` (default 50) convention.
- **`country` is the site edition, not a country filter.** Algolia exposes one index per UI language (`wttj_jobs_production_{fr,en,es,cs,sk}`) over the same jobs. Country filtering is done through **Location**. The `cz` input maps to the site's `cs` locale.
- **Location resolution without geocoding.** The free-text location is matched (accent- and case-insensitively) against the city/district/state/country facet values of the current search, plus country names in six languages via ICU. No third-party geocoding API or key is used. An unmatched location logs suggestions and skips that query instead of silently returning unfiltered results.
- **>1,000 results.** Algolia caps a query at 1,000 reachable hits. Beyond that, the search is recursively split into publication-date slices of at most 1,000 hits each (newest first).
- **Billing safety.** Charges happen only after `Actor.pushData` succeeds. Affordability is checked up front for both the base event and the add-on (serialised by a mutex so concurrent workers can't overshoot). If only the base event fits, the description is stripped instead of emitted uncharged. `actor-start` is guarded by a key-value flag so platform migrations don't double-charge.
- **Residential proxies by default.** The site's AWS WAF challenges datacenter IPs far more often, so `proxyConfiguration` defaults to `RESIDENTIAL`. If the account can't use residential proxies, the run falls back to the default Apify Proxy with a warning instead of failing (`src/utils/proxy.ts`). Payloads are small JSON (~50 KB per job), so residential bandwidth costs stay low (~50 MB per 1,000 jobs).
- **Deduplication** by stable job ID (Algolia `reference`), persisted with `Actor.useState` so a migrated run never re-emits or re-charges a job.
- **Primary location from a single office.** Multi-office jobs take their primary city/country/coordinates from one office (the job's main office), and list every office in `location.allLocations`.

#### Known limitations

- `npm audit` reports a moderate advisory in `stream-json` (GHSA-528h-pc64-c93x), a transitive dependency of `@crawlee/core`. Crawlee only uses it (via `StreamArray`, not the affected pick/ignore/filter/replace filters) to deserialize its own `RequestList` state, which this actor does not use. It is accepted until an upstream Crawlee release bumps it.
- Local pay-per-event simulation in the Apify SDK prices every event at $1 off-platform, so exact budget behaviour is verified by unit tests with a platform-faithful fake instead.
