Scrape job listings from Welcome to the Jungle (France, Spain, Czechia, Slovakia, UK) by keyword, location, contract and remote policy. Get salaries, full descriptions, benefits, apply links and company data as clean JSON.
All notable changes to this actor are documented here. The format follows Keep a Changelog and the actor uses MAJOR.MINOR versions as required by the Apify platform.
[1.0] - 2026-09-27
Added
Search Welcome to the Jungle by keywords, location, site edition (fr/en/es/cz/sk), contract types, remote policy and publication age.
Full job details: description (HTML + text), apply URL, salary, experience/education level, benefits, tools/skills tags, all office locations.
Company data: name, slug, sector(s), size, founding year, logo, website and description.
Run statistics in the STATS key-value record (logged every 30 s), including categorised error counters and budgetReached.
SiteAdapter interface (src/sites/) so other regional boards can be added as one adapter each.
Design decisions
Search via Algolia REST, HTTP only. The site's own search runs on Algolia. The public, search-only app ID and key are read at run start from the site's runtime config (/api/env) instead of being hard-coded, and are re-read once if they rotate mid-run. The key is referer-restricted, so requests send the site's Origin/Referer, exactly as the browser does. analytics: false keeps scraper traffic out of the site's search analytics.
Details from the public JSON API, not __NEXT_DATA__. The spec assumed a Next.js __NEXT_DATA__ payload. The current site has none: job pages embed a React Query cache in window.__INITIAL_DATA__ (a double-encoded JSON string). The primary detail source is the site's public, unauthenticated JSON endpoint api.welcometothejungle.com/api/v1/organizations/{org}/jobs/{slug}, which returns the same object. The HTML __INITIAL_DATA__ parser is kept as a fallback. Both are unit-tested against saved fixtures.
Anti-bot challenges are detected, never solved. Job pages sit behind AWS WAF, which after bursts answers 202 with x-amzn-waf-action: challenge and a JavaScript proof-of-work page. Those responses (and 200-status challenge pages) are classified as blocked and retried with a new session/IP and exponential backoff. If details still fail, the job is emitted with search data only (detailsScraped: false) and no description add-on is charged.
Got Scraping directly instead of CheerioCrawler. All targets are JSON APIs, so no HTML crawling framework is needed. Crawlee's SessionPool and Apify's ProxyConfiguration still provide session and IP rotation. No browser (Playwright) is used.
maxJobs is the "max results" field. The actor spec calls it maxJobs with default 100. It takes the place of the generic maxResults (default 50) convention.
country is the site edition, not a country filter. Algolia exposes one index per UI language (wttj_jobs_production_{fr,en,es,cs,sk}) over the same jobs. Country filtering is done through Location. The cz input maps to the site's cs locale.
Location resolution without geocoding. The free-text location is matched (accent- and case-insensitively) against the city/district/state/country facet values of the current search, plus country names in six languages via ICU. No third-party geocoding API or key is used. An unmatched location logs suggestions and skips that query instead of silently returning unfiltered results.
>1,000 results. Algolia caps a query at 1,000 reachable hits. Beyond that, the search is recursively split into publication-date slices of at most 1,000 hits each (newest first).
Billing safety. Charges happen only after Actor.pushData succeeds. Affordability is checked up front for both the base event and the add-on (serialised by a mutex so concurrent workers can't overshoot). If only the base event fits, the description is stripped instead of emitted uncharged. actor-start is guarded by a key-value flag so platform migrations don't double-charge.
Residential proxies by default. The site's AWS WAF challenges datacenter IPs far more often, so proxyConfiguration defaults to RESIDENTIAL. If the account can't use residential proxies, the run falls back to the default Apify Proxy with a warning instead of failing (src/utils/proxy.ts). Payloads are small JSON (~50 KB per job), so residential bandwidth costs stay low (~50 MB per 1,000 jobs).
Deduplication by stable job ID (Algolia reference), persisted with Actor.useState so a migrated run never re-emits or re-charges a job.
Primary location from a single office. Multi-office jobs take their primary city/country/coordinates from one office (the job's main office), and list every office in location.allLocations.
Known limitations
npm audit reports a moderate advisory in stream-json (GHSA-528h-pc64-c93x), a transitive dependency of @crawlee/core. Crawlee only uses it (via StreamArray, not the affected pick/ignore/filter/replace filters) to deserialize its own RequestList state, which this actor does not use. It is accepted until an upstream Crawlee release bumps it.
Local pay-per-event simulation in the Apify SDK prices every event at $1 off-platform, so exact budget behaviour is verified by unit tests with a platform-faithful fake instead.