# Changelog of Jobindex Scraper — Denmark Job Listings (`blackfalcondata/jobindex-scraper`) Actor

- **URL**: https://apify.com/blackfalcondata/jobindex-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/blackfalcondata/jobindex-scraper.md

## Changelog

### 0.2 — 2026-10-01 (cleaner descriptions, steadier incremental runs, clearer input errors)

#### Fixed

- `description`, `summary` and `descriptionMarkdown` no longer start with the search-page menu text ("Vis interesse, Ikke interesseret, Fejlmeld annonce, Del annoncen, Kopier link, Via apps, Facebook, LinkedIn", the headline and "Se rejsetid"). They begin with the ad text, including on English-language ads whose location is written differently on the search page.
- `deadline` is now only the ad's own application deadline. It used to fall back to the listing's last date, which is not a deadline; it is `null` when the ad has none.
- A location or URL the site does not recognise now ends the run with "A location or URL was not recognised. Check it and try again." instead of a generic "try again later". In a run with several searches, valid searches still deliver and the status says one or more searches were skipped as invalid.
- Salary detection is stricter: an amount in kroner in the ad text now only counts as salary when a pay word (løn, timeløn, honorar, tjen ...) is nearby or it has a period ("pr. måned", "i timen", "/t"), and not when it follows kontingent, bonus, rabat, tillæg, pension and similar. Ads that only mention such amounts no longer get `salaryMin`/`salaryMax`.
- Salaries written with a decimal comma were read wrongly: "162,34 kr. i timen" gave `salaryMin` 34. It now gives 162.34. "220kr.230kr./T", "190kr/t" and "kr pr. måned" are read correctly too.
- Incremental runs with job details switched on: when a known job's detail page could not be read, the job is no longer reported as UPDATED (and billed), and not a second time when the page comes back. Changing `descriptionMaxLength` no longer marks jobs as UPDATED.

#### Changed

- `employmentType` is filled when a job's detail page names one (the site's pages rarely do; it stays `null` otherwise).
- First incremental run after this update: stored jobs are re-baselined silently. Nothing is reported as UPDATED because of the changes above; only real changes after that are.

### 0.2 — 2026-09-30 (search coverage and working-hours fields)

#### Fixed

- A search page that came back empty mid-run is now treated as a missed page instead of the end of results. Before, incremental runs could mark live jobs on that page as expired, and other runs were silently cut short. Nothing changes for searches that genuinely have no results.
- `weeklyHoursMin`/`weeklyHoursMax` are only filled when the ad states hours per week (for example "20 hours per week" or "30 timer om ugen"), and `shiftWork` only when the ad names shift or evening/night/weekend work. Ads that merely mention "within 24 hours" or an "evening event" no longer get these values.

### 0.2 — 2026-09-30

#### Fixed

- Runs with a maximum cost limit now use the full budget. With a limit set, a run stopped early: a $0.03 limit returned 5 jobs where it pays for 10. Incremental runs now leave the jobs a limit did not cover for the next run instead of skipping them. Prices are unchanged.

### 0.2.3 — 2026-08-16

#### Fixed

- A dropped connection no longer fails the whole search. Only server error codes were retried; a request that failed before it got an answer was given up on immediately, so one network blip could end a run with "no results for any search". Those failures now go through the same retry with backoff.

### 0.2.2 — 2026-05-11

#### Added

- `country` and `countryCode` output fields for Denmark-market normalization.
- `summary` output field for compact previews and table views.
- `searchPage`, `searchEmploymentType`, `searchWorkPlace`, and `searchWorkHours` output fields for search provenance.
- Working-hours extraction fields: `workingHoursText`, `workingHoursType`, `weeklyHoursMin`, `weeklyHoursMax`, and `shiftWork`.

#### Changed

- `applyUrl` now falls back to conservative external application-like links found in listing/detail HTML when Jobindex's structured apply fields are empty.
- Compact output now includes `summary`, country fields, and `searchPage`.

### 0.2.1 — 2026-04-25

#### Added — Output fields

- `extractedEmails: string[]` — regex-extracted emails from description+requirements
- `extractedPhones: string[]` — defensive phone-number extraction (strict/lenient mode)
- `extractedUrls: string[]` — URLs in description (excluding self-referential jobindex.dk + it-jobbank.dk + ofir.dk)
- `socialProfiles: { linkedin, twitter, instagram, facebook, youtube, tiktok, github, xing }`
- `postedAt` — alias of `postedDate` for cross-actor consistency
- Incremental lifecycle: `firstSeenAt`, `lastSeenAt`, `previousSeenAt`, `expiredAt`
- Repost detection: `isRepost`, `repostOfId`, `repostDetectedAt`

#### Added — Inputs

- `emitUnchanged`, `emitExpired`, `skipReposts` — incremental emission policy
- 5 notification platforms: Telegram/Discord/Slack/WhatsApp Cloud API/generic webhook
- `webhookUrl`, `webhookHeaders` — n8n/Make/Zapier integration
- `notificationLimit`, `notifyOnlyChanges`, `phoneExtractionMode`

#### Changed

- **Full incremental classification** — `changeType` now emits `NEW`/`UPDATED`/`UNCHANGED`/`EXPIRED`/`REAPPEARED` (uppercase). Previously only `new`/`updated`.
- **State-lock** — concurrent runs sharing `stateKey` refuse with `Actor.fail`.

#### Compliance

- Migrated from hand-rolled `contentHash.ts` + `Record<tid, hash>` map to canonical `_lib/incrementalState.ts` + `_lib/stateLock.ts` + `_lib/notifications.ts`.
- Preserved jobindex's local `salaryParser.ts` (`parseSalaryFromText`/`EMPTY_SALARY`/`SalaryData` exports specific to Jobindex's JSON-LD format) and `htmlToMarkdown.ts`.

### 0.2.0 — 2026-04-22

#### Added

- `queries` input — run multiple keyword searches in one actor run with cross-query deduplication
- `startUrls` input — scrape any Jobindex or IT-Jobbank search URL directly; portal, query, location, and filters parsed automatically from the URL
- `radiusKm` input — limit results to a km radius around the specified location
- `fetchJobDetail` input — fetch each job's detail page to extract structured salary data and full description
- Salary output fields: `salaryText`, `salaryMin`, `salaryMax`, `salaryPeriod`, `salaryCurrency` — parsed from Schema.org data on detail pages, with fallback extraction from search-result snippets (handles Danish thousands separator: 45.000 kr. = 45000 DKK)
- `quickApplyAvailable` output field
- `searchQuery` and `searchUrl` output fields — which query/URL produced each result
- `descriptionMarkdown` now contains real Markdown (headings, lists, bold, links) when `fetchJobDetail: true`; plain text fallback otherwise
- `descriptionHtml` now contains the full detail page HTML when `fetchJobDetail: true`

#### Changed

- Input validation now accepts `queries` or `startUrls` as the primary input — `query` is no longer required when either is provided
- Incremental state key is backward-compatible for single-query runs; multi-task runs use a deterministic scope key

#### Fixed

- fetchWithRetry infinite recursion (0.1.14 regression) that caused `RangeError: Maximum call stack size exceeded`

### 0.1.0 — 2026-03-22

#### Added

- Initial release
- Keyword search across Jobindex.dk and IT-Jobbank.dk
- Company profile enrichment: rating, followers, social media, career page, about text
- Location with GPS coordinates
- Filters: employment type, workplace arrangement (onsite/remote/hybrid)
- Compact output mode for AI-agent and MCP workflows
- Incremental mode with change detection (`new` / `updated` / `unchanged`)
- `contentHash` for content change detection
- Retry on transient 5xx/gateway errors (502–504, 520–524, 590)
