# Changelog of Kununu Scraper \[Just 💰$1] — Employer Reviews & Ratings (DACH) (`blackfalcondata/kununu-scraper`) Actor

- **URL**: https://apify.com/blackfalcondata/kununu-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/blackfalcondata/kununu-scraper.md

## Changelog

### 0.4.3 — 2026-09-29

#### Fixed

- Restored access after kununu changed its human-verification page on 2026-09-28. Since then every run ended with no results and a "temporarily inaccessible" note, whatever the input. Runs now pass the new check and deliver companies and reviews again.

### 0.4.2 — 2026-08-27

#### Fixed

- A run no longer ends on the first momentary connection fault. Retries were decided by matching the wording of the error, so any fault whose wording was not on the list skipped the retry budget entirely and failed the run in seconds. Connection faults are now recognised by what they are rather than by how they are worded, so the full retry budget applies to all of them. A bad URL or a redirect loop still fails immediately, as before.

### 0.4.1 — 2026-05-26

#### Fixed

- Improved resilience of detail-page fetches so transient anti-bot responses or proxy network hiccups are less likely to fail a run outright. Runs that previously failed after a short retry burst now have a higher chance of completing with results.

### 0.4.0 — 2026-05-13

Major data-coverage update — applicant interviews, awards, score trends, structured addresses:

- Added: `includeApplicantReviews` + `maxApplicantReviewPages` — paginate the `/bewerbung` feed for individual applicant/interview reviews. Each emits a `type: "review"` row with `reviewType: "application"`, `applicantYear`, `applicantResult` (e.g. `"hired"` / `"rejected"` / `"offerDeclined"`), and `interviewQuestions[]` (applicant-recalled question text). Default off; auto-enables `includeReviews` + `includeDetails` when set
- Added: `includeRatingBuckets` — 4-bucket star distribution (`excellent` 4-5★, `good` 3-4★, `satisfactory` 2-3★, `subpar` 1-2★) with per-bucket `percentage` + `totalReviews`, for both employee and applicant reviews. Adds 1-2 extra requests per company depending on whether `includeReviews` is also set
- Added: `includeRichContent` — opt-in `richContent` block carrying slogan, narrative HTML (whatIsSpecial, benefitsStatement, whoWeAre, whoWeAreLookingFor, legalInformation), image gallery, video gallery, and jobInsights Q\&A interviews with named employees. HTML emitted verbatim — sanitise on consumer side
- Added: `includeContactData` — toggle (default on) for `extractedEmails` and `extractedPhones`. Disable for stricter PII policies; company social profiles unaffected
- Added: Tier-A enrichment fields populated when `includeDetails: true` — `scoreTrend` (newer vs older 24-month averages with delta), `industryAverageScore`, `firstReviewYear`, `responseTimeInDays`, `simpleName`, `previewType` (`EBP_Pro` / `EBP` / `Claimed` / `Non_EBP`), `awards[]` with per-year badges + `wasPaid` flag, `companyFacts` (`numberOfEmployees`, `revenue`), `officialSocialMedia` (6 platforms), `locationSummary` (main city + state + additional locations count), `profileLocations` (structured addresses with lat/lng), `hasLocationsContent`, `isClaimed`, `isTopCompanyPaid`
- Added: ten review-volume splits — `totalReviewsEmployees`, `totalReviewsCandidates`, `totalReviewsLastTwoYears`, `totalReviewsWithResponseLastTwoYears`, `totalNegativeReviewsLastTwoYears`, `totalReviewsWithText`, `totalCultureReviews`, `totalEmployerReviewsWithText`, `totalApplicationReviewsWithText`, `totalApprenticeshipReviewsWithText`
- Changed: review pagination loop extracted to a pure helper (`src/reviewPaginator.ts`) with unit-tested dedup / bail-out / cap / unlimited / mixed-batch-failure / 404-mid-pagination behaviour. Explicit warning logged when kununu re-serves page 1 as page N
- Changed: incremental-mode content hash now detects same-length text edits (typo fixes, swapped interview answers, narrative HTML changes); the change hash covers every new enrichment field
- Changed: `includeApplicantReviews` / `includeRatingBuckets` / `includeRichContent` / `includeFullSalary` / `includeCulture` auto-enable `includeDetails` (and `includeReviews` where applicable) at normalize-input time with a log warning, so dependent flags can't silently no-op
- Note: discriminator strategy — applicant and employee reviews share `type: "review"`; subtype lives in `reviewType` (`"employer"` / `"application"`). Applicant-only fields (`applicantYear`, `applicantResult`, `interviewQuestions`) are `null` on employee reviews; `recommended` / `former` / `apprenticeshipJob` / `trainee` are `null` on applicant reviews
- Note: all new fields are additive — existing configurations continue to work unchanged. New output fields return `null` when the opt-in flag is off or the source has no data
- Note: existing incremental users will see a one-time UPDATED wave on first v0.4 run as the new enrichment fields enter the change hash

### 0.3.5 — 2026-04-28

Reliability + notifications:

- Fixed: improved reliability on accounts with limited proxy entitlements — the actor now degrades gracefully instead of failing the run when the preferred proxy group is unavailable
- Added: Telegram, Discord, Slack, WhatsApp, and generic webhook notifications. All disabled by default — fill the relevant input fields to enable

### 0.3.4 — 2026-04-17

README savings example now rendered for multi-event pricing:

- Changed: the "recurring monitoring savings" section in README now renders for actors with multi-event pricing. Numbers are computed against the primary event, with a disclosure note naming the secondary events that are billed separately
- No code or pricing change

### 0.3.3 — 2026-04-17

URL-mode internals consolidated into a shared workspace module:

- Changed: URL-mode orchestration, NOT\_FOUND cache, and status-summary helpers now live in a shared module (`urlModeInput`) reused across actors. No behavior change — same parse rules, same NOT\_FOUND key convention (`notfound_url_{cc}_{slug}`), same `setStatusMessage` summary format
- Note: first-run state is unaffected. Runs with a prior `stateKey` continue to classify per-slug exactly as before

### 0.3.2 — 2026-04-17

Incremental robustness:

- Fixed: a transient detail-page fetch failure (SSL handshake, proxy timeout) no longer flips a genuinely unchanged company to `UPDATED`. When the detail fetch fails and incremental mode has a prior hash for the company, the cached hash is preserved and the company is emitted as `UNCHANGED` (filtered from output). Next run retries the fetch
- No schema or behavior change for first-run / companies without cached state

### 0.3.1 — 2026-04-17

URL mode UX polish:

- Added: 30-day NOT\_FOUND cache for URL mode — repeated runs with dead URLs skip the fetch (KV key `notfound_url_{cc}_{slug}`), same pattern as company-name mode
- Changed: `queryCompanyName` on URL-mode output now preserves the exact input URL instead of the resolved company name, so downstream pipelines can join on the user's original input
- Changed: invalid URLs and empty-result runs set a human-readable status message via `Actor.setStatusMessage` instead of failing the run — the run stays SUCCEEDED with a summary like `✓ 2 companies scraped · ✗ 1 unrecognized URL skipped: https://...`

### 0.3.0 — 2026-04-17

Direct URL input mode:

- Added: `companyUrls` input — list of kununu company URLs (e.g. `https://www.kununu.com/de/bosch-gruppe`). Country code and slug are parsed from each URL, so URLs can mix DE/AT/CH in a single run
- Added: URL mode skips search and Jaccard matching — the detail page is fetched directly, which saves one request per company vs `companyNames` mode
- Added: `extractBasicProfile` — builds a `KununuProfile` (name, uuid, score, totalReviews, industry, location, isTopCompany) from the detail page Redux state, replacing the SERP response that URL mode lacks
- Changed: URL mode auto-enables `includeDetails` (the detail page is required to populate the output)
- Note: incremental mode works per-slug exactly as in `companyNames`/`datasetId` modes — first run = all NEW, subsequent runs filter UNCHANGED

### 0.2.5 — 2026-04-17

Reviews page 1 now prefetched alongside detail/salary/culture:

- Changed: `fetchReviewsPage(..., page: 1)` runs in the same Promise.all as detail/salary/culture when `includeReviews: true`. Saves ~2s per company on runs where the company is NEW or UPDATED. The prefetched page is discarded when the company is later classified UNCHANGED (rare in practice — aggregate hash already accounts for review activity)
- No behavior change on individual review emission, pagination, or incremental review tracking

### 0.2.4 — 2026-04-16

Per-company detail fetches now run in parallel:

- Changed: `includeDetails` + `includeFullSalary` + `includeCulture` fetches issue in parallel per company instead of sequentially. Expected ~2-3x speedup when all three flags are enabled
- No behavior change — same output fields, same error handling, same incremental semantics

### 0.2.3 — 2026-04-16

Kulturkompass + review responses:

- Added: `includeCulture` input — fetch the `/kultur` page for the Kulturkompass (profile-vs-industry compass score, MODERN/TRADITIONAL binary classification, 4 culture dimensions, strength/weakness/most-voted factors, company culture statements, and culture-tagged review comments)
- Added: `kulturKompass` output field on company profiles (null when a company has too few culture submissions)
- Added: `responses` field on review items — individual company replies with timestamps, author, and response body (previously only `responseCount` was exposed)
- Changed: review content hash now includes response count + latest response timestamp, so a company reply emits the review as `UPDATED` in incremental mode
- Note: `includeCulture` is opt-in; adds one extra request per company. Requires `includeDetails: true`
- Note: review response data is always included when `includeReviews: true` — no new flag, no extra fetch

### 0.2.2 — 2026-04-16

Full salary enrichment:

- Added: `includeFullSalary` input — fetch the `/gehalt` page for all salary ranges (~20 job titles with min/max/median/average/entry counts) instead of the top 3 shown on the detail summary
- Note: opt-in flag; adds one extra request per company. Requires `includeDetails: true`
- Note: same `SalaryRange` schema as before — downstream consumers need no changes

### 0.2.1 — 2026-04-16

Individual review scraping:

- Added: `includeReviews` input — scrape individual employee reviews per company (requires `includeDetails: true`)
- Added: `maxReviewPages` input — cap review pagination per company
- Added: `reviewSort` input — newest / oldest / relevance / best / worst
- Added: review dataset items with `type: "review"` discriminator, 13 factor ratings per review, pros/cons/suggestions, position/department, reactions, and full company reference
- Added: per-review incremental classification — each review tracked by uuid and content hash; unchanged reviews are filtered from emission
- Added: `review-extracted` PPE event at $0.001 per review. Primary `company-profile` event renamed from the generic dataset-item event so each row is charged exactly once
- Note: review scraping is skipped when a company is UNCHANGED in incremental mode — the aggregate hash covers review activity
- Note: in-session deduplication by review uuid, and early-exit when consecutive pages stop yielding fresh reviews, guard against wasted traffic
- Note: review pagination supports deep scraping (verified 90+ unique reviews at `maxReviewPages: 10`, capped per-company by `totalReviews`)

### 0.2.0 — 2026-04-16

Company profile enrichment — expanded output fields when `includeDetails: true`:

- Added: `scoreBreakdown` — 4 rating categories with 13 factors (salary, career, atmosphere, leadership, etc.)
- Added: `benefits` — list of benefit types with employee endorsement percentages
- Added: `salaryRanges` — top salary ranges per job title (min/max/median/average)
- Added: `competitors` — related companies with scores
- Added: `topCompanyYears` — years the company earned a Top Company badge
- Added: `salarySatisfaction` — aggregated salary satisfaction metrics
- Added: `followerCount` — number of kununu followers
- Added: `recommendationTotalReviews`, `recommendationRecommended`, `recommendationNotRecommended` — full recommendation breakdown
- Added: `type` discriminator field (`"company"`) — enables mixed-item datasets
- Changed: Detail extraction is significantly more complete and accurate than v0.1
- Note: existing incremental users with a `stateKey` will see a one-time UPDATED wave on first v0.2 run as enrichment fields are included in the change hash

### 0.1.x — 2026-04-14

- Added: `descriptionHtml`, `descriptionMarkdown` output fields (triple-format descriptions for RAG/LLM pipelines)
- Added: `contentHash` output field (SHA-256 hash of content-identifying fields)

### \[0.1.0] — 2026-04-14

Initial release.

- Three input modes: keyword search, company name list, dataset enrichment
- Jaccard similarity matching for company name lookup
- Optional detail page enrichment: recommendation rate and company website
- Incremental mode with KV store state tracking (NEW/UPDATED/UNCHANGED)
- NOT\_FOUND caching with 30-day TTL
- Industry filter for keyword search mode
- Compact output mode for AI-agent workflows
- Deduplication support for dataset input via configurable field
- Parallel SERP fetching (5 concurrent pages) and detail fetching (3 concurrent)
- Proxy support for reliable access
- PPE pricing: $0.005/start + $0.001/result
