# Changelog of Glassdoor Salaries Scraper (`devilscrapes/glassdoor-salaries-scraper`) Actor

- **URL**: https://apify.com/devilscrapes/glassdoor-salaries-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/devilscrapes/glassdoor-salaries-scraper.md

## Glassdoor Salaries Scraper — Changelog

### 0.2.0 — 2026-09-09

- Fix: 58% 30-day customer success (12/29 FAILED). Investigated our own
  runs first — all SUCCEEDED since the 0.1.0 proxy fix, and
  `scripts/os/run_usage.py` confirmed RESIDENTIAL is genuinely in use on
  the live build (0.1.2), so the earlier proxy fix is holding, not
  regressed. Most of the 12 FAILED customer runs are pre-fix residue
  (22/30 report-window days predate build 0.1.2).
- Root cause of the remaining/ongoing failures: `src/main.py`'s REQ-11
  zero-rows check (`if stats.parsed_rows_total == 0: raise
  SystemExit(1)`) could not tell a genuine "searched, no match" apart
  from a block. Live recon (Apify RESIDENTIAL proxy + firefox147, no
  Actor-run spend) proved Glassdoor always returns HTTP 200 with a
  structurally well-formed salary-search page — even for an invented job
  title — it just has no `jobTitle` text when it has no data. Any
  customer searching a title Glassdoor doesn't track paid the
  `actor-start` fee and got a FAILED run for a search that worked
  correctly (the exact fleet-wide anti-pattern in
  `ops/os/EMPTY-IS-NOT-A-FAILURE-2026-08-19.md` — this Actor was never on
  that doc's fixed list).
- Fix: `src/scraper.py` `RunStats` gains `searches_completed`, incremented
  only when a search/detail page was fetched AND definitively parsed
  (empty-list counts; `None`/unparseable does not). `src/main.py` extracts
  the decision into a pure `_should_fail_loud(*, parsed_rows_total,
  searches_completed)`: fails loud only when NO target ever completed a
  search (the true block case); zero rows with at least one completed
  search now succeeds with the existing zero-row status-message phrasing.
- New regression tests: `tests/test_rsc_parser.py`
  `test_search_page_with_no_job_title_is_genuine_empty_not_parse_failure`,
  `tests/test_scraper.py`
  `test_title_mode_genuine_no_match_marks_search_completed_zero_rows` +
  `test_title_at_company_no_entries_at_all_still_completed_search`,
  `tests/test_main.py` `test_should_fail_loud_when_nothing_ever_completed`
  - `test_should_not_fail_loud_on_genuine_no_match` +
    `test_should_not_fail_loud_when_rows_were_found` — all fakes/crafted
    fixtures, no monkeypatched parser internals for the new cases.
- Also found and flagged (NOT fixed this session — see
  `docs/specs/glassdoor-salaries-scraper/notes.md` 2026-09-09): the
  `title_at_company` search-mode's company-name-to-employerId resolution
  can never match on the real Glassdoor site, because the title-search
  page never exposes a per-company `employerId` list — confirmed by the
  code's own pre-existing docstrings/test comments. Every customer using
  that documented feature without a pre-known `employerId` will now
  succeed with 0 rows instead of FAILING, which is honest, but the
  feature itself still doesn't resolve companies by name. Needs a
  different Glassdoor endpoint, out of scope for this fix.

### 0.1.0 — 2026-08-25

- Fix: default `proxyConfiguration` now hard-sets
  `apifyProxyGroups: ["RESIDENTIAL"]` (models.py `DEFAULT_PROXY_CONFIGURATION`
  - `.actor/input_schema.json`). Root cause: 30-day customer success rate was
    25% (9/12 FAILED). Our own reproduction run (`fF7WGo9dVLCuO1ObG`,
    2026-08-24) hit a Cloudflare Managed-Challenge 403 on the first request,
    then timed out on every rotated retry — `scripts/os/run_usage.py` billing
    showed **zero** residential-proxy transfer, proving the unset-groups
    default resolves to the account's datacenter tier, not RESIDENTIAL as
    design.md had assumed. Mirrors the existing `bayut-uae-real-estate`
    mandatory-RESIDENTIAL pattern for the same Cloudflare-class target.
- Fix: `BROWSER_PROFILES` in `client.py` drops `chrome131`/`chrome124` and
  leads with `firefox147`/`firefox144` (Safari kept as secondary fallback).
  A live RESIDENTIAL-proxy probe against this Actor's own URLs on
  2026-08-25 (6 trials) found Chrome impersonation 403'd 0/6 while Firefox
  cleared 4/4 — same signal already documented for Reddit, now applied here.

### 0.0.1 — 2026-08-02

- Scaffolded skeleton (T01). Boots, reads raw input, pushes one
  placeholder dataset row. Real crawler pending
  `crawlee-developer`/implementer (`src/models.py`, `src/rsc_parser.py`,
  `src/client.py`, `src/scraper.py`, `src/main.py`).
