# Changelog of Glassdoor Company Reviews Scraper (`devilscrapes/glassdoor-reviews-scraper`) Actor

- **URL**: https://apify.com/devilscrapes/glassdoor-reviews-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/devilscrapes/glassdoor-reviews-scraper.md

### 0.9.0 — 2026-09-19

#### Fixed

- **`resolve_employer_id` was a permanent stub that always returned `None`.**
  `companyNames` is the Actor's only *required* input field; `employerIds` is
  documented as an optional "power-user shortcut". Every real customer who
  followed the primary, documented usage path — typing a company name with no
  `employerIds` — hit `_scrape_company_body`'s `if employer_id is None: return`
  guard, meaning EVERY company in EVERY such run yielded zero rows. `main.py`'s
  REQ-14 zero-rows check then failed the run loud with the status message
  `"0/N companies parsed — Reviews warm-up path may be blocked for this run's
  exit IP"` — a deterministic, 100%-reproducible code defect misdiagnosed by
  its own status message as a probabilistic anti-bot/IP problem. This is very
  likely the dominant contributor to the flagged 30-day customer rate (15/29
  FAILED, 48%): our own QA/dev runs always pass `employerIds` and so never
  exercised this path, which is exactly why `recent_failures` (own runs via
  `/v2/acts/{id}/runs`) showed zero failures on this build while the public
  customer-run stat did not — the two populations take different code paths.
- **Fix**: implemented real company-name resolution against Glassdoor's own
  `/Search/results.htm?keyword={company}` suggestion endpoint. Confirmed live
  2026-09-19 across 6 companies (Netflix→11891, Google→9079 — matches this
  repo's own long-standing QA fixture value, cross-validation not a
  coincidence, Meta→40772 — independently matches `glassdoor-salaries-
  scraper`'s own confirmed value, Deloitte→2763, Bank of America→8874,
  McDonald's→432) plus a nonsense query correctly returning no match. Unlike
  Overview/Reviews, this endpoint 200s on a **cold** navigation — no
  same-context warm-up required (4/4 live probes). New: `browser.py`'s
  `build_search_url`/`fetch_search_results`, `rsc_parser.py`'s
  `parse_company_search_suggestion` (anchored on the `directHit` key — unique
  per probed page, unlike `employerId` which the ratings object also carries).
  `scraper.py::resolve_employer_id` now navigates the passed-in `page` instead
  of receiving an unused `browser` handle, raises `AntiBotChallengeError`
  (RECOVERABLE\_BROWSER\_ERRORS — retried by `main.py`'s REQ-9 driver with a
  fresh proxy/session) when the search page itself is a genuine anti-bot
  interstitial, and returns `None` only for an ordinary no-match.
- Two new live-captured fixtures: `tests/fixtures/glassdoor_search_results.html`
  (Netflix direct hit) and `glassdoor_search_no_match.html` (nonsense query,
  genuine zero-match) — never hand-authored, both real captures per this
  repo's `reference-hand-authored-fixtures-encode-fiction` lesson.
- **Local verification**: `apify run` with `{"companyNames": ["Netflix"]}` (no
  `employerIds` — the realistic customer shape) now completes `exit_code: 0`
  with 3 review rows + 1 company-summary row, after correctly classifying and
  retrying one genuine challenge on the search-resolution step itself. The
  existing `employerIds`-supplied QA fixture (`Google`/`9079`) still passes
  unchanged — no regression on the previously-working path.
- **Known residual, not addressed this release**: a genuinely nonexistent/
  misspelled company name (a true no-match, not a block) still routes through
  `main.py`'s unconditional zero-rows-across-all-companies fail-loud path with
  the same misleading "may be blocked" message. This was already true before
  this fix (indistinguishable from the stub) and is a strict improvement, not
  a regression, but the "empty search should SUCCEED with zero rows" rule
  (`EMPTY-IS-NOT-A-FAILURE-2026-08-19.md`) should eventually be applied here
  too — needs `main.py`/`scraper.py` to propagate a "some companies resolved,
  some genuinely had no match" distinction that doesn't exist yet.

### 0.8.0 — 2026-09-14

#### Fixed

- **The ~10x under-delivery: every SUCCEEDED run capped out at exactly 3 review
  rows regardless of `maxReviewsPerCompany`.** Traced to production, not
  guesswork: 5 cloud runs (`pRgDfifO22exTlnwL`, `z4hA8U69gmT7dc4bu`,
  `4eE1YjYb7kF2uDvsK`, `5hPUeVWddEnvY8c4T`, `EqiMYtgtKQeADnHeH`), all on Apify
  RESIDENTIAL pinned to `US`, each hit exactly 3 rows for the shipped QA
  prefill (`Google`, `maxReviewsPerCompany: 30`). A local recon capture
  reproduced the same 3-row ceiling on a different proxy tier/country
  entirely (WebShare/NL), and it held steady across 60+ seconds of extra
  wait — ruling out a hydration-timing bug. The real cause: Glassdoor bakes
  `pageSize:3` into every anonymous session's review-list filters (confirmed
  in the live RSC payload's own `initFilters`), then blurs anything past
  those 3 behind a login wall (`BlurredOverlay`/`ReviewList_blurOverlayWrapper`
  in the live DOM). Numbered pagination (`_P2.htm`+) doesn't even try to
  blur — it redirects the *entire page* to `Log In | Glassdoor`, for every
  company, every proxy, every geography tested.
- `find_review_entries` was never the bug — it correctly parsed the only 3
  real review objects Glassdoor ever sent an anonymous session. The fix is a
  new pagination source: the Reviews page's own "Users say..." keyword-
  highlight links (`data-test="review-highlight-link"`, e.g.
  `/Reviews/Google-people-Reviews-EI_IE9079...htm`) are real, publicly-linked
  Reviews sub-pages that carry a materially larger anonymous `pageSize` per
  keyword (9 reviews sampled live for two different keywords) — this is now
  `scraper.py`'s actual second-page path (`_rows_from_highlight_pages`,
  `rsc_parser.find_highlight_review_urls`), tried before numbered pagination,
  which remains only as a fallback for companies too small to carry
  highlight links.
- Added `_is_login_wall`, checked ahead of `_is_challenge_page` in every
  pagination walker: retrying the login wall with a fresh proxy/session
  (main.py's REQ-9 driver) can never clear it, so it now stops cleanly
  instead of being mistaken for a retry-worthy anti-bot interstitial or,
  worse, silently falling through `if not entries: return` by accident.
- `_log_page_diagnostics` now runs for every subsequent/highlight page, not
  just page 1 — the pagination path used to be invisible past the first
  fetch.
- Fixed a real infinite-loop risk surfaced by this work's own tests: a page
  that returns non-empty `entries` which are *all* already in `seen` (zero
  new rows) now stops pagination like an empty page would, instead of
  looping forever waiting for a count that can never advance.

**Shipped**: `apify push --force` → version `0.8`, build `POLqouhTfacuMdOcF`
(buildNumber `0.8.1`), tagged `latest`. PPE pricing unchanged from the live
Store listing (per Gate 4b in the preceding dry-run) — no monetization
change this release.

**Post-push cloud verification** (mandatory `actor-qa-engineer` gate, both
runs against `build=latest`):

| run | input | status | rows | row\_type breakdown |
|---|---|---|---|---|
| `Tvt34pAai7zig2DQh` (shipped QA fixture, `maxReviewsPerCompany:3`) | fixture | SUCCEEDED | 4 | 3 review + 1 company\_summary |
| `Ai5TT62o4gvTQRgOp` (pagination-proof, `maxReviewsPerCompany:20`) | override | SUCCEEDED | 21 | 20 review + 1 company\_summary, all 20 `review_id`s unique |

The pagination-proof run's log shows the fix firing live in production —
`reviews-p1` alone wasn't enough to reach 20, so three distinct
`page=reviews-highlight` fetches followed (`"people"`, `"benefit"`,
`"culture"` keyword sub-pages), each `has_next_f_marker=True`, none a
duplicate — this is direct cloud proof the highlight-link pagination path
(new in this version) delivers past the old 3-row ceiling, not just a local
claim. `chargedEventCounts` on both runs matched dataset row counts exactly
(`review-row=20`/`company-summary-row=1` and `review-row=3`/
`company-summary-row=1`), no `ERROR`/`Traceback` in either log, both exited
clean (`exit_code: 0`). QA verdict: PASS.

### 0.7.0 — 2026-09-14

#### Fixed

- **The start fee is owed on delivery, not on boot.** `actor-start` was charged
  immediately after input validation — before proxy resolution and before a
  single page load — so every refused run billed the customer for an empty
  dataset. Glassdoor turns away roughly three attempts in four, which makes
  failing runs a permanent feature of this Actor rather than an anomaly, so
  this was not a rare edge: at the pre-0.6 success rate it meant ~16 of every
  29 customer runs paid for nothing. Live proof, both from 0.6.1 QA:
  `zVX1fjX2IRKBzGqLd` and `mZ9phI4xBBpE136Tl`, FAILED, each
  `{'actor-start': 1, 'review-row': 0, 'company-summary-row': 0}`.

  Billing a start fee for a run that delivered nothing is the shape that got
  `vrbo-vacation-rentals-scraper` delisted at 17 runs. `_StartFee` now fires
  at most once, and only from the first successful row push.

## Glassdoor Company Reviews Scraper — Changelog

### 0.6 — 2026-09-14

CEO daily report flagged UNDER\_MAINTENANCE: 44.8% 30-day customer success
(13/29), live build 0.5.1 (2026-09-09) — matching `main`, so no undeployed
fix was sitting on a branch; this needed a real one.

**Diagnosed from production, not guesswork.** `/v2/acts/{id}/runs` (own
runs) had nothing newer than 2026-09-09 — the 29 customer runs behind the
44.8% figure, including one as recent as 2026-09-13, are invisible to us
(they run under the customer's account, not ours). Read every own-run FAILED
log back to 2026-08-24 instead. All of them, including the two most recent
(`9hqNREpL769oeRmuK`, `gMdegWCzIK9NTKXvK`) — pre-0.5.1 builds — show the
per-company retry loop working exactly as designed: `AntiBotChallengeError`
raised, fresh proxy/session on each attempt, `COMPANY_RETRY_ATTEMPTS`
attempts exhausted, company skipped, run correctly fails loud only when
*every* company came back empty (REQ-14). No fault-isolation leak, no soft
200-as-block miss, no stale shared session — patterns 1-4 in the fixer brief
all check out clean; `_is_challenge_page`'s NEXT\_F\_MARKER/title checks (0.3,
0.5) already cover pattern 2 correctly.

The one own-run sample actually on the live build (`jAnCRbM93NJwRskb2`,
0.5.1, 2026-09-09 21:57) SUCCEEDED — the 0.5 hydration-race fix's target
condition (`rsc_chunks=7` on the very first content read) confirmed live.
That build is code-correct; the residual FAILED share is not a bug riding on
top of the wall, it's the wall itself, and `docs/specs/glassdoor-reviews-scraper/notes.md`'s
2026-09-03 measurement already named the exact number: Glassdoor is
probabilistically reachable at roughly **p≈0.25 per attempt**, and with
`COMPANY_RETRY_ATTEMPTS=2`, `P(at least one clears) = 1-(1-p)^2 ≈ 44%` — a
near-exact match for the live 44.8% that got this Actor flagged. That note
named raising the attempt budget as "the only change measured here that
would move delivery" but left it as a cost/benefit call for whoever owns the
trade-off. That's this session, and UNDER\_MAINTENANCE makes the call: ship
it.

**Fix**: `src/main.py`'s `COMPANY_RETRY_ATTEMPTS` raised `2 -> 4`
(predicted `P(pass) ≈ 68%`). Each retry already launches a fresh
Camoufox browser on a fresh proxy session (REQ-9, unchanged) — this only
widens how many times that already-correct retry gets to run before the
company is skipped. Nothing else touched: `AntiBotChallengeError` (0.1),
the challenge-title poll (0.3), the hydration-race poll (0.5), and REQ-14's
fail-loud-on-zero-rows semantics all stay exactly as documented. Pricing/
`pay_per_event.json` untouched — `actor-start` is still charged once per run
regardless of per-company attempt count, so the customer's cost per run does
not change; only our own compute cost on a company that ultimately fails
anyway goes up, traded for materially more companies succeeding and billing
`review-row`/`company-summary-row` events.

New regression test: `tests/test_main.py::test_retry_driver_uses_widened_attempt_budget`
— fails red against the pre-fix code (`COMPANY_RETRY_ATTEMPTS == 2`,
`open_browser` called twice), asserts both the constant and the actual
`open_browser` call count against an always-failing company, so a future
edit can't silently narrow the budget back.

**Verification**: 100 tests green (up from 99: 1 new). `ruff check` clean.
`ruff format --check` has pre-existing drift in 7 unrelated files (confirmed
via `git stash` diff against this branch's base — not introduced here);
`src/main.py`/`tests/test_main.py` carry the same pre-existing wrapping
drift on lines this change didn't touch. `pyright` 0 errors.

**Shipped**: `apify push` → version `0.6`, build `gkOPijKcEox0R1Eia`
(buildNumber `0.6.1`), tagged `latest`. `verify_ppe_sanity.py --live`
confirmed no live-vs-repo pricing drift.

**Post-push cloud verification** (mandatory QA gate + 4 additional probes,
per this session's brief — "two greens is not a rate; you need 4+" —
against `build=latest` with the unmodified shipped fixture):

| run | status | rows |
|---|---|---|
| `mZ9phI4xBBpE136Tl` (QA gate) | FAILED | 0 |
| `R6tRESMwI2ts4EqPQ` | SUCCEEDED | 4 (real pros/cons/ratings) |
| `BMFhvqSvoO0boMRL5` | SUCCEEDED | 4 |
| `HfJ8fj2XZ9dd7YSUy` | SUCCEEDED | 4 |
| `zVX1fjX2IRKBzGqLd` | FAILED | 0 |

4/5 = 80% — above the ~68% point prediction, well above the pre-fix ~44%
ceiling. Both FAILED runs show 4/4 attempts hitting an unambiguous
`title='Just a moment...'`/`rsc_chunks=0` challenge on every single try
(not the hydration-race shape) — a clean, correctly-classified, correctly
fail-loud genuine block, not a defect. Full run-by-run detail:
`docs/specs/glassdoor-reviews-scraper/notes.md`'s 2026-09-14 entry.

**Next step for whoever reviews this**: per notes.md's own caution, n=5
across this session (1 QA + 4 verification) is still directional, not
statistically sound (that needs n>=20/arm) — treat this as "the fix shipped
and measurably outperforms the pre-fix ceiling," not as final proof the true
rate is exactly 68%. If the 30-day customer number hasn't moved meaningfully
a week out, the next untried lever is `HYDRATION_POLL_ATTEMPTS`/
`CHALLENGE_POLL_ATTEMPTS` (the per-attempt success probability itself) or a
further budget raise (n=6, ~82% predicted) — both should be measured against
fresh evidence, not assumed.

### 0.5 — 2026-09-09

CEO daily report flagged 41% 30-day customer success (14/24 FAILED), 2 users,
live build 0.4.2 (4 days old). `/acts/{id}/runs` (own runs) had nothing newer
than 2026-09-03 — the 24 FAILED/SUCCEEDED customer runs were invisible to us,
so this session reproduced fresh: a free local recon through Apify RESIDENTIAL
(plain `curl`, no Camoufox) got the expected `403 Security | Glassdoor` — a
non-browser client always sees Cloudflare's wall regardless of this Actor's
health, so it only confirms the target is still defended, not whether our code
is broken. Two real cloud probes against live build `0.4.2` (Store-demo input,
`t152VidKGrQlNop2f` SUCCEEDED, `9hqNREpL769oeRmuK` FAILED; ~$0.02 total,
`run_usage.py`-checked RESIDENTIAL) found two distinct, previously
undiagnosed contributors on top of the already-documented probabilistic
anti-bot wall (`docs/specs/glassdoor-reviews-scraper/notes.md`, 2026-09-03):

1. **A false "blocked" classification, not a real block.** `9hqNREpL769oeRmuK`'s
   first attempt: `page=reviews-p1 ... has_next_f_marker=False
   title='Google Reviews (48,967)...'` — a REAL company title (never a
   `CHALLENGE_TITLE_MARKERS` match), yet `_is_challenge_page` scored it as a
   challenge anyway because the RSC hydration marker was absent. Root cause:
   Glassdoor's `<title>` renders before Next.js finishes streaming its RSC
   payload; `_wait_for_challenge_clear` polls only the title, so it returned
   the instant the title looked real, and the very next `page.content()`
   call captured a still-streaming ~18KB shell instead of the ~250KB+
   hydrated page. That burned a whole REQ-9 retry attempt (half of
   `COMPANY_RETRY_ATTEMPTS=2`) on a page that was never actually blocked —
   confirmed as a recurring pattern, not a one-off: the identical signature
   (~18-19KB, real title, zero RSC chunks) already appears in
   `PtHaK788kRd2QwwcN`'s log from the 2026-09-03 investigation, unnoticed at
   the time because that session's read attributed the whole failure class to
   the probabilistic Cloudflare wall.
   Fix: `src/browser.py` gains `_read_hydrated_target_content` — after the
   title-based challenge-clear poll returns, this polls `page.content()`
   itself (bounded, `HYDRATION_POLL_ATTEMPTS=4` / 1s apart) until
   `NEXT_F_MARKER` appears or the short budget is exhausted, then hands off
   to the existing classification unchanged. `NEXT_F_MARKER` moved from
   `scraper.py` into `browser.py` as the single source of truth (same pattern
   as `CHALLENGE_TITLE_MARKERS`'s 2026-09-03 move) so both the hydration wait
   and the challenge classifier agree on one signal.
2. **`ENABLE_GEOIP=True` (flipped on in 0.4, unproven) burned a retry attempt
   on an unrelated transient.** `t152VidKGrQlNop2f`'s successful run still
   logged `attempt=1/2 cause=Failed to get IP address: ...ipecho.net...
   SSLEOFError` before attempt 2 succeeded — Camoufox's own geoip IP-echo
   probe sweep (through the residential proxy, before Firefox even launches)
   failing on an unrelated third-party service, consuming one of only two
   attempts against Glassdoor's actual wall. The 2026-09-03 notes already
   judged geoip=True "not shown to help" at n=9 (aggregate 2/9 pass, same as
   geoip=False); this session found a concrete cost that measurement hadn't
   surfaced. Reverted `src/main.py`'s `ENABLE_GEOIP` to `False` (its own
   documented default before the 0.4 experiment) — an unproven flag that can
   measurably halve the effective retry budget doesn't stay on by a coin
   flip.

Both fixes are additive to the existing `AntiBotChallengeError` retry chain
(0.1) and the challenge-title poll (0.3) — neither is touched, both stay as
documented. This does **not** change the fundamental probabilistic nature of
Glassdoor's block (`docs/specs/glassdoor-reviews-scraper/notes.md`'s
attempt-budget math still applies); it removes two sources of *false*
failures riding on top of it.

Also cleared Gate 7 for publish: the pre-flight `/publish-actor` dry-run (this
session) found the 4 pyright errors carried since 0.3 in `test_models.py`/
`test_rsc_parser.py`/`test_scraper.py` now block publish outright (Gate 7 is
binary on zero errors, regardless of blame). Fixed all 3 sites — typing-only,
no behavior change: `test_models.py`'s unknown-field-rejection test now goes
through `ResultRow.model_validate(dict)` instead of a static kwarg pyright
can check against the model's real fields; `test_rsc_parser.py` narrows a
`str | None` before passing it to `extract_json_object`; `test_scraper.py`'s
`_cfg` helper builds its merged dict via `{**base, **overrides}` instead of
`dict.update()` against a narrower inferred type. `uv run pyright` is now
0 errors (was 4).

New regression tests (fakes, no monkeypatch):
`tests/test_browser.py::test_navigate_with_warmup_waits_for_hydration_when_title_clears_before_content_streams`
reproduces the exact live signature (title clears, first content() read has
no marker, second read does);
`test_read_hydrated_target_content_returns_immediately_when_already_hydrated`
and `test_read_hydrated_target_content_gives_up_after_poll_budget_without_raising`
cover the two boundary cases; three existing `navigate_with_warmup` tests
updated to carry `NEXT_F_MARKER` in their target fixtures so the new poll
doesn't change their unrelated assertions. `tests/test_main.py::
test_enable_geoip_defaults_off` locks the revert. 99 tests green (up from
95\), ruff clean, pyright clean on every file this change touches
(`browser.py`/`scraper.py`/`main.py`/`test_browser.py`/`test_main.py`) — the
4 pre-existing pyright errors in `test_models.py`/`test_rsc_parser.py`/
`test_scraper.py` are unrelated drift, documented since 0.3, untouched here.

Pushed as build `0.5.1` and cloud-QA'd by `actor-qa-engineer` immediately
after: run `jAnCRbM93NJwRskb2` (Store-demo fixture) SUCCEEDED first attempt —
`page=reviews-p1 ... rsc_chunks=7 has_next_f_marker=True` on the very first
read this time, `geoip=False` confirmed in the launch log, no `ERROR`/
`Traceback`. 3 review rows + 1 company-summary row landed, RESIDENTIAL proxy
confirmed via `PROXY_RESIDENTIAL_TRANSFER_GBYTES > 0`, `chargedEventCounts`
`actor-start=1 review-row=3 company-summary-row=1` (PPE active, pricing
unchanged from the live Store listing). Both `output_schema.json` template
URLs (JSON + CSV) returned HTTP 200. QA verdict: PASS.

### 0.3 — 2026-09-03

CEO daily report flagged this listing's Store demo run as failing (evidence:
run `bq6vCeZriEHrScB07`, build 0.1.7, 2026-09-03 08:08Z — `PROXY_RESIDENTIAL_
TRANSFER_GBYTES` on that run confirms RESIDENTIAL via `run_usage.py`, ruling
out a datacenter/degraded-proxy misread). Checking the live Actor detail
first showed 0.2 (challenge-clear polling + bounded SDK awaits) was already
pushed and live as build `0.2.1` since 2026-09-02 09:26Z — the flagged run
had been pinned to the older `0.1.7` build, not `latest`.

Three fresh cloud probes against the *actual* live build (`0.2.1`, the exact
Store-demo input, 1 company/3 reviews): 1 SUCCEEDED (`FgboMQdjOBUPgVYKy`, 4
rows), 2 FAILED (`sB0pIGOu5Nk1BMshl`, `dvahR5AXaySh43Jwl`) — all three
RESIDENTIAL-confirmed via billing. Reading the FAILED logs found the real bug
0.2's poll fix left in place: the Reviews page's own challenge interstitial
titles itself **"Security | Glassdoor"**, not "Just a moment..." (that title
belongs to the Overview warm-up page's Cloudflare-branded challenge).
`_wait_for_challenge_clear`'s `CHALLENGE_TITLE_MARKER` was a single-string
constant that only matched "just a moment" — so on the Reviews-page challenge
variant the poll's very first title check already looked "clear" and
returned after one call (`title_calls == 1`), never actually waiting out the
24s budget the 0.2 fix was built to spend. `scraper.py::_is_challenge_page`
already carried the full marker list (it correctly logged "reviews-p1 served
a challenge page" in both FAILED runs) — the poll just never shared it.

Fix: `browser.py` now owns one `CHALLENGE_TITLE_MARKERS` tuple (5 known
variants) as the single source of truth; `_wait_for_challenge_clear` checks
against all of them, and `scraper.py` imports the same tuple instead of
keeping its own copy. New regression test
(`test_wait_for_challenge_clear_recognizes_glassdoor_security_title`)
reproduces the bug against the pre-fix code (failed with `1 == 3`) and
passes after. 92 tests green, ruff clean, pyright clean on every file this
change touches (browser.py/scraper.py/tests/test\_browser.py) — three
pre-existing pyright errors in `test_models.py`/`test_rsc_parser.py`/
`test_scraper.py` are unrelated drift, present before this change too, out
of scope here.

Also investigated per the fixer brief's lead: `geoip=True` (the Camoufox
`LeakWarning` every run emits). `browser.py`'s own module docstring already
forbids it — a diagnosed 2026-08-06 crash bug (`ai-overview-citations`) where
`geoip=True`'s IP-echo sweep through the proxy raises Camoufox's own
`InvalidIP`/`InvalidProxy`/`LocaleError` before the browser even launches.
`ops/reports/VRBO-GEOIP-CONTRADICTION-2026-08-25.md` records that this
rationale was measured on the FREE-tier pool and this account moved to
STARTER on 2026-08-20, so it needs re-proving, not assuming — and per
`ops/os/CAP-LIFT-CHECKLIST.md` §6b, proving it means testing the launch-
survival question on its own, not bundled with anything else. Live
production evidence from `aliexpress-products-scraper` (runs
`YSBcad4xebfIygDll`, `RqruyZxtXodLrDEdj`, 2026-08-31, same STARTER account,
same RESIDENTIAL pool, `geoip_probe=True` wired unconditionally in
`detail_browser.py`) shows `Launching Camoufox (... geoip_probe=True)` with
no `InvalidIP`/`InvalidProxy`/`LocaleError` crash on either run — the
launch-crash half of the old lesson does not currently reproduce on this
account's tier. That is NOT the same as proving it helps *this* target: no
cloud run against glassdoor-reviews-scraper's own code with `geoip=True` set
was possible in this session (deploy is explicitly out of scope — see
notes.md), so it is wired as an opt-in parameter (`open_browser(..., geoip:
bool = False)`, default OFF, matching the established `yelp-business-
reviews-scraper` pattern) rather than defaulted on. Proving it for Glassdoor
specifically is the next step, once pushed.

country\_code was already pinned to `US` by default (`ActorInput.country_
code`, REQ-10) before this change — the other half of the fixer brief's lead
was already correct.

Zero-row exit semantics (`_finalize`) were also reviewed per the brief: on
zero rows the Actor still raises `SystemExit(EXIT_FAILURE)` (FAILED), and
`EVENT_ACTOR_START` still charges before any scrape attempt (unconditional,
same as every PPE Actor in this fleet — see pricing skill's "Actor start...
covers warmup"). Judgment: this is NOT the `vrbo`/`opentable` "silently
succeed with 0 rows" bug — Glassdoor's per-company retry loop already
distinguishes a genuine block (`AntiBotChallengeError`, raised, retried,
FAILS the run when exhausted) from a real empty parse, so REQ-14's
fail-loud-on-zero-rows is the correct call here, not a bug to flip. It is
also NOT the `vrbo` 100%-blocked pattern that justified delisting: 1/3 fresh
probes SUCCEEDED with real rows on the current live build, consistent with
the ~24-33% probabilistic success rate `notes.md` already documents for
RESIDENTIAL against this target. Charging the actor-start warmup fee on a
failed attempt is the disclosed, fleet-standard PPE contract, not a
mis-implementation — the lever that actually matters is raising the success
rate (this fix + the geoip experiment above), not changing what gets
charged. Pricing is out of scope for this fixer invocation regardless.

### 0.2 — 2026-09-02

Store demo (`companyNames: ["Google"], employerIds: [9079]`) was still
failing on live build 0.1.7 (`demo_health.py`, 2026-09-02) despite 0.1's
`AntiBotChallengeError` retry. Own-run history stayed FAILED-heavy: 13 of 17
30-day runs failed (`fleet_status.py`), and `qa-ledger.jsonl` shows the
latest own probe (run `7I9V8euj7UmHcwxuU`, 2026-09-01) FAILED with the
Reviews-page title/html byte-identical to the pre-navigation challenge page.

This lands two fixes that were prepared and locally tested on
`fix/glassdoor-reviews-scraper-challenge-poll` (2026-09-01) but never
pushed — the branch was intentionally paused pending a fresh evaluation
window for 0.1's fix, and that window is now the demo-health failure above:

1. **Poll for the Cloudflare challenge to clear instead of a flat 3s
   sleep** (`src/browser.py::_wait_for_challenge_clear`). The flat
   `WARMUP_SETTLE_MS=3000` wait after navigating to the Reviews page was
   too short — `ziprecruiter-jobs-scraper`'s recon (2026-08-31) showed the
   same challenge takes 6-18s to self-clear. Now polls up to
   `CHALLENGE_POLL_ATTEMPTS * CHALLENGE_POLL_INTERVAL_MS` (24s) before
   reading page content, same pattern as that Actor.
2. **Bounded every unwrapped `Actor.*` SDK await** (`src/main.py`) —
   fifth fleet instance of the `ops/os/UNBOUNDED-AWAIT-TIMEOUTS-2026-09-01.md`
   pattern. `set_status_message`/`charge` now swallow-and-log past a 60s
   bound; `push_data` raises `PushStalledError` past that bound, caught by
   `_scrape_one_target` so a stall skips only the rest of one company
   instead of the whole run. This actor's own failures were logged FAILED
   not TIMED-OUT, so this is hardening against the fleet-wide pattern, not
   the diagnosed cause of the demo failure — landed together because both
   were already tested and reviewed as one unit.

Both changes are additive to 0.1's `AntiBotChallengeError` retry, which
stays in place unchanged. Glassdoor RESIDENTIAL-proxy access remains
probabilistic by nature (`docs/specs/glassdoor-reviews-scraper/notes.md`)
— this raises the odds a genuinely-clearing challenge is read correctly
instead of read too early, it does not make every run succeed.

`glassdoor-salaries-scraper/src/main.py` has the identical unwrapped-SDK-
await shape — still flagged, still out of scope for this invocation.

### 0.1 — 2026-08-25

Fix for 33% customer success rate (6/10 30-day FAILED runs). Investigation
correction mid-fix, recorded here so the next person doesn't re-walk it:

- **WebShare is currently NOT the answer, despite CLOUD-RECON-RESULT.md's
  2026-08-12 GO finding.** Tried it first (it was never wired as an Actor
  secret — fixed that operationally via the Apify API), but live evidence
  13 days later shows Glassdoor now serves an explicit challenge on the
  exact same technique: warm-up page titled "Security | Glassdoor",
  Reviews page titled "Just a moment..." — 2/2 WebShare probes blocked.
  `WEBSHARE_PROXY_URL` was removed again; **Apify RESIDENTIAL is back to
  being the default proxy** (unchanged from 0.0 — `src/browser.py` already
  preferred it whenever WebShare is unset).
- **Real root cause: a 200-status anti-bot challenge page was silently
  treated as "zero rows found," not as a failure worth retrying.**
  RESIDENTIAL exits intermittently serve that same interstitial instead of
  the real Reviews page — no timeout, no `PWError`, `page.goto` succeeds
  cleanly — so the existing per-company retry (`COMPANY_RETRY_ATTEMPTS`,
  REQ-9) never triggered; the company was just marked empty on the first
  attempt. Live evidence same day: two probe runs back-to-back, one hit
  the challenge, the very next (fresh proxy session) SUCCEEDED with 4 real
  rows and confirmed PPE charges. Added `AntiBotChallengeError` (`src/
  browser.py`, joins `RECOVERABLE_BROWSER_ERRORS`) plus `_is_challenge_page`
  (`src/scraper.py`, anchored on the confirmed-absent `self.__next_f.push(`
  wire-format marker, title as a supplementary signal) — a challenge page
  now raises instead of silently returning empty, so REQ-9's existing
  fresh-session retry actually gets a second exit IP to try.
- Also fixed along the way: `page.content()` could race a still-in-flight
  client-side redirect on the warm-up page (`_read_content_with_retry`,
  `src/browser.py`); added `_log_page_diagnostics` (html length, RSC-chunk
  count, marker presence, `<title>` snippet — never the body) so a future
  zero-rows run is diagnosable from the log alone; added the fleet's
  missing `tests/test_smoke.py` (Gate 7 was a silent no-op without it).

### 0.0 — 2026-08-12

- Initial implementation: Camoufox + WebShare/Apify-Proxy + mandatory
  same-context warm-up navigation (Overview page, then Reviews page)
  unblocks the target — see
  `docs/specs/glassdoor-reviews-scraper/CLOUD-RECON-RESULT.md`. Emits
  `review` and `company_summary` rows. Local test suite green.
