Scrape Glassdoor company reviews by company name: rating breakdowns, pros/cons, job title, employment status, review date, and CEO/recommend-to-friend signals, plus a company rating summary. Camoufox-powered to clear the target's bot defenses. Pay only per row scraped.
resolve_employer_id was a permanent stub that always returned None.companyNames is the Actor's only required input field; employerIds is
documented as an optional "power-user shortcut". Every real customer who
followed the primary, documented usage path — typing a company name with no
employerIds — hit _scrape_company_body's if employer_id is None: return
guard, meaning EVERY company in EVERY such run yielded zero rows. main.py's
REQ-14 zero-rows check then failed the run loud with the status message
"0/N companies parsed — Reviews warm-up path may be blocked for this run's exit IP"
— a deterministic, 100%-reproducible code defect misdiagnosed by
its own status message as a probabilistic anti-bot/IP problem. This is very
likely the dominant contributor to the flagged 30-day customer rate (15/29
FAILED, 48%): our own QA/dev runs always pass employerIds and so never
exercised this path, which is exactly why recent_failures (own runs via
/v2/acts/{id}/runs) showed zero failures on this build while the public
customer-run stat did not — the two populations take different code paths.
Fix: implemented real company-name resolution against Glassdoor's own
/Search/results.htm?keyword={company} suggestion endpoint. Confirmed live
2026-09-19 across 6 companies (Netflix→11891, Google→9079 — matches this
repo's own long-standing QA fixture value, cross-validation not a
coincidence, Meta→40772 — independently matches
glassdoor-salaries- scraper
's own confirmed value, Deloitte→2763, Bank of America→8874,
McDonald's→432) plus a nonsense query correctly returning no match. Unlike
Overview/Reviews, this endpoint 200s on a cold navigation — no
same-context warm-up required (4/4 live probes). New: browser.py's
build_search_url/fetch_search_results, rsc_parser.py's
parse_company_search_suggestion (anchored on the directHit key — unique
per probed page, unlike employerId which the ratings object also carries).
scraper.py::resolve_employer_id now navigates the passed-in page instead
of receiving an unused browser handle, raises AntiBotChallengeError
(RECOVERABLE_BROWSER_ERRORS — retried by main.py's REQ-9 driver with a
fresh proxy/session) when the search page itself is a genuine anti-bot
interstitial, and returns None only for an ordinary no-match.
Two new live-captured fixtures: tests/fixtures/glassdoor_search_results.html
(Netflix direct hit) and glassdoor_search_no_match.html (nonsense query,
genuine zero-match) — never hand-authored, both real captures per this
repo's reference-hand-authored-fixtures-encode-fiction lesson.
Local verification: apify run with {"companyNames": ["Netflix"]} (no
employerIds — the realistic customer shape) now completes exit_code: 0
with 3 review rows + 1 company-summary row, after correctly classifying and
retrying one genuine challenge on the search-resolution step itself. The
existing employerIds-supplied QA fixture (Google/9079) still passes
unchanged — no regression on the previously-working path.
Known residual, not addressed this release: a genuinely nonexistent/
misspelled company name (a true no-match, not a block) still routes through
main.py's unconditional zero-rows-across-all-companies fail-loud path with
the same misleading "may be blocked" message. This was already true before
this fix (indistinguishable from the stub) and is a strict improvement, not
a regression, but the "empty search should SUCCEED with zero rows" rule
(EMPTY-IS-NOT-A-FAILURE-2026-08-19.md) should eventually be applied here
too — needs main.py/scraper.py to propagate a "some companies resolved,
some genuinely had no match" distinction that doesn't exist yet.
0.8.0 — 2026-09-14
Fixed
The ~10x under-delivery: every SUCCEEDED run capped out at exactly 3 review
rows regardless of maxReviewsPerCompany. Traced to production, not
guesswork: 5 cloud runs (pRgDfifO22exTlnwL, z4hA8U69gmT7dc4bu,
4eE1YjYb7kF2uDvsK, 5hPUeVWddEnvY8c4T, EqiMYtgtKQeADnHeH), all on Apify
RESIDENTIAL pinned to US, each hit exactly 3 rows for the shipped QA
prefill (Google, maxReviewsPerCompany: 30). A local recon capture
reproduced the same 3-row ceiling on a different proxy tier/country
entirely (WebShare/NL), and it held steady across 60+ seconds of extra
wait — ruling out a hydration-timing bug. The real cause: Glassdoor bakes
pageSize:3 into every anonymous session's review-list filters (confirmed
in the live RSC payload's own initFilters), then blurs anything past
those 3 behind a login wall (BlurredOverlay/ReviewList_blurOverlayWrapper
in the live DOM). Numbered pagination (_P2.htm+) doesn't even try to
blur — it redirects the entire page to Log In | Glassdoor, for every
company, every proxy, every geography tested.
find_review_entries was never the bug — it correctly parsed the only 3
real review objects Glassdoor ever sent an anonymous session. The fix is a
new pagination source: the Reviews page's own "Users say..." keyword-
highlight links (data-test="review-highlight-link", e.g.
/Reviews/Google-people-Reviews-EI_IE9079...htm) are real, publicly-linked
Reviews sub-pages that carry a materially larger anonymous pageSize per
keyword (9 reviews sampled live for two different keywords) — this is now
scraper.py's actual second-page path (_rows_from_highlight_pages,
rsc_parser.find_highlight_review_urls), tried before numbered pagination,
which remains only as a fallback for companies too small to carry
highlight links.
Added _is_login_wall, checked ahead of _is_challenge_page in every
pagination walker: retrying the login wall with a fresh proxy/session
(main.py's REQ-9 driver) can never clear it, so it now stops cleanly
instead of being mistaken for a retry-worthy anti-bot interstitial or,
worse, silently falling through if not entries: return by accident.
_log_page_diagnostics now runs for every subsequent/highlight page, not
just page 1 — the pagination path used to be invisible past the first
fetch.
Fixed a real infinite-loop risk surfaced by this work's own tests: a page
that returns non-empty entries which are all already in seen (zero
new rows) now stops pagination like an empty page would, instead of
looping forever waiting for a count that can never advance.
Shipped: apify push --force → version 0.8, build POLqouhTfacuMdOcF
(buildNumber 0.8.1), tagged latest. PPE pricing unchanged from the live
Store listing (per Gate 4b in the preceding dry-run) — no monetization
change this release.
Post-push cloud verification (mandatory actor-qa-engineer gate, both
runs against build=latest):
20 review + 1 company_summary, all 20 review_ids unique
The pagination-proof run's log shows the fix firing live in production —
reviews-p1 alone wasn't enough to reach 20, so three distinct
page=reviews-highlight fetches followed ("people", "benefit",
"culture" keyword sub-pages), each has_next_f_marker=True, none a
duplicate — this is direct cloud proof the highlight-link pagination path
(new in this version) delivers past the old 3-row ceiling, not just a local
claim. chargedEventCounts on both runs matched dataset row counts exactly
(review-row=20/company-summary-row=1 and review-row=3/
company-summary-row=1), no ERROR/Traceback in either log, both exited
clean (exit_code: 0). QA verdict: PASS.
0.7.0 — 2026-09-14
Fixed
The start fee is owed on delivery, not on boot.actor-start was charged
immediately after input validation — before proxy resolution and before a
single page load — so every refused run billed the customer for an empty
dataset. Glassdoor turns away roughly three attempts in four, which makes
failing runs a permanent feature of this Actor rather than an anomaly, so
this was not a rare edge: at the pre-0.6 success rate it meant ~16 of every
29 customer runs paid for nothing. Live proof, both from 0.6.1 QA:
zVX1fjX2IRKBzGqLd and mZ9phI4xBBpE136Tl, FAILED, each
{'actor-start': 1, 'review-row': 0, 'company-summary-row': 0}.
Billing a start fee for a run that delivered nothing is the shape that got
vrbo-vacation-rentals-scraper delisted at 17 runs. _StartFee now fires
at most once, and only from the first successful row push.
Glassdoor Company Reviews Scraper — Changelog
0.6 — 2026-09-14
CEO daily report flagged UNDER_MAINTENANCE: 44.8% 30-day customer success
(13/29), live build 0.5.1 (2026-09-09) — matching main, so no undeployed
fix was sitting on a branch; this needed a real one.
Diagnosed from production, not guesswork./v2/acts/{id}/runs (own
runs) had nothing newer than 2026-09-09 — the 29 customer runs behind the
44.8% figure, including one as recent as 2026-09-13, are invisible to us
(they run under the customer's account, not ours). Read every own-run FAILED
log back to 2026-08-24 instead. All of them, including the two most recent
(9hqNREpL769oeRmuK, gMdegWCzIK9NTKXvK) — pre-0.5.1 builds — show the
per-company retry loop working exactly as designed: AntiBotChallengeError
raised, fresh proxy/session on each attempt, COMPANY_RETRY_ATTEMPTS
attempts exhausted, company skipped, run correctly fails loud only when
every company came back empty (REQ-14). No fault-isolation leak, no soft
200-as-block miss, no stale shared session — patterns 1-4 in the fixer brief
all check out clean; _is_challenge_page's NEXT_F_MARKER/title checks (0.3,
0.5) already cover pattern 2 correctly.
The one own-run sample actually on the live build (jAnCRbM93NJwRskb2,
0.5.1, 2026-09-09 21:57) SUCCEEDED — the 0.5 hydration-race fix's target
condition (rsc_chunks=7 on the very first content read) confirmed live.
That build is code-correct; the residual FAILED share is not a bug riding on
top of the wall, it's the wall itself, and docs/specs/glassdoor-reviews-scraper/notes.md's
2026-09-03 measurement already named the exact number: Glassdoor is
probabilistically reachable at roughly p≈0.25 per attempt, and with
COMPANY_RETRY_ATTEMPTS=2, P(at least one clears) = 1-(1-p)^2 ≈ 44% — a
near-exact match for the live 44.8% that got this Actor flagged. That note
named raising the attempt budget as "the only change measured here that
would move delivery" but left it as a cost/benefit call for whoever owns the
trade-off. That's this session, and UNDER_MAINTENANCE makes the call: ship
it.
Fix: src/main.py's COMPANY_RETRY_ATTEMPTS raised 2 -> 4
(predicted P(pass) ≈ 68%). Each retry already launches a fresh
Camoufox browser on a fresh proxy session (REQ-9, unchanged) — this only
widens how many times that already-correct retry gets to run before the
company is skipped. Nothing else touched: AntiBotChallengeError (0.1),
the challenge-title poll (0.3), the hydration-race poll (0.5), and REQ-14's
fail-loud-on-zero-rows semantics all stay exactly as documented. Pricing/
pay_per_event.json untouched — actor-start is still charged once per run
regardless of per-company attempt count, so the customer's cost per run does
not change; only our own compute cost on a company that ultimately fails
anyway goes up, traded for materially more companies succeeding and billing
review-row/company-summary-row events.
New regression test: tests/test_main.py::test_retry_driver_uses_widened_attempt_budget
— fails red against the pre-fix code (COMPANY_RETRY_ATTEMPTS == 2,
open_browser called twice), asserts both the constant and the actual
open_browser call count against an always-failing company, so a future
edit can't silently narrow the budget back.
Verification: 100 tests green (up from 99: 1 new). ruff check clean.
ruff format --check has pre-existing drift in 7 unrelated files (confirmed
via git stash diff against this branch's base — not introduced here);
src/main.py/tests/test_main.py carry the same pre-existing wrapping
drift on lines this change didn't touch. pyright 0 errors.
Shipped: apify push → version 0.6, build gkOPijKcEox0R1Eia
(buildNumber 0.6.1), tagged latest. verify_ppe_sanity.py --live
confirmed no live-vs-repo pricing drift.
Post-push cloud verification (mandatory QA gate + 4 additional probes,
per this session's brief — "two greens is not a rate; you need 4+" —
against build=latest with the unmodified shipped fixture):
run
status
rows
mZ9phI4xBBpE136Tl (QA gate)
FAILED
0
R6tRESMwI2ts4EqPQ
SUCCEEDED
4 (real pros/cons/ratings)
BMFhvqSvoO0boMRL5
SUCCEEDED
4
HfJ8fj2XZ9dd7YSUy
SUCCEEDED
4
zVX1fjX2IRKBzGqLd
FAILED
0
4/5 = 80% — above the ~68% point prediction, well above the pre-fix ~44%
ceiling. Both FAILED runs show 4/4 attempts hitting an unambiguous
title='Just a moment...'/rsc_chunks=0 challenge on every single try
(not the hydration-race shape) — a clean, correctly-classified, correctly
fail-loud genuine block, not a defect. Full run-by-run detail:
docs/specs/glassdoor-reviews-scraper/notes.md's 2026-09-14 entry.
Next step for whoever reviews this: per notes.md's own caution, n=5
across this session (1 QA + 4 verification) is still directional, not
statistically sound (that needs n>=20/arm) — treat this as "the fix shipped
and measurably outperforms the pre-fix ceiling," not as final proof the true
rate is exactly 68%. If the 30-day customer number hasn't moved meaningfully
a week out, the next untried lever is HYDRATION_POLL_ATTEMPTS/
CHALLENGE_POLL_ATTEMPTS (the per-attempt success probability itself) or a
further budget raise (n=6, ~82% predicted) — both should be measured against
fresh evidence, not assumed.
0.5 — 2026-09-09
CEO daily report flagged 41% 30-day customer success (14/24 FAILED), 2 users,
live build 0.4.2 (4 days old). /acts/{id}/runs (own runs) had nothing newer
than 2026-09-03 — the 24 FAILED/SUCCEEDED customer runs were invisible to us,
so this session reproduced fresh: a free local recon through Apify RESIDENTIAL
(plain curl, no Camoufox) got the expected 403 Security | Glassdoor — a
non-browser client always sees Cloudflare's wall regardless of this Actor's
health, so it only confirms the target is still defended, not whether our code
is broken. Two real cloud probes against live build 0.4.2 (Store-demo input,
t152VidKGrQlNop2f SUCCEEDED, 9hqNREpL769oeRmuK FAILED; ~$0.02 total,
run_usage.py-checked RESIDENTIAL) found two distinct, previously
undiagnosed contributors on top of the already-documented probabilistic
anti-bot wall (docs/specs/glassdoor-reviews-scraper/notes.md, 2026-09-03):
A false "blocked" classification, not a real block.9hqNREpL769oeRmuK's
first attempt:
— a REAL company title (never a
CHALLENGE_TITLE_MARKERS match), yet _is_challenge_page scored it as a
challenge anyway because the RSC hydration marker was absent. Root cause:
Glassdoor's <title> renders before Next.js finishes streaming its RSC
payload; _wait_for_challenge_clear polls only the title, so it returned
the instant the title looked real, and the very next page.content()
call captured a still-streaming ~18KB shell instead of the ~250KB+
hydrated page. That burned a whole REQ-9 retry attempt (half of
COMPANY_RETRY_ATTEMPTS=2) on a page that was never actually blocked —
confirmed as a recurring pattern, not a one-off: the identical signature
(~18-19KB, real title, zero RSC chunks) already appears in
PtHaK788kRd2QwwcN's log from the 2026-09-03 investigation, unnoticed at
the time because that session's read attributed the whole failure class to
the probabilistic Cloudflare wall.
Fix: src/browser.py gains _read_hydrated_target_content — after the
title-based challenge-clear poll returns, this polls page.content()
itself (bounded, HYDRATION_POLL_ATTEMPTS=4 / 1s apart) until
NEXT_F_MARKER appears or the short budget is exhausted, then hands off
to the existing classification unchanged. NEXT_F_MARKER moved from
scraper.py into browser.py as the single source of truth (same pattern
as CHALLENGE_TITLE_MARKERS's 2026-09-03 move) so both the hydration wait
and the challenge classifier agree on one signal.
ENABLE_GEOIP=True (flipped on in 0.4, unproven) burned a retry attempt
on an unrelated transient.t152VidKGrQlNop2f's successful run still
logged
attempt=1/2 cause=Failed to getIPaddress:...ipecho.net... SSLEOFError
before attempt 2 succeeded — Camoufox's own geoip IP-echo
probe sweep (through the residential proxy, before Firefox even launches)
failing on an unrelated third-party service, consuming one of only two
attempts against Glassdoor's actual wall. The 2026-09-03 notes already
judged geoip=True "not shown to help" at n=9 (aggregate 2/9 pass, same as
geoip=False); this session found a concrete cost that measurement hadn't
surfaced. Reverted src/main.py's ENABLE_GEOIP to False (its own
documented default before the 0.4 experiment) — an unproven flag that can
measurably halve the effective retry budget doesn't stay on by a coin
flip.
Both fixes are additive to the existing AntiBotChallengeError retry chain
(0.1) and the challenge-title poll (0.3) — neither is touched, both stay as
documented. This does not change the fundamental probabilistic nature of
Glassdoor's block (docs/specs/glassdoor-reviews-scraper/notes.md's
attempt-budget math still applies); it removes two sources of false
failures riding on top of it.
Also cleared Gate 7 for publish: the pre-flight /publish-actor dry-run (this
session) found the 4 pyright errors carried since 0.3 in test_models.py/
test_rsc_parser.py/test_scraper.py now block publish outright (Gate 7 is
binary on zero errors, regardless of blame). Fixed all 3 sites — typing-only,
no behavior change: test_models.py's unknown-field-rejection test now goes
through ResultRow.model_validate(dict) instead of a static kwarg pyright
can check against the model's real fields; test_rsc_parser.py narrows a
str | None before passing it to extract_json_object; test_scraper.py's
_cfg helper builds its merged dict via {**base, **overrides} instead of
dict.update() against a narrower inferred type. uv run pyright is now
0 errors (was 4).
New regression tests (fakes, no monkeypatch):
tests/test_browser.py::test_navigate_with_warmup_waits_for_hydration_when_title_clears_before_content_streams
reproduces the exact live signature (title clears, first content() read has
no marker, second read does);
test_read_hydrated_target_content_returns_immediately_when_already_hydrated
and test_read_hydrated_target_content_gives_up_after_poll_budget_without_raising
cover the two boundary cases; three existing navigate_with_warmup tests
updated to carry NEXT_F_MARKER in their target fixtures so the new poll
doesn't change their unrelated assertions.
locks the revert. 99 tests green (up from
95), ruff clean, pyright clean on every file this change touches
(browser.py/scraper.py/main.py/test_browser.py/test_main.py) — the
4 pre-existing pyright errors in test_models.py/test_rsc_parser.py/
test_scraper.py are unrelated drift, documented since 0.3, untouched here.
Pushed as build 0.5.1 and cloud-QA'd by actor-qa-engineer immediately
after: run jAnCRbM93NJwRskb2 (Store-demo fixture) SUCCEEDED first attempt —
page=reviews-p1 ... rsc_chunks=7 has_next_f_marker=True on the very first
read this time, geoip=False confirmed in the launch log, no ERROR/
Traceback. 3 review rows + 1 company-summary row landed, RESIDENTIAL proxy
confirmed via PROXY_RESIDENTIAL_TRANSFER_GBYTES > 0, chargedEventCountsactor-start=1 review-row=3 company-summary-row=1 (PPE active, pricing
unchanged from the live Store listing). Both output_schema.json template
URLs (JSON + CSV) returned HTTP 200. QA verdict: PASS.
0.3 — 2026-09-03
CEO daily report flagged this listing's Store demo run as failing (evidence:
run bq6vCeZriEHrScB07, build 0.1.7, 2026-09-03 08:08Z —
PROXY_RESIDENTIAL_TRANSFER_GBYTES
on that run confirms RESIDENTIAL via run_usage.py, ruling
out a datacenter/degraded-proxy misread). Checking the live Actor detail
first showed 0.2 (challenge-clear polling + bounded SDK awaits) was already
pushed and live as build 0.2.1 since 2026-09-02 09:26Z — the flagged run
had been pinned to the older 0.1.7 build, not latest.
Three fresh cloud probes against the actual live build (0.2.1, the exact
Store-demo input, 1 company/3 reviews): 1 SUCCEEDED (FgboMQdjOBUPgVYKy, 4
rows), 2 FAILED (sB0pIGOu5Nk1BMshl, dvahR5AXaySh43Jwl) — all three
RESIDENTIAL-confirmed via billing. Reading the FAILED logs found the real bug
0.2's poll fix left in place: the Reviews page's own challenge interstitial
titles itself "Security | Glassdoor", not "Just a moment..." (that title
belongs to the Overview warm-up page's Cloudflare-branded challenge).
_wait_for_challenge_clear's CHALLENGE_TITLE_MARKER was a single-string
constant that only matched "just a moment" — so on the Reviews-page challenge
variant the poll's very first title check already looked "clear" and
returned after one call (title_calls == 1), never actually waiting out the
24s budget the 0.2 fix was built to spend. scraper.py::_is_challenge_page
already carried the full marker list (it correctly logged "reviews-p1 served
a challenge page" in both FAILED runs) — the poll just never shared it.
Fix: browser.py now owns one CHALLENGE_TITLE_MARKERS tuple (5 known
variants) as the single source of truth; _wait_for_challenge_clear checks
against all of them, and scraper.py imports the same tuple instead of
keeping its own copy. New regression test
(test_wait_for_challenge_clear_recognizes_glassdoor_security_title)
reproduces the bug against the pre-fix code (failed with 1 == 3) and
passes after. 92 tests green, ruff clean, pyright clean on every file this
change touches (browser.py/scraper.py/tests/test_browser.py) — three
pre-existing pyright errors in test_models.py/test_rsc_parser.py/
test_scraper.py are unrelated drift, present before this change too, out
of scope here.
Also investigated per the fixer brief's lead: geoip=True (the Camoufox
LeakWarning every run emits). browser.py's own module docstring already
forbids it — a diagnosed 2026-08-06 crash bug (ai-overview-citations) where
geoip=True's IP-echo sweep through the proxy raises Camoufox's own
InvalidIP/InvalidProxy/LocaleError before the browser even launches.
ops/reports/VRBO-GEOIP-CONTRADICTION-2026-08-25.md records that this
rationale was measured on the FREE-tier pool and this account moved to
STARTER on 2026-08-20, so it needs re-proving, not assuming — and per
ops/os/CAP-LIFT-CHECKLIST.md §6b, proving it means testing the launch-
survival question on its own, not bundled with anything else. Live
production evidence from aliexpress-products-scraper (runs
YSBcad4xebfIygDll, RqruyZxtXodLrDEdj, 2026-08-31, same STARTER account,
same RESIDENTIAL pool, geoip_probe=True wired unconditionally in
detail_browser.py) shows Launching Camoufox (... geoip_probe=True) with
no InvalidIP/InvalidProxy/LocaleError crash on either run — the
launch-crash half of the old lesson does not currently reproduce on this
account's tier. That is NOT the same as proving it helps this target: no
cloud run against glassdoor-reviews-scraper's own code with geoip=True set
was possible in this session (deploy is explicitly out of scope — see
notes.md), so it is wired as an opt-in parameter (
open_browser(..., geoip: bool = False)
, default OFF, matching the established
yelp-business- reviews-scraper
pattern) rather than defaulted on. Proving it for Glassdoor
specifically is the next step, once pushed.
country_code was already pinned to US by default (
ActorInput.country_ code
, REQ-10) before this change — the other half of the fixer brief's lead
was already correct.
Zero-row exit semantics (_finalize) were also reviewed per the brief: on
zero rows the Actor still raises SystemExit(EXIT_FAILURE) (FAILED), and
EVENT_ACTOR_START still charges before any scrape attempt (unconditional,
same as every PPE Actor in this fleet — see pricing skill's "Actor start...
covers warmup"). Judgment: this is NOT the vrbo/opentable "silently
succeed with 0 rows" bug — Glassdoor's per-company retry loop already
distinguishes a genuine block (AntiBotChallengeError, raised, retried,
FAILS the run when exhausted) from a real empty parse, so REQ-14's
fail-loud-on-zero-rows is the correct call here, not a bug to flip. It is
also NOT the vrbo 100%-blocked pattern that justified delisting: 1/3 fresh
probes SUCCEEDED with real rows on the current live build, consistent with
the ~24-33% probabilistic success rate notes.md already documents for
RESIDENTIAL against this target. Charging the actor-start warmup fee on a
failed attempt is the disclosed, fleet-standard PPE contract, not a
mis-implementation — the lever that actually matters is raising the success
rate (this fix + the geoip experiment above), not changing what gets
charged. Pricing is out of scope for this fixer invocation regardless.
0.2 — 2026-09-02
Store demo (companyNames: ["Google"], employerIds: [9079]) was still
failing on live build 0.1.7 (demo_health.py, 2026-09-02) despite 0.1's
AntiBotChallengeError retry. Own-run history stayed FAILED-heavy: 13 of 17
30-day runs failed (fleet_status.py), and qa-ledger.jsonl shows the
latest own probe (run 7I9V8euj7UmHcwxuU, 2026-09-01) FAILED with the
Reviews-page title/html byte-identical to the pre-navigation challenge page.
This lands two fixes that were prepared and locally tested on
fix/glassdoor-reviews-scraper-challenge-poll (2026-09-01) but never
pushed — the branch was intentionally paused pending a fresh evaluation
window for 0.1's fix, and that window is now the demo-health failure above:
Poll for the Cloudflare challenge to clear instead of a flat 3s
sleep (src/browser.py::_wait_for_challenge_clear). The flat
WARMUP_SETTLE_MS=3000 wait after navigating to the Reviews page was
too short — ziprecruiter-jobs-scraper's recon (2026-08-31) showed the
same challenge takes 6-18s to self-clear. Now polls up to
CHALLENGE_POLL_ATTEMPTS * CHALLENGE_POLL_INTERVAL_MS (24s) before
reading page content, same pattern as that Actor.
Bounded every unwrapped Actor.* SDK await (src/main.py) —
fifth fleet instance of the ops/os/UNBOUNDED-AWAIT-TIMEOUTS-2026-09-01.md
pattern. set_status_message/charge now swallow-and-log past a 60s
bound; push_data raises PushStalledError past that bound, caught by
_scrape_one_target so a stall skips only the rest of one company
instead of the whole run. This actor's own failures were logged FAILED
not TIMED-OUT, so this is hardening against the fleet-wide pattern, not
the diagnosed cause of the demo failure — landed together because both
were already tested and reviewed as one unit.
Both changes are additive to 0.1's AntiBotChallengeError retry, which
stays in place unchanged. Glassdoor RESIDENTIAL-proxy access remains
probabilistic by nature (docs/specs/glassdoor-reviews-scraper/notes.md)
— this raises the odds a genuinely-clearing challenge is read correctly
instead of read too early, it does not make every run succeed.
glassdoor-salaries-scraper/src/main.py has the identical unwrapped-SDK-
await shape — still flagged, still out of scope for this invocation.
0.1 — 2026-08-25
Fix for 33% customer success rate (6/10 30-day FAILED runs). Investigation
correction mid-fix, recorded here so the next person doesn't re-walk it:
WebShare is currently NOT the answer, despite CLOUD-RECON-RESULT.md's
2026-08-12 GO finding. Tried it first (it was never wired as an Actor
secret — fixed that operationally via the Apify API), but live evidence
13 days later shows Glassdoor now serves an explicit challenge on the
exact same technique: warm-up page titled "Security | Glassdoor",
Reviews page titled "Just a moment..." — 2/2 WebShare probes blocked.
WEBSHARE_PROXY_URL was removed again; Apify RESIDENTIAL is back to
being the default proxy (unchanged from 0.0 — src/browser.py already
preferred it whenever WebShare is unset).
Real root cause: a 200-status anti-bot challenge page was silently
treated as "zero rows found," not as a failure worth retrying.
RESIDENTIAL exits intermittently serve that same interstitial instead of
the real Reviews page — no timeout, no PWError, page.goto succeeds
cleanly — so the existing per-company retry (COMPANY_RETRY_ATTEMPTS,
REQ-9) never triggered; the company was just marked empty on the first
attempt. Live evidence same day: two probe runs back-to-back, one hit
the challenge, the very next (fresh proxy session) SUCCEEDED with 4 real
rows and confirmed PPE charges. Added AntiBotChallengeError (
src/ browser.py
, joins RECOVERABLE_BROWSER_ERRORS) plus _is_challenge_page
(src/scraper.py, anchored on the confirmed-absent self.__next_f.push(
wire-format marker, title as a supplementary signal) — a challenge page
now raises instead of silently returning empty, so REQ-9's existing
fresh-session retry actually gets a second exit IP to try.
Also fixed along the way: page.content() could race a still-in-flight
client-side redirect on the warm-up page (_read_content_with_retry,
src/browser.py); added _log_page_diagnostics (html length, RSC-chunk
count, marker presence, <title> snippet — never the body) so a future
zero-rows run is diagnosable from the log alone; added the fleet's
missing tests/test_smoke.py (Gate 7 was a silent no-op without it).
0.0 — 2026-08-12
Initial implementation: Camoufox + WebShare/Apify-Proxy + mandatory
same-context warm-up navigation (Overview page, then Reviews page)
unblocks the target — see
docs/specs/glassdoor-reviews-scraper/CLOUD-RECON-RESULT.md. Emits
review and company_summary rows. Local test suite green.