# Changelog of SEO Keyword & Search-Demand Intel (`seibs.co/keyword-demand-intel`) Actor

- **URL**: https://apify.com/seibs.co/keyword-demand-intel/changelog.md
- **Full Actor documentation**: https://apify.com/seibs.co/keyword-demand-intel.md

## Changelog

### 0.2 - 2026-09-27 - Reliability: shared browser, JSON legs off the browser, honest partial runs

- **Browser tier rebuilt around one shared Chromium per run** (`src/browser_session.py`): a fresh isolated context per SERP, pages capped by run memory (1 below 2 GB, 2 at 2 GB, 3 at 4 GB+), a 60-second cap per fetch, and a circuit breaker that turns the tier off for the rest of the run after 3 consecutive failures or challenge pages. Previously every escalated request launched its own browser, up to four at once, which exhausts a 1 GB container. Headful runs start a private Xvfb when the image ships one and no display is running; `--disable-dev-shm-usage` keeps renderers alive on the 64 MB container /dev/shm; if patchright's own Chromium build is missing the Playwright build is used.
- **Autosuggest and Trends never escalate to the browser.** They are JSON endpoints; a browser only wraps the JSON in a viewer page. Their ladder is now httpx -> curl_cffi.
- **Trends works again:** the client warms the public Trends explore page once per run so Google issues the NID cookie `/trends/api` requires (without it every call answered 429). Verified live: 3/3 seeds returned 54-point weekly series. The CONSENT cookie moved from a raw header into the cookie jar so issued cookies are sent. A Trends breaker skips the rest of the run's Trends lookups after 2 consecutive misses.
- **SERP parser reads the JavaScript-rendered SERP:** organic results now link through opaque `/goto?url=` redirects, so the result's site is read from its `<cite>` breadcrumb (origin only; breadcrumb labels are not URL paths). Verified live: 8 organic results with domains on "crm software" (previously 0). New `ai_overview` SERP feature.
- **Time-aware runs:** the compute cap shrinks to fit the platform timeout (`ACTOR_TIMEOUT_AT`) minus a 40-second emit reserve; fetches still running at the deadline are cancelled and the run finishes SUCCEEDED with what finished. SERP terms not reached are listed as uncharged `fetch_error` records.
- **Charging honesty:** a seed whose every autosuggest prefix was blocked emits no keyword record (and no charge) - a `fetch_error` says why. A soft-blocked SERP (no parseable content) emits no snapshot and no charge. Every emitted keyword is charged exactly once; the time cap no longer skips accounting for records already built.
- **Run status message** summarizes what came back (keywords, SERPs, Trends signals) and what was blocked or not reached; a run with nothing extracted says so, pushes an `availability_notice`, and costs $0.
- `access_notes.anti_bot_escalation.browser_session` reports the browser mode, launches, pages served, and breaker state.
- Default run memory 1024 -> 2048 MB (a headful Chromium rendering Google's SERP peaks near 1 GB).

### 0.1 - 2026-06-12 - Initial build

- First release of **SEO Keyword & Search-Demand Intel**. Four modes: `keyword_expansion`, `serp_analysis`, `keyword_metrics`, `trend_signal`.
- **Sources** (all logged-out, public, keyless): Google autosuggest (`/complete/search`) for long-tail expansion + demand-rank signal; the logged-out web SERP (`/search`) for features, top results, People-Also-Ask, related searches, and difficulty; Google Trends (unofficial `/trends/api`, best-effort) for seasonality.
- **Estimation model** (`estimation.py`, versioned `kdi-demand-v1`): transparent, deterministic intent classifier (5 classes), banded volume estimator (autosuggest rank + ubiquity + query shape + SERP corroboration) with explicit confidence, and a 0-100 difficulty score (SERP-backed when a SERP is fetched, else a documented proxy). Estimates are clearly labeled as model output, not Google's gated figures.
- **Value layer**: search-intent classification, topic clustering (stop-word-insensitive near-dup grouping), and the related/PAA mining that turns one SERP into a fan-out of bonus keywords.
- **Anti-bot escalation** for the SERP leg: `httpx (datacenter)` -> `curl_cffi Chrome TLS impersonation (residential)` -> `Playwright stealth browser (residential)` -> fail-soft. Optional warm-browser CDP endpoint (`browser_cdp_url` / `BROWSER_CDP_URL`) and an opt-in CAPTCHA solver hook (off by default). Autosuggest clears on plain httpx, so a run is never empty.
- **Cost control** (portfolio 5-layer defense): pre-flight input caps, demo-mode soft-fail on any error (run finishes SUCCEEDED), `_RunBudget` compute-vs-revenue guard, in-loop budget checks, hardcoded PPE prices.
- **Monitor mode**: scheduled-run change digest (rank/difficulty/SERP-feature deltas) with optional Slack webhook; charges `scheduled_delta_run`.
- Normalized dataset schema with `overview` / `keywords` / `clusters` / `serp` views, MCP-compatible output link, CSV + per-keyword content-worksheet artifacts.
- Paired MCP twin (`mcp-keyword-demand-intel`) with five agent tools and x402 / Skyfire agentic-payment hooks.
- 28-test offline smoke suite (`scripts/smoke_test_keyword_demand.py`) covering the estimation models, autosuggest/SERP/Trends parsing against real-shape fixtures, the escalation ladder, budget/preflight guards, demo routing, config integrity, and the MCP catalog.
