# Changelog of Kick Chat Scraper & Real-Time Archive (`devilscrapes/kick-chat-archive`) Actor

- **URL**: https://apify.com/devilscrapes/kick-chat-archive/changelog.md
- **Full Actor documentation**: https://apify.com/devilscrapes/kick-chat-archive.md

## Kick.com Chat Archive — Changelog

### 0.15 — 2026-09-01 (actor-fixer)

- Daily fleet triage flagged this Actor at 75% 30-day customer success
  (24/32 succeeded, 1 ABORTED, **7 TIMED-OUT**), `users_30d=20`. Build
  `0.14.1` (2026-08-25) was 7 days old at measurement time, so 23 of the
  30 trailing days predate it — the same "stale window" shape this
  Actor's history already documents twice (v0.7, v0.8, and the 2026-08-26
  verification entry below). Own runs (`/acts/{id}/runs`) and the
  account-wide `/v2/actor-runs` feed both showed zero failures — confirmed
  directly that **both** endpoints only ever return runs started by our
  own token (10 total, all `API`/`CLI` origin, all `SUCCEEDED`), never the
  32 customer (`WEB`/scheduler-origin) runs `publicActorRunStats30Days`
  counts — so, per `reference-fleet-health-signal`, individual customer
  run timestamps/logs are **not obtainable via the API at all**, not just
  inconvenient to fetch. Could not settle "still failing on 0.14.1?" from
  run timestamps as a result; treated that as inconclusive rather than a
  clean bill of health and instead re-audited the current source for any
  remaining unbounded-stall path of the same shape as the bug 0.14.1 fixed.
- **Found one.** `PUSH_TIMEOUT_S` (introduced 2026-08-25) bounds
  `Actor.push_data()`/`Actor.charge()` against apify\_client's own
  legitimately-unbounded (~48-minute) retry policy on a platform-side
  stall — but every `Actor.set_status_message()`/`Actor.set_value()` call
  in `main.py` was left unwrapped, sharing the identical HTTP layer and
  retry policy. Several of these are the **last** call almost every run
  path makes — most importantly the plain success message at the bottom
  of `main()`, reached only *after* every row is already pushed and every
  PPE event already charged. A stall there is indistinguishable, from the
  customer's and the fleet-health dashboard's point of view, from the
  exact TIMED-OUT defect class `PUSH_TIMEOUT_S` was built to fix: a
  functionally-complete run — data landed, customer billed correctly —
  reported TIMED-OUT by the platform anyway because the very last SDK
  call of the run hung on a stalled run-update API call.
- **Fix**: added `_set_status_message()` and `_set_value()` — thin
  `asyncio.wait_for(..., timeout=PUSH_TIMEOUT_S)` wrappers mirroring the
  existing `_charge()` guard (swallow-and-log on stall, never raise —
  callers on the fail-loud path still raise `SystemExit` themselves
  afterward). Replaced all 7 `Actor.set_status_message(...)` call sites
  and the 1 `Actor.set_value(...)` call site in `main.py` with the bounded
  wrappers. No behavioural change on the happy path; only bounds the
  worst-case wall-clock a stall can consume.
- Added 5 new tests (`test_set_status_message_swallows_a_stall`,
  `test_set_status_message_succeeds_within_timeout`,
  `test_set_value_swallows_a_stall`, `test_set_value_succeeds_within_timeout`,
  `test_fail_on_silent_parse_failure_still_exits_when_set_value_stalls`)
  mirroring the existing `_push_batch`/`_charge` stall-coverage pattern.
  Full suite: 68/68 green, `ruff check` clean, `pyright` 0 errors.
- **Not verified**: could not confirm this fix against a live customer
  run — the API gives no visibility into customer-run logs/timestamps
  (see above), and this session did not `apify push` (live Actor, 20
  real users; push + cloud QA belongs to the human/publish lane per the
  fixer brief). Local reproduction of a genuine platform-side
  `set_status_message` stall isn't practical without cooperation from the
  platform itself, so this fix is proven by unit test (the stall is
  bounded and swallowed) and by code inspection (it closes the one
  remaining unwrapped SDK round-trip of the same shape as the already-
  confirmed 0.14.1 fix), not by a fresh cloud reproduction of the original
  symptom.

### 0.14 — 2026-08-26 (verification session, no code change)

- Daily signal: 30-day customer health 60% (13/33 runs failed, **all 13
  TIMED-OUT**), `users_30d=20` — this Actor's highest-traffic listing,
  flagged as the single biggest customer-facing defect by users affected.
  Own runs (`/acts/{id}/runs`) showed zero failures — the standard
  asymmetry (`reference-fleet-health-signal`): customer runs fail, ours
  don't, because we're testing tiny fixtures against a bug that only
  shows up at real scale/duration.
- **Found the repo was diverged from the platform again, independently.**
  This worktree's `main` was still stuck at version `0.11` — the same
  silent-revert bug documented in the 0.14 entry below had recurred: a
  real fix (bounding `Actor.push_data`/`Actor.charge` with
  `PUSH_TIMEOUT_S`, see below) had already been found, tested, and pushed
  live as build `0.14.1` on 2026-08-25 from branch
  `fix/kick-chat-archive-timeout-repro`, but that branch was never merged.
  Recovered it file-for-file into this branch before doing anything else
  (see `docs/specs/kick-chat-archive/notes.md`) — confirmed independently
  via `GET /v2/acts/DevilScrapes~kick-chat-archive` that `taggedBuilds.
  latest.buildNumber = "0.14.1"`, `finishedAt = 2026-08-25T10:41:45Z`.
- **Verdict: already fixed, 30-day window is stale, no code change
  needed.** Build `0.14.1` has been live for ~24h on a low-traffic Actor
  (~1 run/day); the 30-day window's 13 TIMED-OUT failures necessarily
  span the pre-0.14.1 builds (0.11.2, 0.10.1, ...) still aging out — the
  exact "stale window, not an active bug" pattern already documented on
  this Actor in the v0.7 and v0.8 entries below. Proved this two ways
  against the live `0.14.1` build via fresh cloud runs (not local —
  `POST /v2/acts/.../runs?build=latest`), both **SUCCEEDED**:
  - Small/QA-shaped: `channelSlugs=["xqc"]`, 30 s, cap 10 →
    run `ztYEvOlrt7q8DwpDK`, SUCCEEDED in 37.5 s, 1 row.
  - Large/real-sized (the actual asymmetry this brief asked to close):
    `channelSlugs=["n3on"]` — genuinely live at request time, ~18 200
    viewers — 300 s window, cap 3000 → run `4QUk9FISiqIcRPzSx`,
    SUCCEEDED in 309.1 s (i.e. ran the full requested window, no early
    kill), **1103 rows** archived via ~23 incremental `push_data` batches
    of 50, `statusMessage = "Archived 1103 chat message(s) across 1
    channel(s)."`. No TIMED-OUT, no truncation, no stall.
- No source change this session — `main.py`/`ws_client.py` already carry
  the fix (unbounded `Actor.push_data()`/`Actor.charge()` round-trip
  wrapped in `asyncio.wait_for(..., timeout=PUSH_TIMEOUT_S=60)`,
  `PushStalledError` → graceful partial-success exit instead of a
  platform force-kill). Local gates re-verified green on the recovered
  source: `ruff check` clean, `pyright` 0 errors, `pytest tests/` 63/63.
  Did not `apify push` — the live build is already byte-identical to
  this branch's `actors/kick-chat-archive` tree, and pushing an
  unchanged tree would only consume another build/version slot for no
  benefit (versions are already at the platform's 10-slot cap — see the
  "Discovered mid-push" note below).

### 0.14 — 2026-08-25

- **Recovered a silent regression before it could compound.** While
  preparing this push, `apify api GET .../versions` showed the platform
  already had versions `0.12` and `0.13` (pushed 2026-08-20, build
  `0.13.1`) containing a real fix — `ws_client.py::listen()` skipping the
  connect attempt entirely when the computed budget is already `<= 0`, and
  `main.py::compute_listen_budget_s` reserving `CONNECT_OVERHEAD_RESERVE_S`
  (`2 * PUSHER_CONNECT_TIMEOUT_S`) against the platform deadline so a
  small-but-insufficient budget doesn't spend the connect handshake's cost
  out of the flush/close margin. **Neither fix was ever committed to
  git** — they were pushed straight to the Apify platform from a checkout
  that was never merged into `main`. A day later (2026-08-21), this repo's
  own git history pushed version `0.11` (build `0.11.2`) from a checkout
  that predated both fixes, and because Apify tags `latest` by push
  recency rather than by comparing version numbers, that push silently
  **reverted** `0.12`/`0.13`'s protection back out of the `latest` build —
  invisibly, with no error, no warning, and no record in git. `0.11.2` is
  what customers have been running ever since, missing two real fixes.
  Both are reinstated in this release (see `docs/specs/kick-chat-archive/
  notes.md` for the full recovery trail). Version jumps `0.11` → `0.14`
  (skipping `0.12`/`0.13`, which already exist as version numbers on the
  platform) to guarantee `0.14` is unambiguously newer than everything
  ever pushed, closing the exact ordering gap that caused this revert.
- Daily signal: `publicActorRunStats30Days = {FAILED: 0, SUCCEEDED: 18,
  TIMED-OUT: 14, TOTAL: 32}` (56% success) on live build 0.11.2 — the 0.11
  connect/handshake crash fix helped (FAILED 21→0, TIMED-OUT share 75%→44%)
  but a large TIMED-OUT tail persisted. Three separate real-traffic cloud
  reproductions this session (offline channels/900s multi-channel, and a
  genuinely live 744-message channel over 600s) all completed cleanly and
  exited exactly on schedule, ruling out every previously-audited WS I/O
  site (connect, recv, send) all over again.
- **Root cause found:** `main.py::_push_batch()` called
  `Actor.push_data()`/`Actor.charge()` with no timeout of their own — the
  one write path in this Actor that was never bounded, even though every
  WS-layer I/O call in `ws_client.py` already is. `apify_client`'s own
  retry policy can legitimately take up to ~48 minutes under a
  platform-side stall (8 attempts x up to 360s each, its library default),
  which is unbounded relative to *any* customer's `maxDurationSeconds` and
  invisible to every deadline check `compute_listen_budget_s` performs,
  because it happens on the *consumer* side of the `listen()` generator —
  never inside the WS recv loop that `end_at` actually governs. A single
  stalled push held the whole run hostage until the platform force-killed
  it as TIMED-OUT, with nothing in the run's own logs to explain it and
  total loss of every row not yet flushed. This matches the observed
  `FAILED: 0` / `TIMED-OUT`-only failure shape exactly: the container is
  killed mid-`await`, before our own code ever gets a chance to raise or
  log anything.
- Fix: `main.py` — `_push_batch()` and `_charge()` now wrap
  `Actor.push_data()`/`Actor.charge()` in `asyncio.wait_for(...,
  timeout=PUSH_TIMEOUT_S)` (60s). A stalled push raises the new
  `PushStalledError`; `_stream_and_archive()` catches it, stops the listen
  loop immediately (no more accumulating an ever-larger unflushed batch
  behind a dead write), and returns a `push_stalled` flag. `main()` reports
  this as an honest partial success — "Archived N chat message(s) ...
  before a dataset write stalled — stopped early on purpose" — instead of
  letting the platform force-kill the run with nothing to show for it. A
  stalled `Actor.charge()` alone stays non-fatal, matching the existing PPE
  failure-handling contract.
- Also fixed a stale test assertion (`test_pay_per_event_declared`) left
  over from the 2026-08-20 fleet-wide start-fee bump (`actor-start`
  $0.002 → $0.20) that was failing the local suite on every commit since.

### 0.11 — 2026-08-17

- Daily signal: `publicActorRunStats30Days = {FAILED: 1, SUCCEEDED: 6,
  TIMED-OUT: 21, TOTAL: 28}` (78.6% failure) — the fleet-wide health sweep's
  \#3 worst Actor by lost customer runs (`ops/FLEET-HEALTH-TRIAGE-2026-08-17.md`).
  Build `0.10.1` (the dynamic-deadline fix) had already been live since
  2026-08-15 08:25, and the 2026-08-15/16 CEO reports show the failure rate
  did **not** clear afterward (08-16: "22/23 failures are HANGS, not
  errors") — direct evidence the 0.10 fix, while real, was not the whole
  story. Own runs (3 total, all SUCCEEDED) give no visibility into the
  failing customer runs (`reference-fleet-health-signal`), so this session
  is pure local reproduction against real Kick.com traffic, per
  `reference-fleet-fault-isolation-pattern`.
- **Root cause found — CONFIRMED by direct local reproduction, three
  variants:** `src/ws_client.py::listen()`'s connect + initial-handshake
  phase (`async with websockets.connect(...) as ws:` and the immediately
  following `ws.recv()` for Pusher's `connection_established` frame) had
  **no exception handling at all** — the one connection phase in this
  module that v0.5/v0.7/v0.9 never hardened, even though every other I/O
  site (send, recv-loop, subscribe) already has a bounded, graceful-failure
  contract. Reproduced with a local `websockets.serve()` mock Pusher server
  standing in for the real one:

  1. Peer accepts the connection then closes without sending anything →
     unhandled `websockets.ConnectionClosedOK`.
  2. Peer accepts the connection but never sends the first frame → unhandled
     `TimeoutError` (bounded by `PUSHER_CONNECT_TIMEOUT_S`, but not caught).
  3. Endpoint unreachable/refused (DNS failure, network blip, Pusher-side
     outage) → unhandled `OSError`.

  All three crash the **entire Actor run** with an uncaught exception
  instead of the graceful outcome the code clearly intends: `stats.
  connected` stays `False`, `main.py`'s existing REQ-13 check fires, and the
  run exits 1 with "Could not establish a live chat subscription" — a clear,
  billed-but-honest failure message instead of a raw traceback. Since
  `Actor.charge("actor-start", ...)` fires before any of this runs, every
  one of these was already a customer billed for nothing before the crash.
- **Honest scope — what this does and does not explain:** an uncaught
  exception inside `async with Actor:` produces a fast `FAILED` exit, not
  the `TIMED-OUT` (platform force-kill after exceeding the run's wall-clock
  budget) that dominates this Actor's current 30-day stats (21 of 22
  failures). This fix is **CONFIRMED** to close a real, previously-uncaught
  crash class with three concretely reproduced trigger conditions, and
  HYPOTHESIS (not proven) that it explains any share of the `TIMED-OUT`
  majority specifically — plausible if the Apify SDK's own exception-path
  shutdown inside `async with Actor:` doesn't itself return quickly, but
  that mechanism was not directly observed this session. `websockets`'
  connect-side WS traffic also carries no browser-like fingerprint (default
  `User-Agent: Python/websockets`), unlike every other network surface in
  this Actor (`api_client.py` rotates curl-cffi browser impersonation) and
  the fleet's anti-blocking stack — a plausible, not confirmed, reason a
  datacenter-IP handshake to Pusher gets rejected/dropped more often than
  a browser's would.
- **What was tried and did not explain the TIMED-OUT majority:** re-audited
  every wall-clock cap in `main.py`/`ws_client.py`
  (`RESOLVE_PHASE_MAX_S`, `RECV_CHUNK_S`, `SEND_TIMEOUT_S`,
  `PUSHER_CONNECT_TIMEOUT_S`, `DURATION_MAX_S`,
  `DEADLINE_SAFETY_MARGIN_S`) and the 0.10 `Actor.configuration.timeout_at`
  plumbing directly against the installed `apify==3.4.0` SDK source
  (`_configuration.py`) — the env-var aliases (`ACTOR_TIMEOUT_AT`/
  `APIFY_TIMEOUT_AT`) are real and resolve correctly when set, so the 0.10
  fix is not a silent no-op. No local reproduction attempt (see below)
  triggered an actual local hang past its expected bound.
- **What was tried and could not be reproduced locally:** live Kick.com
  traffic (real network calls against `kick.com`/Pusher, not mocked) —
  `xqc`, `kaicenat`, `adinross`, `ninja`, `trainwreckstv`, `nadia`,
  `stableronaldo`, `buddha`, `n3on`, `iceposeidon`, `fousey`, `agent00`,
  `duke_dennis`, `sketch`, `mikelong`, `elonmusk`, `amouranth`,
  `asmongold`, `ibai`, `loltyler1`, plus 404s (`some-definitely-fake-slug`,
  `sneakobrown`, `plaqueboymax`, `yeezuz`) to exercise REQ-1 slug-miss
  handling. None were live at request time (2026-08-17 ~11:00 UTC), so the
  "huge chat volume" and "non-English chat" scenarios in this session's
  brief could not be driven with real traffic; Kick's public API surface
  offers no reliable "currently live, high-viewer-count channel" listing
  endpoint to target one deterministically. A real `apify run` against
  `tests/fixtures/input.qa.json` (`xqc`, offline) confirms the happy/quiet
  path is unaffected by this fix: resolved <1s, connected, subscribed, ran
  the full 30s window, exited 0 with the expected message.
- Fix: `src/ws_client.py` — new `CONNECTION_ERRORS` tuple
  (`OSError`, `TimeoutError`, `websockets.WebSocketException`) wraps the
  entire `async with websockets.connect(...)` block in `listen()`; the
  initial-handshake logic is extracted to `_establish()` so `listen()`
  stays under the 40-line function ceiling. Any connection-establishment
  failure is logged and swallowed exactly like a same-effect "never
  connected" outcome, leaving `stats.connected=False` for the caller's
  existing fail-loud gate.
- New regression tests (`tests/test_ws_client.py`, 3 new, all verified to
  fail against the pre-fix source with the exact three exception types
  above, and pass after): `test_listen_survives_peer_closing_before_any_frame`,
  `test_listen_survives_peer_never_sending_first_frame`,
  `test_listen_survives_connection_refused` — each spins up a real local
  `websockets.serve()` mock Pusher server (not a fake object) so the
  regression exercises the actual `websockets` connect/handshake code path,
  not just a mocked substitute.
- Local gates: `ruff check src/ tests/` clean. `pyright` (via `uvx`, not
  added as a project dependency) — 1 pre-existing error only
  (`tests/test_models.py:63`, the documented Pydantic-`Field`-default false
  positive since 2026-07-10). `pytest tests/`: 57/58 green (54 pre-existing
  - 3 new); the 1 failure (`test_pay_per_event_declared`) is the same
    pre-existing, unrelated `$0.05` fleet-pricing-standard gap flagged in
    every prior session on this Actor — not fixed here (repricing has its own
    14-day-notice constraint).
- **Twitch sibling check (`actors/twitch-vod-chat-archive`, not modified):**
  does **not** share this flaw. It has no WebSocket/live-connection code at
  all — it archives already-recorded VOD chat via paginated Twitch GQL HTTP
  POSTs (`src/client.py::gql_post`), and that retry loop already wraps
  every network call in `try/except (RequestsError, OSError)` with
  backoff, returning `(None, reason)` instead of raising. The two Actors
  share a *name pattern* ("chat archive") but not an architecture — Kick
  has no chat-history API, so this Actor is necessarily live-only over a
  persistent WebSocket; Twitch's VOD replay is a fundamentally different,
  already-hardened, one-shot-HTTP-per-page design. No action needed on
  Twitch from this session's finding.
- **Budget note:** no cloud run was used this session (mid-session
  instruction withdrew the previously-allowed single small cloud run,
  account at $3.79/$5.00). This fix is local-reproduction-verified only;
  it has not been cloud-QA'd. The specific thing a cloud run would need to
  prove that local reproduction cannot: whether the `TIMED-OUT` majority
  itself clears once this build is live, since local runs never experience
  an actual Apify platform force-kill.

### 0.10 — 2026-08-13

- Daily signal (measured 2026-08-13, `GET /v2/acts/Sb8LSyiJ0UZfwdxad`):
  `publicActorRunStats30Days = {FAILED: 1, SUCCEEDED: 6, TIMED-OUT: 25,
  TOTAL: 32}` — 78% of customer runs TIMED-OUT. Compared against the prior
  six sessions' snapshots (see `docs/specs/kick-chat-archive/notes.md`),
  this is **not** simply the same stale window continuing to age out:
  `stats.lastRunStartedAt = 2026-08-11T14:45:16Z`, i.e. a run (not ours —
  our own `/acts/{id}/runs` shows only 3 owner runs, none since 2026-08-06)
  landed *after* build 0.8.1 (finished 2026-08-06, the live build for this
  whole window) went live. The "aging legacy timeouts, no new failures"
  verdict from the five prior sessions (2026-07-10 through 2026-08-12)
  cannot be extended to cover that run without more direct evidence, and
  this Actor's near-zero user count (`totalUsers30Days: 1`) means a single
  customer's usage pattern dominates the whole stat — worth re-auditing for
  a class of hang none of the prior five sessions checked for.
- **Root cause found:** every prior TIMED-OUT fix on this Actor (v0.4-v0.9)
  bounded the WebSocket listen window against a *static* assumption of the
  Apify platform's kill timer — `DURATION_MAX_S=3540s` vs. a hardcoded
  belief that `defaultRunOptions.timeoutSecs=3660s` always applies. That
  assumption breaks whenever the *actual* run is started with a tighter
  platform timeout than the Actor's own default — routine for schedulers,
  webhooks, and third-party no-code integrations (Zapier/Make/n8n-style
  connectors commonly default an actor-run timeout far below an hour). This
  Actor's normal use case (archive a whole stream) means most successful
  runs already try to use the *entire* requested `maxDurationSeconds`
  window — the single highest-exposure design in the fleet for exactly this
  class of caller-timeout mismatch, and invisible to every prior fix
  because none of them ever read the platform's *actual* per-run deadline.
  Apify exposes it via the `APIFY_TIMEOUT_AT`/`ACTOR_TIMEOUT_AT` env var,
  surfaced by the SDK as `Actor.configuration.timeout_at` — unused by this
  Actor's code until now.
- Fix: `src/main.py::_seconds_until_deadline()` reads
  `Actor.configuration.timeout_at` and returns the real seconds remaining
  until the platform force-kills the run (`None` on local runs — no env var
  set, so local behaviour is unchanged). `compute_listen_budget_s()` gained
  a third, independent ceiling — `deadline_remaining_s` minus a new
  `DEADLINE_SAFETY_MARGIN_S=30s` flush buffer — so the listen window can
  never outlive the *actual* platform deadline, regardless of what
  `maxDurationSeconds` was requested or what this Actor's own
  `defaultRunOptions.timeoutSecs` default is. A run whose caller supplied a
  tight timeout now exits cleanly via the Actor's own bounded exit path
  well before the platform force-kills it — trading a small amount of
  listen time for a guaranteed `SUCCEEDED` (or the existing REQ-13
  fail-loud path) instead of a `TIMED-OUT`.
- New regression tests (`tests/test_main.py`, 6 new):
  `test_listen_budget_clips_to_platform_deadline_when_shorter_than_
  requested_window`, `test_listen_budget_deadline_never_negative_even_
  when_already_overdue`, `test_listen_budget_ignores_deadline_when_none_
  local_run`, `test_listen_budget_prefers_requested_window_when_deadline_
  is_generous`, `test_seconds_until_deadline_none_when_no_platform_
  timeout`, `test_seconds_until_deadline_computes_remaining_seconds` — all
  six verified to fail against the pre-fix source (missing
  `_seconds_until_deadline` / `deadline_remaining_s` kwarg entirely) before
  the fix, and pass after.
- **Honest scope caveat:** this closes a real, previously-unaudited gap and
  is architecturally the best explanation found for a TIMED-OUT that
  survives every static wall-clock cap already in place — but the account's
  monthly usage hard limit (resets 2026-08-15) blocks any cloud run this
  session, so it is **not proven against a live caller-timeout-override
  run**. Local `apify run` against `tests/fixtures/input.qa.json` confirms
  the happy path (no platform timeout env var set locally) is unaffected:
  resolved in <1s, connected, subscribed, ran the full 30s window, exited 0
  with the expected "captured 0 messages" status. Needs a cloud QA run
  once the cap resets, ideally one that explicitly passes a short `timeout`
  query param on `POST /v2/acts/.../runs` alongside a much larger
  `maxDurationSeconds` in the body, to directly exercise the clipped path
  end-to-end (not just unit-tested).
- Unrelated, pre-existing finding (not touched — out of scope for a
  TIMED-OUT fix): `tests/test_input_schema.py::test_pay_per_event_declared`
  fails on `main` too (`actor-start` priced `$0.002` in
  `.actor/pay_per_event.json` vs. the fleet's `$0.05` flat-fee standard the
  test expects) — looks like this Actor was missed by the recent
  fleet-wide `$0.05` start-fee repricing pass. Flagging for a separate,
  dedicated pricing fix (repricing has its own 14-day-notice constraint —
  not safe to bundle into this change).

### 0.9 — 2026-08-12

- Daily report flagged 17% 30-day success rate (28/34 runs failed, mostly
  `TIMED-OUT`). Could not read the failing customer runs directly (own-run
  API scope only ever shows this Actor's own 3 runs — none TIMED-OUT — the
  documented pattern: `recent_failures` was empty). Cloud runs/pushes were
  also blocked today by the account's monthly usage hard limit, so this
  fix is diagnosed from live log-reading via `GET`-only REST calls plus
  local reproduction, not a fresh cloud run.
- Per the fleet fault-isolation pattern (`reference-fleet-fault-isolation-
  pattern` — the #1 cause of low-success actors fleet-wide, and the exact
  bug class v0.8 already partially closed for this Actor), re-audited
  `src/parser.py::build_row` and `src/ws_client.py::_route_frame` for any
  remaining "one bad message crashes the whole run" gap and found two,
  both reproduced locally before fixing:
  1. **`build_row` — non-dict `sender` / `sender.identity`.** `sender =
     inner.get("sender") or {}` and the subsequent `_badge_types(sender)`
     call ran *before* the function's own try/except, so a payload where
     Kick sends `sender` (or `sender.identity`) as something other than an
     object (e.g. a bare string) raised an uncaught `AttributeError` —
     directly contradicting `build_row`'s documented "never raises"
     contract, and would have crashed every subscribed channel's stream on
     one malformed message. Fix: new `_as_dict()` helper coerces `sender`
     and `identity` to `{}` on any non-dict shape *before* they're used,
     eliminating the crash instead of merely catching it after the fact;
     `AttributeError` also added to the existing defense-in-depth
     `except` clause around `ResultRow(...)` construction.
  2. **`_route_frame` — non-string `data` field.** `json.loads(frame.get
     ("data") or "{}")` only caught `json.JSONDecodeError`. Pusher's
     protocol normally encodes `data` as a JSON string (documented at the
     top of `ws_client.py`), but a frame where `data` arrives already
     decoded (a dict) raises `TypeError` on `json.loads`, uncaught,
     crashing the whole run identically. Also hardened the mirror case —
     a `data` string that decodes to a non-object (e.g. a JSON array) —
     which would otherwise pass a non-dict `inner` payload downstream into
     `build_row`, itself an `AttributeError` risk. Both cases now log a
     warning and skip the one frame; `chat_frames_seen` still increments
     so the existing "chat arrived but nothing parsed" fail-loud path
     (v0.8) stays correct.
- Customer-visible effect: a single malformed/reshaped chat message from
  Kick no longer has any code path that can crash the entire run for every
  subscribed channel — it's logged and skipped, matching every other
  malformed-payload path already hardened in this Actor (v0.8's
  non-numeric-`chatroom_id` fix, this release's two siblings).
- New regression tests: `test_build_row_survives_non_dict_sender_in_payload`,
  `test_build_row_survives_non_dict_identity_in_payload` (`test_parser.py`);
  `test_route_frame_survives_non_string_data_field`,
  `test_route_frame_survives_data_decoding_to_non_object`
  (`test_ws_client.py`) — all four verified to fail against the pre-fix
  source and pass against the fix.
- **Scope note:** this closes a genuine, reproducible fault-isolation gap,
  but the dominant failure status in the 30-day window is `TIMED-OUT`
  (platform-forced kill), not `FAILED` — an uncaught exception like the
  ones fixed here produces a fast `FAILED` exit, not a `TIMED-OUT` hang.
  This fix is therefore very likely a real but partial contributor; the
  `TIMED-OUT` majority may still trace to a cause outside this Actor's own
  wall-clock budgeting (all internal caps — `RESOLVE_PHASE_MAX_S`,
  `RECV_CHUNK_S`, `SEND_TIMEOUT_S`, `DURATION_MAX_S` — were re-verified
  correct and bound total wall-clock to ≤3540s against a 3660s platform
  timeout). See CHANGELOG entry discipline in `docs/specs/kick-chat-
  archive/notes.md` for the full trail of prior TIMED-OUT investigations
  on this Actor (v0.4–v0.8) — none reproduced a live bug from this
  session's evidence gathering. Flagged for cloud QA + a fresh customer-
  facing check once the account's usage hard limit resets (2026-08-15).

### 0.8 — 2026-08-06

- Daily report re-flagged 16% 30-day success rate (29/37 TIMED-OUT, 37
  total). Live-verified this is the **same stale 30-day-window snapshot**
  already diagnosed on 2026-08-05: `publicActorRunStats30Days` is
  byte-identical to the pre-v0.7-push reading (`TIMED-OUT: 29, SUCCEEDED:
  6, FAILED: 1, ABORTED: 1, TOTAL: 37`), the live build is already `0.7.1`
  (finished `2026-08-05T07:54:31Z`, contains the send-timeout hang fix),
  and our own last run (`RH4CPSPeJop6wzq8H`) SUCCEEDED that same day. Zero
  customer runs have landed in the ~24h since — traffic on this Actor is
  very low (`users_30d` historically 1-2) — so the trailing window hasn't
  moved at all. Not a new TIMED-OUT bug; the wall-clock-budget class (REQ:
  resolve-phase cap, listen-phase cap below platform `timeoutSecs`,
  recv/send chunk timeouts) was already fully addressed across v0.4-v0.7
  and re-confirmed live and correct today (`defaultRunOptions.timeoutSecs = 3660s` on the platform vs. `DURATION_MAX_S = 3540s` in code).
- Per this task's explicit brief, independently applied the two other
  fleet fault-isolation requirements regardless of the TIMED-OUT verdict:
  1. **Fault isolation** (`src/parser.py::build_row`) — found and fixed a
     real, previously-uncaught bug: a malformed (non-numeric)
     `chatroom_id` field inside one incoming `ChatMessageEvent` payload
     raised an unhandled `ValueError` (from `int(...)`) that propagated
     out of `build_row`, through `_stream_and_archive`'s async generator,
     and would have crashed the ENTIRE run for every subscribed channel —
     not just skipped the one bad message. Reproduced locally before
     fixing. Fix: `_int_or_default()` coerces safely (falls back to the
     already-known chatroom\_id), and the whole `ResultRow(...)`
     construction is wrapped in `try/except (ValidationError, ValueError,
     TypeError)` as defense-in-depth, matching `build_row`'s own
     documented "skip, don't crash" contract for malformed payloads.
  2. **Never a silent empty success** — added `ListenStats.chat_frames_seen`
     (`src/ws_client.py`), incremented whenever a `ChatMessageEvent` frame
     for a chatroom we subscribed to is identified (regardless of parse
     outcome). `main.py` now distinguishes a genuinely quiet/offline
     channel (`chat_frames_seen == 0` — REQ-13's existing valid-empty-
     success path, unchanged) from chat activity that arrived but 100%
     failed to parse (`chat_frames_seen > 0, kept == 0` — most likely
     Kick changed its payload shape): the new case dumps up to 5 raw
     failed payloads to the KVS (`RAW_PAYLOAD_SAMPLES` key) via
     `Actor.set_value` and fails loud (`SystemExit(1)`) with a clear
     status message instead of silently reporting "0 messages captured."
- New regression tests: `test_build_row_survives_non_numeric_chatroom_id_in_payload`,
  `test_build_row_catches_unexpected_row_construction_failure` (`test_parser.py`);
  `test_route_frame_increments_chat_frames_seen_*` x3 (`test_ws_client.py`);
  `test_should_fail_on_silent_parse_failure_*` x3,
  `test_stream_and_archive_captures_raw_samples_on_parse_failure`,
  `test_stream_and_archive_caps_raw_samples_at_max`,
  `test_fail_on_silent_parse_failure_dumps_kvs_and_exits` (`test_main.py`).

### 0.7 — 2026-08-05

- Daily report flagged 16% 30-day success rate (29/37 TIMED-OUT). Triage
  confirmed the TIMED-OUT count (29) has been flat across three checks
  spanning 18 days (2026-07-18, 2026-07-20, 2026-08-05) while total runs
  and successes both grew — the v0.6 (2026-07-10) resolve-phase fix is
  holding; every run since is a legacy pre-fix TIMED-OUT sitting in the
  30-day trailing window, expected to fully age out ~2026-08-09. Not a
  new active bug. See `docs/specs/kick-chat-archive/notes.md`.
- Independent code audit for *any remaining* unbounded-hang vector (the
  brief for this fix asked for one) found a real, if narrower, gap:
  `ws_client.py`'s `ws.send()` calls (subscribe frames in
  `_send_subscriptions`, the ping->pong reply in `_handle_control_frame`)
  had no timeout at all — the mirror-image of the `RECV_CHUNK_S` hang
  fixed in v0.5, just on the write path. A congested proxy or a peer
  that stops reading (no RST/FIN) would leave any of these `await
  ws.send(...)` calls blocked forever, invisible to the existing
  `RECV_CHUNK_S` / `RESOLVE_PHASE_MAX_S` / `DURATION_MAX_S` caps because
  none of them bound writes.
- Fix: `_send_with_timeout()` wraps every `ws.send()` in
  `asyncio.wait_for(..., timeout=SEND_TIMEOUT_S=10.0)` and raises
  `ConnectionStalledError` on a stall. `_send_subscriptions` stops at the
  first stalled chatroom rather than retrying the remaining ones against
  a dead pipe (bounds total subscribe-phase stall cost to one
  `SEND_TIMEOUT_S`, not `channel_count * SEND_TIMEOUT_S`). The listen
  loop (`listen()`, `_iter_messages()`) treats `ConnectionStalledError`
  the same as `websockets.ConnectionClosed` — log + clean early return —
  so a stalled write now produces a fast, bounded exit (existing REQ-13
  fail-loud path fires when zero subscriptions succeeded) instead of an
  unbounded hang.
- Added `tests/test_ws_client.py` (6 tests) — locks in the stall-detection
  behavior on both the subscribe path and the ping/pong path, plus the
  happy-path pass-through.

### 0.6 — 2026-07-10

- Fix the actual TIMED-OUT root cause behind the "Under maintenance"
  flag (30-day customer success rate 9%, 29/33 runs TIMED-OUT). Build
  0.5.1's `RECV_CHUNK_S` fix addressed a *different* hang vector inside
  the WebSocket listen loop and was already live, but customer runs
  kept TIMED-OUT — because the real bug is upstream of the listen loop
  entirely.
- Root cause: `_resolve_all` (REST slug -> chatroom\_id lookup, up to 20
  channels x 4 retries x 30s HTTP timeout + exponential backoff each in
  `api_client.py`) had no wall-clock budget of its own, and the time it
  took was never subtracted from the WebSocket listen window
  (`max_duration_seconds`). A customer with several slow/rate-limited
  slugs (routine on Apify's shared proxy IPs hitting Kick's
  Cloudflare-protected REST API) could burn most or all of the run's
  time budget just resolving channels — *before* listening even
  started — pushing total wall-clock past the platform's
  `timeoutSecs=3660` buffer and getting killed as TIMED-OUT instead of
  exiting cleanly via the Actor's own duration cap. Single-slug,
  low-latency local QA (the existing `tests/fixtures/input.qa.json`)
  resolves in well under a second and never exercised this path.
- Fix: cap the resolve phase at `RESOLVE_PHASE_MAX_S=90s` wall-clock
  (`src/main.py::_resolve_all`, `_resolve_into`) — proceeds with
  whatever channels resolved in time rather than blocking indefinitely
  — and deduct the resolve phase's actual elapsed time from the
  WebSocket listen budget (`compute_listen_budget_s`) so total Actor
  wall-clock (resolve + listen) never exceeds the customer's own
  `maxDurationSeconds`, regardless of how long resolution took.
- README: corrected `maxDurationSeconds` doc from "5–3,600 seconds" to
  the actual enforced range "5–3,540 seconds (59 min)" — the stated
  ceiling didn't match `DURATION_MAX_S` in `src/models.py`, which may
  have nudged customers toward duration values close to the platform
  timeout with no headroom.

### 0.5 — 2026-06-13

- Add `RECV_CHUNK_S=45` cap on each `ws.recv()` call in `_iter_messages`.
  Previously the recv timeout was set to the full remaining window duration
  (up to 3540 s). In cloud environments where NAT/firewall infrastructure
  silently drops idle TCP connections (no RST/FIN), the recv would hang
  until the window expired rather than detecting the dead connection quickly.
  With the 45-second chunk cap, a stale TCP connection is detected within
  one chunk window; a chunk timeout on a still-live connection is benign
  (the loop continues for another chunk). This closes a residual hang vector
  on channels with no chat activity.

### 0.4 — 2026-06-09

- Fix 10/13 TIMED-OUT customer runs. Root cause: `maxDurationSeconds` cap
  (3600 s) matched the Apify platform `timeoutSecs` exactly, so the platform
  timer raced the Actor's listen window during container startup (~5-10 s
  overhead) and won → the run reported TIMED-OUT instead of SUCCEEDED.
  Fix: reduce `DURATION_MAX_S` to 3540 s (59 min) and set the platform
  `defaultRunOptions.timeoutSecs` to 3660 s (60 s buffer over the new cap).
- Replace `asyncio.get_event_loop()` with `asyncio.get_running_loop()` in
  `ws_client.py` (Python 3.10+ best practice; `get_event_loop()` is
  deprecated when called outside a running event loop context).
- Add `ping_interval=None` to `websockets.connect()` to disable library-level
  WebSocket pings. Pusher drives its own heartbeat via `pusher:ping` /
  `pusher:pong` at the application layer; the library's concurrent ping
  occasionally interfered in cloud environments causing spurious
  `ConnectionClosed` during quiet channels.

### 0.3 — 2026-06-01

- Fix "Under maintenance" flag. A run that connected to Kick's chat
  socket and subscribed successfully but saw zero messages (every
  channel offline/silent for the window) used to exit 1 — so Apify's
  automated daily QA, which prefills a single channel that is often
  offline, failed three days running and unlisted the Actor.
- A clean connect + subscribe with zero messages is now a successful
  empty run (exit 0) with an explanatory status message. Fail-loud
  (exit 1) is reserved for a genuinely broken integration: the socket
  never connects, or no channel accepts the subscription
  (`ListenStats.connected` / `subscriptions_succeeded`).

### 0.1.0 — 2026-05-16

- Initial release. Subscribes to Kick.com chat via the public Pusher
  WebSocket, archives one row per chat message until the duration or
  message-count cap fires. Real-time only — Kick exposes no historical
  chat retrieval.
