# Changelog of YouTube Transcript Scraper - Transcriber (`prodiger/youtube-transcript-scraper---transcriber`) Actor

- **URL**: https://apify.com/prodiger/youtube-transcript-scraper---transcriber/changelog.md
- **Full Actor documentation**: https://apify.com/prodiger/youtube-transcript-scraper---transcriber.md

## Changelog

All notable changes to this project are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).

### \[0.5.7] — 2026-04-27

#### Fixed

- **Prefilled `videoUrls` shrunk from 3 URLs to 1 short video** to fit Apify's 5-minute automated-test budget. The previous prefill (`[Me at the zoo, @example channel, 200-video playlist]`) timed out auto-tests in ~10+ minutes, accumulating 8 TIMED-OUT runs in the trailing 30-day window and re-triggering the "Under Maintenance" auto-flag despite v0.5.6's success-rate fix. Channel and playlist URL examples now live in the field's description text instead.

### \[0.5.6] — 2026-04-26

#### Fixed

- **`Actor.fail()` on user-input errors → `Actor.exit()` with skip record.** Both rejection paths (`assertValidInput` throw + post-expansion empty queue) used to call `Actor.fail()`, which counts as a `FAILED` run on Apify's success-rate metric. With marketplace-vetting / curiosity probes routinely calling the actor with junk like `videoUrls: ["https://x"]`, the actor accumulated 5 FAILED runs out of 10 trailing → Apify auto-flagged the actor as "Under Maintenance". v0.5.6 reframes both as user errors: log a clear message, push a single skip record (so the caller still sees what went wrong in the dataset), set the run's `statusMessage`, and exit with status `SUCCEEDED`. Actor-side errors (Whisper/yt-dlp/proxy failures) still count as failures as before.

#### Operational

- After deploy, manually clear the "Under Maintenance" flag in the Apify Console (Publication tab). Once trailing 30-day success rate recovers, the auto-flag won't re-fire from input-rejection runs.

### \[0.5.2] — 2026-04-20

#### Reverted

- **Audio-client swap from \[0.5].** `downloadAudio` returned to `buildBaseArgs` (shared `default` client + PoToken). The 0.5 change to `player_client=android_vr,tv,tv_downgraded` was based on a misread of v0.4 stderr (we lacked stderr logging then). The new stderr-tail logging added in 0.5 revealed the actual failure signature is `Sign in to confirm you're not a bot` — the bot challenge — which is precisely what PoToken defeats. The TV and android\_vr clients don't use PoToken, so the swap traded SABR-for-web → bot-challenge-for-tv on the same videos. Net audio success rate was unchanged.

#### Kept from \[0.5]

- **Long-audio support via ffmpeg chunking** (unchanged). Files above Whisper's 25 MB limit are split, transcribed chunk-by-chunk, and re-merged. `maxDurationMinutes` can safely run up to 300.
- **Audio-download failure log includes stderr tail** (unchanged). This is how we caught the real root cause.

#### Known limitations (carried from \[0.4])

- **Bot-challenge on some audio downloads remains.** PoToken mitigates most cases on the `default` client but not all — some videos on flagged IPs still return the `Sign in to confirm you're not a bot` page despite PoToken being present. This is a yt-dlp / PoToken upstream issue, not something the actor can fix without cookies. Captions and metadata paths are unaffected. Mitigation: captions-first routing means only videos without English captions in the caller's language ever hit this path.

### \[0.5] — 2026-04-20 \[withdrawn — see 0.5.2]

#### Added

- **Long-audio support via ffmpeg chunking.** Audio files above Whisper's 25 MB per-request limit are split with ffmpeg's segment muxer (`-c copy`, no re-encoding), transcribed chunk-by-chunk, then merged with correct timestamp offsets. Raw-audio download ceiling raised from 24 MB to 200 MB (`MAX_DOWNLOAD_BYTES`). `maxDurationMinutes` can now safely run up to its 300-minute (5 h) maximum — the `audio-exceeds-whisper-limit` skip now fires only for truly huge files (> 200 MB, e.g., archived multi-hour livestreams).
- **New `src/whisper/chunker.ts`** — byte-bounded splitter with validation (every chunk verified under the cap before Whisper is called) and best-effort temp-dir cleanup.

#### Changed

- **SABR audio-download bypass \[reverted in 0.5.2].** `downloadAudio` used `--extractor-args youtube:player_client=android_vr,tv,tv_downgraded` instead of the shared `default` client. Cloud smoke testing revealed the theory didn't hold — see 0.5.2.
- **Audio-download failure log now includes the last 500 chars of yt-dlp stderr** so format-selection failures surface cleanly without needing debug logging.

#### Claim retracted in 0.5.2

- 0.5 claimed it "closed the known limitation noted in \[0.4]" for SABR-flagged videos. Cloud smoke testing on `@starterstory/videos` cap=3 showed identical 2/3 success as v0.4, with the failing video's error message changing (SABR → bot challenge) but the outcome unchanged. 0.5.2 reverts the faulty half of 0.5 and carries forward the honestly beneficial half (chunking).

### \[0.4] — 2026-04-20

#### Added

- **PO Token support via `bgutil-ytdlp-pot-provider` plugin** (v1.3.1, actively maintained). Installed in the Dockerfile and run as a background Node.js server on `127.0.0.1:4416`. yt-dlp's plugin system auto-discovers the server and mints a fresh PO Token per video request. This is the actually-correct 2026 bypass for YouTube's "Sign in to confirm you're not a bot" challenge.
- New `src/youtube/potoken-server.ts` module manages the server lifecycle: startup before first yt-dlp call with 15s readiness poll, clean shutdown in the `aborting` handler and at normal exit.

#### Changed

- `--extractor-args youtube:player_client=default,tv,mweb` → `youtube:player_client=web`. The earlier multi-client fallback caused false `DRM protected` errors from the TV client and random 360p-only responses from `android_vr`. With PoToken support, the web client is the supported path.

#### Known limitations

- **SABR streaming enforcement remains unaddressed**: on flagged proxy IPs, YouTube sometimes returns only storyboard images or 360p video instead of downloadable audio streams. yt-dlp has no SABR implementation (upstream issue #16082). Captions and metadata operations are unaffected; only the Whisper fallback path (audio download) is exposed to SABR. Mitigation: captions-first routing means only videos without English captions hit this path.

### \[0.3] — 2026-04-20

#### Added

- **Channel and playlist URLs accepted in `videoUrls`.** `/@handle`, `/channel/UC…`, `/c/<name>`, `/user/<name>`, and `/playlist?list=…` are all recognized. Channel URLs expand the `/videos` tab only (Shorts and Streams deferred). Playlists return videos in playlist-creator-defined order (NOT necessarily newest-first) — documented prominently.
- **`maxVideosPerChannel` input field** (default 200, max 5000, `0` = unlimited). Caps per-source enumeration.
- **`channel-summary` dataset record** emitted once per channel or playlist source with `discovered`/`returned` counts. Lets users detect silent truncation on large channels. Fires no PPE charge.
- **Specific skip reasons for enumeration failures**: `channel-not-found`, `channel-private`, `channel-geo-blocked`, `channel-rate-limited`, `channel-unavailable` fallback, `channel-empty` (plus `playlist-*` mirrors). Parsed from yt-dlp stderr with first-match-wins order.
- **`sourceUrl` and `sourceKind` fields** on dataset records — populated on enumeration-failure skip records so users can tell which channel failed.
- **`type` discriminator field** on dataset records: `'transcript' | 'channel-summary'`.
- **First test files** for the actor: canonical extractors, yt-dlp stderr mapping, parse, and format builders. 73 test scenarios total.

#### Changed

- **Default `maxWhisperMinutesPerRun` bumped from 60 → 240** (4 hours of audio). Sized for bulk channel runs; caps OpenAI cost at ~$1.44/run.
- **Default actor timeout bumped from 1h → 4h** via `defaultRunOptions.timeoutSecs: 14400` in `actor.json`. A default 200-video channel run with mixed captions/Whisper needs more than 1h.
- **Records now written to the run's default dataset** instead of the account-level named `transcripts` dataset. Fixes the same bug pattern that was fixed in reddit-scraper v0.2 (Apify Console UI and standard dataset readers were seeing 0 items despite successful runs).
- **`buildSkipRecord` videoUrl is empty (not `https://www.youtube.com/watch?v=`)** when no videoId is provided — fixes a malformed-URL bug that only manifested for enumeration-failure records introduced in this version.

#### Security

- **`--no-config` flag added to every yt-dlp invocation** — prevents config-file side-loading RCE class (an attacker who could plant `~/.config/yt-dlp/config` would otherwise inject `--exec` into every subprocess call).
- **`--` argument terminator before every URL positional** — defends against yt-dlp reparsing URL content as flags.
- **Per-kind character allowlists for channel/playlist extractors** mirror the existing `VIDEO_ID_RE` precedent. Rejects percent-encoded slashes, unicode handles (v0.3 scope), credentials-in-authority, and path-traversal-shaped inputs.
- **Newline-stripping on the watch+list `log.info` URL** to prevent log injection.
- **yt-dlp stderr truncated to 500 chars** before storage; not emitted into dataset records (only to `log.warning`).

### \[0.2] — 2026-04-18

#### Changed

- `openaiApiKey` is now **optional**. Captions-only workflows no longer require a key. With `transcriptMethod=auto` and no key, videos without captions in the requested language are skipped with reason `no-openai-key-no-fallback` instead of failing the run. Required only when `transcriptMethod=whisper`.

### \[0.1] — 2026-04-19

#### Added

- Initial release. YouTube transcriber covering single video URLs.
- Captions-first transcription via yt-dlp + json3 caption parsing.
- OpenAI Whisper API fallback for videos without captions in the requested language. BYOK (bring your own OpenAI API key).
- Two output formats: `text` (with optional timestamps), `json` (segments array).
- yt-dlp + ffmpeg pinned in Docker image (yt-dlp 2026.3.17). Audio path validated at 87% real-video success rate during pre-build spike.
- RESIDENTIAL proxy default (YouTube blocks datacenter IPs in 2026).
- `maxDurationMinutes` cap (default 18) to keep audio under Whisper's 25 MB limit at typical bitrates.
- `maxWhisperMinutesPerRun` cap (default 60) to bound user OpenAI bill per run.
- Pay-per-event pricing matching the YouTube transcriber market: `transcript-captions` $0.0005 (matches `supreme_coder/youtube-transcript-scraper`), `transcript-whisper` $0.05 (covers proxy + compute COGS, still ~5-6× cheaper than codepoetry's bundled $0.012/min). Plus standard `apify-actor-start` $0.003.
- Strict SSRF defense via hostname allowlist + canonical URL reconstruction before any subprocess invocation.
- OpenAI key sanitization in error logs (`sk-*` masking).
- Audio temp-file lifecycle (mkdtemp → try/finally cleanup → abort handler best-effort sweep).
- Per-run summary log with per-reason skip counts and estimated user OpenAI cost.
- Skip records emitted to dataset (transcript empty, skipReason filled) so users have in-product visibility into per-video outcomes.
