# Changelog of Youtube Channel Scraper (`newbs/youtube-channel`) Actor

- **URL**: https://apify.com/newbs/youtube-channel/changelog.md
- **Full Actor documentation**: https://apify.com/newbs/youtube-channel.md

### 2026-10-03 - Overlapping processing stages

- Feed discovered videos into bounded detail and recovery queues during collection and pagination.
- Start eligible recovery without waiting for every initial detail job; preserve the run-wide caption coverage threshold, retry cap, and shared proxy safeguards.
- Save finalized rows immediately and release saved row payloads from memory when streaming without debug diagnostics.
- Add queue occupancy and overlapping stage activity to OUTPUT, plus regressions for backpressure, early delivery, coverage decisions, deduplication, and fatal write handling.
- Keep current event prices and one charge per successfully saved video.

### 2026-10-03 - Detail and caption latency

- Prioritize caption/player/comment requests ahead of queued watch-page downloads so in-progress proxy results can finish.
- Stream completed adaptive retries immediately and avoid retrying exhausted video budgets.
- Recover captions through lightweight player requests before downloading another proxy watch page; retain the watch-page fallback when needed.
- Try alternative player caption sources before exhausting many empty watch-page language tracks; stop repeating an HTTP-blocked player endpoint on the same route.
- Reuse collected channel profiles, overlap independent comments/profile work, cache duplicate caption attempts, and cancel redundant alternatives after a usable transcript is found.
- Default adaptive caption recovery to eight video jobs. Normal detail extraction stays at 20 workers.
- Add first-result and per-phase timings to OUTPUT. Existing prices, result charging rules, and per-video/shared spending safeguards remain in place.

### 2026-10-03 - Concurrency tuning

- Restore 20 simultaneous video-detail jobs by default at 512 MB and 1,024 MB after hosted comparisons with 5 and 10 workers.
- Respect explicitly lower concurrency settings, retaining the shared two-request proxy limit, recovery budgets, and failure diagnostics.
- Replace the earlier estimated memory-based detail cap with the measured maximum of 20; Basic mode is unaffected.

### 2026-10-03 - Result quality and runtime safeguards

- Exclude advertising links and sponsored blocks before counting or exporting results.
- Add basic/full collection mode, details status, a Basic results view, and sanitized row error fields without removing existing fields.
- Limit proxy work across the run using successful-result allowances, transfer cancellation, request concurrency, and repeated-failure detection.
- Cap detail concurrency for the available memory, stop new work after fatal worker errors, and retain completed results.
- Persist progress and failure diagnostics in OUTPUT; log sanitized stacks and launch Node directly.
- Keep direct and proxy event prices and one-charge-per-successful-video behavior unchanged.

### 2026-09-16 - Channel reliability fixes

- Resolve UC channel IDs directly and require exact identity matches for search recovery.
- Normalize channel tabs and video URLs; reject unsupported hosts and malformed inputs.
- Recover transport errors and HTTP-200 challenge responses, keeping proxy transport for pagination.
- Preserve collected videos on pagination failure, honor zero-page and transcript retry limits, and skip About fetches in list-only mode.
- Save source diagnostics to OUTPUT and propagate genuine failures through the Actor lifecycle.
- Add deterministic regressions and six bounded live Actor smoke checks.

## Changelog

### 2026-09-15

- Prepared two exclusive result prices: $0.50 per 1,000 direct video results and $8 per 1,000 proxy-assisted video results, including run platform usage. The Pricing tab determines when the new rates become effective.

- Added per-video proxy-use accounting and mixed-price spending-limit checks. Each successful video is charged once; failed and partial rows remain free.

- Bounded per-video work across retries, reduced duplicate caption-format downloads, and lowered default memory to 512 MB with a 1,024 MB maximum.

- Prepared `video-result` pay-per-event billing for successful videos, preserving existing rental behavior until the pricing migration takes effect.

- Kept partial results, failed rows, and Shorts-only guidance free, with available captions included in each successful video result.

- Serialized result writes and charges, prevented duplicate video billing, and added graceful stopping at the run spending limit.

- Added SDK-backed billing tests covering failed writes, concurrent workers, batch limits, and rental compatibility.

### 2026-08-09

- Removed source diagnostic rows from the default dataset so exports contain video records only.
- Removed `resultType`, `message`, and `sourceStatusCode` from video results and dataset views.
- Kept Shorts-only and source failure explanations in run logs and the final run status.
- Added one actionable `shorts_only` guidance row when a channel contains Shorts but no normal videos, including a direct link to the separate YouTube Shorts Scraper.
- Normalized Markdown-wrapped channel links such as `[URL](URL)` before collection.

### 2026-06-11

- Added visible `source_status` dataset rows for Shorts-only and failed channel sources, including a readable message and machine-readable status code.
- Added `resultType: "video"` to normal video rows so integrations can separate videos from source diagnostics.
- Added explicit diagnostics for Shorts-only channels so a successful empty run no longer looks like an unexplained scraper failure. The actor still excludes Shorts by design.
- Empty HTTP 200 channel pages now retry through the canonical channel ID before being classified as having no normal videos.
- Added proxy-backed source collection retry for `proxyMode: "auto"` when YouTube throttles channel or search listing pages with statuses such as 429.
- Added a channel search resolver for broken YouTube handle/vanity redirects, so URLs like `@mktv` can recover through the stable channel ID URL instead of returning zero rows when YouTube redirects the handle to a dead custom path.
- Raised the default HTTP video-detail concurrency for hosted runs and added `adaptiveTranscriptRetryConcurrency` so transcript-heavy multi-channel runs can finish faster.
- Batched final dataset writes for rows held during transcript recovery, reducing the chance that a short external run timeout leaves users with only early no-transcript rows.
- Added `videoDetailProxyStrategy` so hosted runs can test proxy-first video detail scraping without proxying channel/search collection.
- Added `transcriptMode` with standard and aggressive modes, while preserving the older residential transcript fallback option for compatibility.
- Added adaptive transcript proxy recovery in standard mode so batches with poor direct transcript coverage can retry missing captions without making every successful run pay the aggressive fallback cost.
- Bounded adaptive transcript recovery on larger runs with a round-robin retry sample across channels, preventing 100+ result runs from spending the whole timeout retrying every missing transcript while still retrying up to 80 missing transcript rows by default.
- Added `transcriptStatus`, `transcriptSource`, and `transcriptLanguage` to make transcript availability easier to diagnose in Apify dataset views.
- Stream completed dataset rows during processing so large runs preserve completed videos instead of waiting until the very end.
- Improved publish-date parsing from exact watch-page date text before falling back to relative `publishedTimeText`.
- Kept this actor focused on normal YouTube videos; Shorts collection belongs in the separate YouTube Shorts Scraper actor.
- Made `proxyMode: "auto"` cheaper by stopping default residential retries for transcript-only misses; transcript proxy fallback now only runs when explicitly enabled.
- Added stronger direct retries for incomplete video metadata before proxy fallback, and avoid caching failed channel About lookups as permanent misses.
- Added hard HTTP request timeouts and relative publish-date recovery so slow proxy requests cannot hang a run and hosted rows do not lose `publishedAt` when YouTube only exposes `publishedTimeText`.
- Improved direct video URL handling for `youtu.be`, `/embed/`, and `/live/` formats.
- Added direct retry recovery for transient high-concurrency video detail failures.
- Added channel About/profile retry recovery for transient channel link and metadata misses.
- Made multi-source runs resilient when one channel URL fails; valid sources still return rows and source errors are reported in run metadata.

### 2026-05-01

- Improved video detail speed with parallel channel and video processing.
- Added automatic proxy fallback mode to reduce residential proxy cost.
- Improved transcript extraction using direct caption tracks before slower UI fallback.
- Added richer channel metadata extraction for subscriber count, total videos, total views, profile image, banner image, and verification status.
- Added dataset and output schemas so Apify users, APIs, and AI agents can understand the returned fields.
- Reworked README and input descriptions for clearer setup, use cases, speed guidance, and transcript expectations.
