# Changelog of Youtube Text & Metadata Scraper (`ilborso/youtube-text-scraper`) Actor

- **URL**: https://apify.com/ilborso/youtube-text-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/ilborso/youtube-text-scraper.md

## Changelog

All notable changes to this project are documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

### \[Unreleased]

#### Added

- **`PARALLEL` environment variable** - Parallel video processing support
  - Type: integer, Default: 1
  - Controls the number of videos to process in parallel
  - **Implementation**: Uses `asyncio.Semaphore` to limit concurrent operations + `asyncio.to_thread()` for I/O-bound operations
  - **Location**: [src/main.py:105-160](src/main.py#L105-L160) and [main() function](src/main.py#L430)
  - **Usage**: Set `PARALLEL=N` where N is the number of concurrent videos (e.g., PARALLEL=3, PARALLEL=20)
  - **Note**: Transcript extraction runs in thread pool via `asyncio.to_thread()` to prevent event loop blocking
  - **Performance impact**: Reduces total execution time significantly with concurrent requests; e.g., 20 parallel jobs ~3-4x faster than sequential
  - **Reason**: Allow users to leverage multiple concurrent requests to speed up large batch processing

- **`enableMetadata` flag** - New optional input to fetch video metadata and subtitles
  - Type: boolean, Default: false
  - When `enableMetadata=false` (default): behavior identical to previous version, backward compatible
  - When `enableMetadata=true`: fetches complete metadata and list of available subtitles
  - **Event emission**: when enabled, emits an Apify `apify-actor-metadata` event (via `Actor.charge(event_name='apify-actor-metadata')`) at the end of execution
  - Event emitted only once per run, after completion of all processing
  - **Implementation date**: 2026-06-29
  - **Reason**: allow users to choose between optimized mode (default) and complete metadata mode
  - **Performance impact**: ~3-5x increase in proxy credit consumption when enabled

#### Changed

- **Async implementation** - All video processing now uses async/await pattern
  - **`process_direct_urls()`** - Now async, returns awaitable coroutine
  - **`search_videos()`** - Now async, returns awaitable coroutine
  - **`_process_single_video()`** - New async method for processing individual videos
  - **`extract_transcript()` calls** - Now wrapped with `asyncio.to_thread()` to run in thread pool
  - **Reason**: Enable concurrent processing with `asyncio.Semaphore` for better parallelization; thread pool prevents event loop blocking on I/O operations
  - **Backward compatibility**: Changes require updating all call sites with `await` keyword

#### Restored (Feature Reactivation)

- **`get_video_metadata()`** - Method restored and made controllable via `enableMetadata` flag
  - **Location**: [src/main.py:207-245](src/main.py#L207-L245)
  - Fetches: channel name/URL, views, duration, release\_date, thumbnail, hashtags
  - Resilient: fallback to None if error occurs
  - Execution: only if `enableMetadata=true`

- **`get_available_subtitles()`** - Method restored and made controllable via `enableMetadata` flag
  - **Location**: [src/main.py:182-205](src/main.py#L182-L205)
  - Fetches: list of available languages with metadata (language, code, generated/manual, translatable)
  - Resilient: fallback to \[] if error occurs
  - Execution: only if `enableMetadata=true`

- **Restored imports**: `import asyncio`, `from datetime import datetime`, `from pytubefix import YouTube`, `from youtube_transcript_api.proxies import GenericProxyConfig`
  - Required for `get_video_metadata()`, `get_available_subtitles()` methods, and async processing

#### Changed

- **Input schema** - Added optional `enableMetadata` field
  - Type: boolean
  - Default: false
  - Editor: checkbox
  - Description: "If true, also fetches video metadata and available subtitles (consumes more proxy credits)..."

- **Documentation (README.md)** - New "Metadata Retrieval Strategy" section
  - Explains default behavior vs metadata mode
  - Notes on performance and credit consumption
  - Usage example
  - Note in "What data does the Actor extract?" section

- **Method signatures**:
  - `process_direct_urls()` - Added optional parameter `enable_metadata=False`
  - `search_videos()` - Added optional parameter `enable_metadata=False`
  - Helper `_get_optional_metadata()` - New method to reduce duplication

- **main() function** - Added support for reading and passing the `enableMetadata` flag
  - Normalized reading to boolean with default false
  - Flag propagation to processing functions (`process_direct_urls()`, `search_videos()`)
  - Apify event emission via `Actor.charge(event_name='apify-actor-metadata')` at end of execution
  - Event emitted only if `enableMetadata=true`, after results saved to dataset and key-value store
  - Resilient error handling: warning log if charging fails, execution not interrupted

#### Disabled (Token Optimization)

- **`get_video_metadata()`** - Method disabled to save Scrape.do credits
  - Previously fetched: views, duration, release\_date, thumbnail, hashtags, channel info
  - **Location**: [src/main.py:199-237](src/main.py#L199-L237)
  - **Disable date**: Before 2026-06-29
  - **Impact**: Significant reduction in proxy requests

- **`get_available_subtitles()`** - Method disabled to save Scrape.do credits
  - Previously fetched: metadata on available subtitles (language, generated/manual, translatable)
  - **Location**: [src/main.py:170-193](src/main.py#L170-L193)
  - **Disable date**: Before 2026-06-29
  - **Impact**: Reduction in proxy requests per video

- **Commented import**: `from datetime import datetime`
  - No longer necessary since video metadata is disabled
  - **Location**: [src/main.py:5](src/main.py#L5)

#### Changed

- **Proxy Configuration**: Full Scrape.do proxy usage restored for all contexts
  - `extract_transcript()`: Proxy enabled via requests session
  - `search_videos()`: Proxy enabled for YouTube searches
  - `process_direct_urls()`: Proxy enabled for transcription
  - **Date**: 2026-06-29
  - **Reason**: Ensure stability and prevent blocking on all sensitive operations

### \[1.0.0] - Baseline

#### Initial Release

- Core implementation of YouTube scraper with Apify Actor
- Support for video search and direct URLs
- Transcript extraction with multi-language support
- Scrape.do proxy configuration
- Input/Output schema for Apify Console
