# Changelog of Audio & Video to Text - Whisper Transcription API (`zenomastro/audio-video-to-text`) Actor

- **URL**: https://apify.com/zenomastro/audio-video-to-text/changelog.md
- **Full Actor documentation**: https://apify.com/zenomastro/audio-video-to-text.md

## Changelog

### 2026-10-03

- First release: transcribe audio and video files from direct URLs (mp3, wav, m4a, ogg, opus, flac, webm, mp4, mov, mkv) with Whisper large-v3-turbo.
- Output per file: plain text, timestamped segments, SRT, WebVTT (also saved as files in the key-value store with direct links) and RAG chunks with a configurable size.
- Language auto-detection or a chosen language, transcribe or translate to English, optional voice activity detection and context prompt.
- Long files are converted to 16 kHz mono mp3 and cut at natural pauses into parts of about 10 minutes that are transcribed in parallel and merged with correct timestamps. Up to 180 minutes and 2 GB per file.
- Every row has a status: COMPLETE, PARTIAL, VALID\_EMPTY, INVALID\_INPUT, UPSTREAM\_FAILED or TOO\_LONG. Only COMPLETE and PARTIAL files are charged, $0.004 per started audio minute actually transcribed; everything else is free.
- YouTube and social-media page links are rejected with a hint to the YouTube Transcript Actor. Private and local network addresses are blocked on every download and redirect.
- If the daily transcription capacity is used up, remaining files are returned as UPSTREAM\_FAILED and are not charged.
- Platform usage included in the price; subscriber discounts on paid Apify plans.
