Transcribe | Transcribe any video or audio avatar

Transcribe | Transcribe any video or audio

Pricing

from $5.00 / 1,000 results

Go to Apify Store
Transcribe | Transcribe any video or audio

Transcribe | Transcribe any video or audio

Transcribe any video or audio from YouTube, TikTok, Instagram, Twitter, and 1000+ sites

Pricing

from $5.00 / 1,000 results

Rating

5.0

(1)

Developer

REXREUS D.O

REXREUS D.O

Maintained by Community

Actor stats

1

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

๐ŸŽ™๏ธ Universal Video/Audio Transcriber

Transcribe any video or audio from YouTube, TikTok, Instagram, Twitter, and 1000+ sites

Powered by OpenAI ยท Groq ยท AssemblyAI ยท Deepgram โ€” your choice of AI engine

Quick Start Apify Node.js LICENSE


โœจ Why This Actor?

FeatureWhat You Get
๐ŸŒ 1000+ platformsYouTube, TikTok, Instagram, Twitter, Vimeo, SoundCloud, and virtually any site with video/audio
๐Ÿค– 4 AI providersOpenAI Whisper, Groq (free tier!), AssemblyAI, Deepgram โ€” pick the best fit
๐Ÿ’ฐ Free tier optionGroq offers generous free API credits โ€” start transcribing for $0
๐Ÿ“ Multiple formatsJSON (full metadata), SRT (subtitles), VTT (web captions), TXT (plain text)
๐ŸŒ 99+ languagesAuto-detect or specify language โ€” from English to Japanese to Arabic
๐Ÿ”Š Speaker diarizationIdentify who said what with AssemblyAI or Deepgram
๐Ÿ”„ Batch processingPaste 1 URL or 100 โ€” the actor handles them all with per-URL error isolation
๐Ÿ“ No length limitsAutomatic compression + smart chunking handles hours of audio
๐Ÿ”— Direct media URLsBypass extraction entirely for direct MP3/MP4/WAV links
๐Ÿ›ก๏ธ Resilient by designAutomatic retries, graceful error handling, state persistence across migrations

๐Ÿš€ Quick Start

  1. Open the Actor on Apify Console
  2. Paste your URLs โ€” YouTube, TikTok, Instagram, etc.
  3. Choose a provider โ€” OpenAI, Groq, AssemblyAI, or Deepgram
  4. Enter your API key โ€” Get one from your chosen provider
  5. Click Start โ€” Results appear in the Dataset within minutes

Option 2: Local Development

# Clone the repository
git clone https://github.com/your-username/universal-transcriber.git
cd universal-transcriber
# Install dependencies
npm install
# Set up your API key
cp .env.example .env
# Edit .env and add your API key
# Run
npm start

๐Ÿ“Š Provider Comparison

Choose the provider that fits your needs:

ProviderPrice/minBest ForFile LimitTimestampsTranslationDiarization
OpenAI$0.006All-around accuracy25 MB*โœ…โœ…โŒ
GroqFREE tierSpeed + budget25 MB*โœ…โœ…โŒ
AssemblyAI$0.0035Long videos + accuracyโˆžโœ…โŒโœ…
Deepgram$0.0043Low latencyโˆžโœ…โŒโœ…

*Files over 25 MB are automatically compressed and chunked โ€” no manual work needed.

Available Models

ModelProviderHighlights
whisper-1OpenAIClassic, reliable, all output formats
gpt-4o-transcribeOpenAINewer architecture, better at context
gpt-4o-mini-transcribeOpenAICheapest OpenAI option
whisper-large-v3GroqHigh accuracy, free tier available
whisper-large-v3-turboGroqFastest option, free tier
universal-3-proAssemblyAIBest overall accuracy, no file limit
universal-2AssemblyAI99 languages, no file limit
nova-3DeepgramLatest Deepgram, best accuracy
nova-2DeepgramProven and stable

๐Ÿ”ง How It Works

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Input: URLs + Provider โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”‚
โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ 1. EXTRACTION (yt-dlp) โ”‚
โ”‚ โ€ข Downloads audio from 1000+ platforms โ”‚
โ”‚ โ€ข Extracts metadata (title, duration, platform) โ”‚
โ”‚ โ€ข Retries with exponential backoff on failure โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”‚
โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ 2. COMPRESSION (ffmpeg) โ”‚
โ”‚ โ€ข Converts to 16kHz mono 64kbps MP3 โ”‚
โ”‚ โ€ข ~50 minutes of audio fits in 25 MB โ”‚
โ”‚ โ€ข Reduces API costs by minimizing file size โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”‚
โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ 3. CHUNKING (if needed) โ”‚
โ”‚ โ€ข Silence-aware splitting at 25 MB boundaries โ”‚
โ”‚ โ€ข Preserves word boundaries (no mid-word cuts) โ”‚
โ”‚ โ€ข Tracks time offsets for seamless merging โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”‚
โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ 4. TRANSCRIPTION (AI Provider) โ”‚
โ”‚ โ€ข Sends to OpenAI / Groq / AssemblyAI / Deepgram โ”‚
โ”‚ โ€ข Handles provider-specific API formats โ”‚
โ”‚ โ€ข Merges chunks with correct time offsets โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”‚
โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ 5. OUTPUT FORMATTING โ”‚
โ”‚ โ€ข JSON โ€” Full metadata + word/segment timestamps โ”‚
โ”‚ โ€ข SRT โ€” Standard subtitle format โ”‚
โ”‚ โ€ข VTT โ€” Web Video Text Tracks โ”‚
โ”‚ โ€ข TXT โ€” Clean plain text โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”‚
โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ 6. STORAGE (Apify Platform) โ”‚
โ”‚ โ€ข Results โ†’ Dataset (searchable, exportable) โ”‚
โ”‚ โ€ข Status โ†’ Key-Value Store (real-time progress) โ”‚
โ”‚ โ€ข Per-URL error isolation (one failure โ‰  batch failure) โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐ŸŒ YouTube Extraction

This actor automatically uses Apify residential proxy for YouTube extraction to bypass bot detection. No manual proxy configuration is needed โ€” it's enabled by default in the actor input.


๐Ÿ“ฅ Input

Basic Usage

{
"urls": [
{ "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ" },
{ "url": "https://www.tiktok.com/@user/video/1234567890" }
],
"provider": "openai",
"model": "whisper-1",
"openaiApiKey": "sk-...",
"outputFormats": ["json", "srt"]
}

Full Options

{
"urls": [{ "url": "https://www.youtube.com/watch?v=..." }],
"directMediaUrls": ["https://example.com/audio.mp3"],
"provider": "groq",
"model": "whisper-large-v3-turbo",
"groqApiKey": "gsk_...",
"outputFormats": ["json", "srt", "vtt", "txt"],
"language": "en",
"translateToEnglish": false,
"prompt": "This is a technical discussion about Kubernetes and Docker.",
"speakerDiarization": false,
"audioFormat": "mp3",
"maxDurationMinutes": 0,
"debugMode": false,
"proxyConfiguration": {
"useApifyProxy": false
}
}

Input Fields Reference

FieldTypeRequiredDefaultDescription
urlsArrayโœ… Yesโ€”List of video/audio URLs to transcribe
directMediaUrlsArrayNoโ€”Direct links to media files (bypasses extraction)
providerStringNoopenaiopenai, groq, assemblyai, or deepgram
modelStringNowhisper-1Transcription model (see provider comparison)
openaiApiKeyString*โ€”Required when provider is openai
groqApiKeyString*โ€”Required when provider is groq
assemblyaiApiKeyString*โ€”Required when provider is assemblyai
deepgramApiKeyString*โ€”Required when provider is deepgram
outputFormatsArrayNo["json"]Output formats: json, srt, vtt, txt
languageStringNoautoISO-639-1 code (e.g., en, id, ja)
translateToEnglishBooleanNofalseTranslate non-English to English
promptStringNoโ€”Context prompt for better accuracy
speakerDiarizationBooleanNofalseIdentify speakers (AssemblyAI/Deepgram only)
audioFormatStringNomp3Extraction format: mp3, m4a, wav, webm, flac
maxDurationMinutesIntegerNo0Skip videos longer than X min (0 = no limit)
debugModeBooleanNofalseVerbose logging for troubleshooting

๐Ÿ“ค Output

Each URL produces a record in the Dataset:

{
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Music Video)",
"provider": "openai",
"model": "whisper-1",
"language": "en",
"duration": 213,
"status": "success",
"transcript": {
"text": "We're no strangers to love...",
"segments": [
{
"start": 0.0,
"end": 4.5,
"text": "We're no strangers to love"
}
]
},
"formats": {
"json": { "..." : "..." },
"srt": "1\n00:00:00,000 --> 00:00:04,500\nWe're no strangers to love\n\n...",
"txt": "We're no strangers to love..."
},
"metadata": {
"platform": "youtube",
"extractionTime": 3.2,
"compressionRatio": 0.85,
"totalChunks": 1
},
"error": null
}

Output Formats

JSON โ€” Full structured data with timestamps, confidence scores, and segments:

{
"text": "Full transcription text...",
"segments": [
{ "start": 0.0, "end": 4.5, "text": "First segment..." },
{ "start": 4.5, "end": 8.2, "text": "Second segment..." }
]
}

SRT โ€” Standard subtitle format for video players:

1
00:00:00,000 --> 00:00:04,500
We're no strangers to love
2
00:00:04,500 --> 00:00:08,200
You know the rules and so do I

VTT โ€” Web-compatible subtitle format:

WEBVTT
00:00:00.000 --> 00:00:04.500
We're no strangers to love
00:00:04.500 --> 00:00:08.200
You know the rules and so do I

TXT โ€” Clean plain text:

We're no strangers to love
You know the rules and so do I

๐Ÿ—๏ธ Architecture

src/
โ”œโ”€โ”€ main.js # Entry point โ€” orchestrates the full pipeline
โ”œโ”€โ”€ config.js # Provider definitions, validation, constants
โ”œโ”€โ”€ extractor.js # yt-dlp wrapper with retry + error classification
โ”œโ”€โ”€ processor.js # ffmpeg compression + silence-aware chunking
โ”œโ”€โ”€ formatter.js # Output format converters (JSON/SRT/VTT/TXT)
โ”œโ”€โ”€ storage.js # Apify Dataset + KVS persistence
โ”œโ”€โ”€ state.js # Resumable state tracking across migrations
โ”œโ”€โ”€ utils.js # Retry logic, temp dirs, platform detection
โ””โ”€โ”€ providers/
โ”œโ”€โ”€ base.js # Abstract provider interface
โ”œโ”€โ”€ openai.js # OpenAI Whisper + GPT-4o transcription
โ”œโ”€โ”€ groq.js # Groq Whisper (OpenAI-compatible API)
โ”œโ”€โ”€ assemblyai.js # AssemblyAI (URL-based, no file upload)
โ”œโ”€โ”€ deepgram.js # Deepgram (URL-based, streaming-capable)
โ””โ”€โ”€ index.js # Provider factory

Key Design Decisions

  • Compression-first: Audio is compressed to 16kHz mono 64kbps before sending to APIs โ€” this means ~50 min of audio fits in 25 MB, dramatically reducing costs
  • Silence-aware chunking: When files exceed provider limits, we split at silence boundaries instead of cutting mid-sentence
  • Per-URL isolation: One failed URL never aborts the entire batch โ€” errors are captured and stored alongside successful results
  • Provider abstraction: Adding a new provider requires only implementing the BaseProvider interface โ€” no changes to the pipeline
  • State persistence: Progress is saved to Apify KVS, so if the actor migrates servers mid-batch, it resumes where it left off

๐Ÿ”‘ Getting API Keys

OpenAI

  1. Go to platform.openai.com
  2. Navigate to API Keys
  3. Create a new key starting with sk-...
  4. Pricing: $0.006/min (whisper-1), $0.02/min (gpt-4o-transcribe)
  1. Go to console.groq.com
  2. Sign up for a free account
  3. Create an API key
  4. Pricing: Free tier available with generous rate limits

AssemblyAI

  1. Go to assemblyai.com
  2. Sign up and get your API key
  3. Pricing: $0.0035/min โ€” best value for long videos

Deepgram

  1. Go to deepgram.com
  2. Sign up for $200 free credit
  3. Create an API key
  4. Pricing: $0.0043/min

๐Ÿ› ๏ธ Development

Prerequisites

  • Node.js โ‰ฅ 20
  • ffmpeg (for audio compression)
  • yt-dlp (for audio extraction from URLs)

Local Setup

# Install dependencies
npm install
# Create environment file
cp .env.example .env
# Add your API key to .env
# Run
npm start
# Run tests
npm test

Adding a New Provider

  1. Create src/providers/your-provider.js extending BaseProvider
  2. Implement transcribe(audioBuffer, options) and get supportsTimestamps()
  3. Add your provider to PROVIDERS and MODEL_TO_PROVIDER in config.js
  4. Update the factory in providers/index.js
// src/providers/your-provider.js
import { BaseProvider } from './base.js';
export class YourProvider extends BaseProvider {
constructor(apiKey) {
super(apiKey);
// Initialize your SDK client
}
async transcribe(audioBuffer, options = {}) {
// Call your provider's API
return { text: '...', segments: [...] };
}
get supportsTimestamps() {
return true;
}
}

๐Ÿณ Docker

The actor runs in a Docker container with all dependencies pre-installed:

FROM apify/actor-node:22-alpine
# Install ffmpeg + yt-dlp
RUN apk add --no-cache ffmpeg curl \
&& curl -L https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp \
-o /usr/local/bin/yt-dlp \
&& chmod +x /usr/local/bin/yt-dlp
# Install Node.js dependencies
COPY package*.json ./
RUN npm install --omit=dev
# Copy source code
COPY . ./

Build and run locally:

docker build -t universal-transcriber .
docker run --env-file .env universal-transcriber

๐Ÿ“‹ Supported Platforms

yt-dlp supports 1000+ sites. Here are the most popular:

PlatformStatusNotes
YouTubeโœ…Videos, Shorts, Playlists
TikTokโœ…Public videos
Instagramโœ…Reels, Posts, Stories
Twitter/Xโœ…Public tweets with video
Facebookโœ…Public videos
Vimeoโœ…All videos
SoundCloudโœ…Tracks, playlists
Twitchโœ…VODs, clips
Redditโœ…Video posts
Dailymotionโœ…All videos
LinkedInโœ…Public videos
Direct URLsโœ…MP3, MP4, WAV, M4A, WEBM, OGG, FLAC
...and 990+ moreโœ…See yt-dlp supported sites

โ“ FAQ

Q: How much does it cost? A: Depends on the provider. Groq offers a free tier. OpenAI charges $0.006/min. For a 1-hour video, expect ~$0.36 with OpenAI or $0 with Groq's free tier.

Q: How long does transcription take? A: Extraction + compression typically takes 1-3 minutes. API transcription is usually 10-30% of the audio duration. A 10-minute video is typically done in under 5 minutes total.

Q: Can I transcribe private/unlisted videos? A: Yes, if you provide a proxy configuration or the platform allows access via the URL alone. For geo-restricted content, use the proxy settings.

Q: What about very long videos (2+ hours)? A: The actor automatically compresses and chunks long audio. AssemblyAI and Deepgram have no file size limits, making them ideal for very long content.

Q: Can I translate non-English audio to English? A: Yes! Enable translateToEnglish with OpenAI whisper-1 or Groq whisper-large-v3.

Q: How do I identify different speakers? A: Enable speakerDiarization with AssemblyAI or Deepgram. This adds speaker labels (SPEAKER 1, SPEAKER 2, etc.) to the output.

Q: Can I run this locally without Apify? A: Yes! Install ffmpeg and yt-dlp, set your API key in .env, and run npm start.


๐Ÿค Contributing

Contributions are welcome! Here's how:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

๐Ÿ“„ License

This project is licensed under the MIT License โ€” see the LICENSE file for details.


๐Ÿ™ Acknowledgments

  • yt-dlp โ€” The incredible media downloader powering extraction
  • ffmpeg โ€” The Swiss army knife for audio processing
  • Apify โ€” The platform that makes deployment effortless
  • OpenAI, Groq, AssemblyAI, Deepgram โ€” The AI engines making transcription possible

โญ Star this repo if you find it useful!

Made with โค๏ธ for the transcription community