Audio Extractor โ€” Video to MP3, WAV, FLAC (Speech Ready) avatar

Audio Extractor โ€” Video to MP3, WAV, FLAC (Speech Ready)

Pricing

from $2.10 / 1,000 audio minute extracteds

Go to Apify Store
Audio Extractor โ€” Video to MP3, WAV, FLAC (Speech Ready)

Audio Extractor โ€” Video to MP3, WAV, FLAC (Speech Ready)

Extract audio from videos in bulk as MP3, M4A, Opus, WAV or FLAC. One-click speech-to-text format (WAV 16 kHz mono) for Whisper and speech APIs, loudness normalization, metadata removed. Pay per minute.

Pricing

from $2.10 / 1,000 audio minute extracteds

Rating

0.0

(0)

Developer

Leandro Zanatta

Leandro Zanatta

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

14 hours ago

Last modified

Categories

Share

Audio Extractor โ€” Video to MP3, WAV, FLAC (Speech-to-Text Ready)

Pull the audio track out of videos in bulk and save it as MP3, M4A (AAC), Opus, WAV or FLAC. One switch gives you the format speech recognition engines like Whisper work best with (WAV, mono, 16 kHz), and optional loudness normalization fixes recordings that are too quiet or too loud. Metadata is removed from every file.

Waveform of a quiet audio track before and after loudness normalization, and output size per hour for each format

Real output of this Actor: a quiet street recording raised from -36 to -17 LUFS, and measured file sizes per format. Audio: Qviri, CC BY-SA 4.0, Wikimedia Commons.

Why use it

  • ๐Ÿ—ฃ๏ธ Better transcripts, lower cost: speech-to-text APIs charge by upload size or duration and work best with 16 kHz mono audio. Sending a 1 GB video when 110 MB of WAV (or 25 MB of Opus) carries the same speech wastes bandwidth and time.
  • ๐Ÿ”Š Even volume: loudness normalization to the podcast standard (-16 LUFS) makes quiet phone recordings and loud webinars sound the same.
  • ๐ŸŽง Every common format: MP3 for compatibility, M4A for Apple and podcasts, Opus for the smallest files, FLAC/WAV for lossless archives and editing.
  • ๐Ÿ“ฆ Bulk and automated: hundreds of files per run, from URLs, Google Drive or Dropbox links, on a schedule or from a webhook.
  • ๐Ÿ”’ Metadata removed: no device, software or location tags in the output.
  • ๐Ÿงฉ No FFmpeg setup: send URLs, get audio links back.

Who uses it

SectorTypical use
AI & speech-to-text pipelinesPrepare audio for Whisper, Deepgram, AssemblyAI, Google or Azure Speech
Podcasters & content creatorsTurn video episodes, lives and interviews into podcast audio
Education & e-learningAudio versions of lectures and courses for listening on the go
Companies & legalArchive the audio of meetings, hearings, depositions and calls
Media monitoring & researchExtract speech from news, social and broadcast videos for analysis
Call centers & QANormalize recorded video calls before transcription and scoring
Data teamsBuild audio datasets (speech, sound events) from video collections

Formats

Measured on real output (size per hour of audio):

FormatSettingSize / hourBest for
Opus64 kbps~25 MBSmallest files, excellent for speech
MP3 (default)128 kbps~55 MBPlays on any device
M4A (AAC)128 kbps~56 MBApple devices, podcast platforms
Speech-ready WAV16 kHz, mono~110 MBSpeech-to-text engines (Whisper & co.)
FLAClossless~540 MBArchive without quality loss
WAVoriginal rate~660 MB (48 kHz stereo)Audio editing software

Lossy formats use the Bitrate you choose (32โ€“320 kbps).

Input

FieldTypeDefaultDescription
Media URLs (mediaUrls)arrayโ€”Links to videos or audio files. Public Google Drive and Dropbox share links are converted automatically. Up to 4 GB per file.
Speech-to-text ready (speechReady)booleanfalseOverrides format, rate and channels with WAV, 16 kHz, mono
Format (format)stringmp3mp3, m4a, opus, wav or flac
Bitrate (kbps) (bitrateKbps)integer128For MP3, M4A and Opus (32โ€“320)
Sample rate (sampleRate)stringoriginaloriginal, 16000, 22050, 44100 or 48000
Channels (channels)stringoriginaloriginal, mono or stereo
Normalize loudness (normalizeLoudness)booleanfalseEBU R128 normalization to -16 LUFS
Max minutes per file (maxDurationMinutes)integer180Only the first N minutes are extracted and billed

Example:

{
"mediaUrls": [
"https://example.com/webinar.mp4",
"https://www.dropbox.com/s/abc123/interview.mov?dl=0"
],
"format": "mp3",
"bitrateKbps": 96,
"channels": "mono",
"normalizeLoudness": true
}

Supported inputs: MP4, MOV, WebM, MKV, AVI, FLV, MPEG-TS, and audio files such as M4A, OGG, WAV, FLAC and MP3. The first audio track of each file is used.

Output

One dataset row per file. The audio file is stored in the run's key-value store and linked in audioUrl.

{
"sourceUrl": "https://example.com/webinar.mp4",
"audioUrl": "https://api.apify.com/v2/key-value-stores/.../records/audio-0001.mp3",
"format": "mp3",
"durationSeconds": 3605.2,
"sizeMB": 55.1,
"sourceAudioCodec": "aac",
"truncated": false,
"billedMinutes": 61
}
FieldMeaning
audioUrlDownload link of the audio file
format, durationSeconds, sizeMBWhat you got
sourceAudioCodecAudio codec of the input (aac, opus, mp3, pcmโ€ฆ)
truncatedtrue when only the first Max minutes per file were extracted
billedMinutesStarted minutes billed for this file
errorPresent only when a file failed. Files without an audio track return a clear error and are not billed

Pricing

Pay per started minute of audio extracted, plus a tiny per-run start fee. No subscription. Failed files are free. Apify plan discounts apply automatically (see the Pricing tab).

ExampleMinutes billed
30-second clip1
1-hour webinar60
100 short videos of 2 minutes200

Use Max minutes per file or the run's maximum cost setting to cap spending; the Actor stops cleanly when the limit is reached.

Use it from code or an AI agent

Python: extract, then transcribe

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("adorable_partial/audio-extractor").call(
run_input={"mediaUrls": ["https://example.com/webinar.mp4"], "speechReady": True}
)
audio_urls = [i["audioUrl"] for i in client.dataset(run["defaultDatasetId"]).iterate_items() if "audioUrl" in i]
# Optional: send the audio to the Whisper transcriber Actor
t = client.actor("adorable_partial/whisper-audio-video-transcriber").call(run_input={"mediaUrls": audio_urls})
for item in client.dataset(t["defaultDatasetId"]).iterate_items():
print(item.get("text", "")[:200])

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('adorable_partial/audio-extractor').call({
mediaUrls: ['https://example.com/episode-12.mp4'],
format: 'mp3',
normalizeLoudness: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].audioUrl);

HTTP

curl -X POST "https://api.apify.com/v2/acts/adorable_partial~audio-extractor/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"mediaUrls": ["https://example.com/webinar.mp4"], "format": "opus", "bitrateKbps": 48}'

No-code and agents: use the Apify modules in Make, Zapier or n8n, trigger runs from webhooks or schedules, or let an AI agent call it through the Apify MCP server.

FAQ

Which format should I pick for transcription? Turn on Speech-to-text ready (WAV, 16 kHz, mono). If your transcription service limits upload size, use Opus 32โ€“64 kbps mono instead: about 4โ€“6ร— smaller with practically the same accuracy for speech.

Does normalization change the content? No. It only adjusts the volume so the whole file sits around -16 LUFS without clipping peaks.

Can it extract from YouTube or social media links? It needs a direct link to the file (or a public Google Drive / Dropbox link). Page URLs of streaming sites are not supported.

What about videos with several audio tracks? The first audio track is extracted.

How long can files be? Up to 4 GB per file; up to Max minutes per file (default 180) are extracted.

Is my data kept? Files are processed inside your run and stored only in your own Apify storage, under your account's data-retention settings.

Support

Need another format or setting? Open an issue in the Issues tab.