# Audio Extractor — Video to MP3, WAV, FLAC (Speech Ready) (`adorable_partial/audio-extractor`) Actor

Extract audio from videos in bulk as MP3, M4A, Opus, WAV or FLAC. One-click speech-to-text format (WAV 16 kHz mono) for Whisper and speech APIs, loudness normalization, metadata removed. Pay per minute.

- **URL**: https://apify.com/adorable\_partial/audio-extractor.md
- **Developed by:** [Leandro Zanatta](https://apify.com/adorable_partial) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.10 / 1,000 audio minute extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Audio Extractor — Video to MP3, WAV, FLAC (Speech-to-Text Ready)

Pull the **audio track out of videos** in bulk and save it as **MP3, M4A (AAC), Opus, WAV or FLAC**. One switch gives you the format speech recognition engines like **Whisper** work best with (**WAV, mono, 16 kHz**), and optional **loudness normalization** fixes recordings that are too quiet or too loud. Metadata is removed from every file.

![Waveform of a quiet audio track before and after loudness normalization, and output size per hour for each format](https://api.apify.com/v2/key-value-stores/59pZpQChkp7amu9f4/records/demo.png?signature=1rvs0Q8YzOaPzPBz9yz4D)

*Real output of this Actor: a quiet street recording raised from -36 to -17 LUFS, and measured file sizes per format. Audio: Qviri, CC BY-SA 4.0, Wikimedia Commons.*

### Why use it

- 🗣️ **Better transcripts, lower cost**: speech-to-text APIs charge by upload size or duration and work best with 16 kHz mono audio. Sending a 1 GB video when 110 MB of WAV (or 25 MB of Opus) carries the same speech wastes bandwidth and time.
- 🔊 **Even volume**: loudness normalization to the podcast standard (-16 LUFS) makes quiet phone recordings and loud webinars sound the same.
- 🎧 **Every common format**: MP3 for compatibility, M4A for Apple and podcasts, Opus for the smallest files, FLAC/WAV for lossless archives and editing.
- 📦 **Bulk and automated**: hundreds of files per run, from URLs, Google Drive or Dropbox links, on a schedule or from a webhook.
- 🔒 **Metadata removed**: no device, software or location tags in the output.
- 🧩 **No FFmpeg setup**: send URLs, get audio links back.

### Who uses it

| Sector | Typical use |
|---|---|
| **AI & speech-to-text pipelines** | Prepare audio for Whisper, Deepgram, AssemblyAI, Google or Azure Speech |
| **Podcasters & content creators** | Turn video episodes, lives and interviews into podcast audio |
| **Education & e-learning** | Audio versions of lectures and courses for listening on the go |
| **Companies & legal** | Archive the audio of meetings, hearings, depositions and calls |
| **Media monitoring & research** | Extract speech from news, social and broadcast videos for analysis |
| **Call centers & QA** | Normalize recorded video calls before transcription and scoring |
| **Data teams** | Build audio datasets (speech, sound events) from video collections |

### Formats

Measured on real output (size per hour of audio):

| Format | Setting | Size / hour | Best for |
|---|---|---|---|
| **Opus** | 64 kbps | ~25 MB | Smallest files, excellent for speech |
| **MP3** (default) | 128 kbps | ~55 MB | Plays on any device |
| **M4A (AAC)** | 128 kbps | ~56 MB | Apple devices, podcast platforms |
| **Speech-ready WAV** | 16 kHz, mono | ~110 MB | Speech-to-text engines (Whisper & co.) |
| **FLAC** | lossless | ~540 MB | Archive without quality loss |
| **WAV** | original rate | ~660 MB (48 kHz stereo) | Audio editing software |

Lossy formats use the **Bitrate** you choose (32–320 kbps).

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| **Media URLs** (`mediaUrls`) | array | — | Links to videos or audio files. Public Google Drive and Dropbox share links are converted automatically. Up to 4 GB per file. |
| **Speech-to-text ready** (`speechReady`) | boolean | `false` | Overrides format, rate and channels with WAV, 16 kHz, mono |
| **Format** (`format`) | string | `mp3` | `mp3`, `m4a`, `opus`, `wav` or `flac` |
| **Bitrate (kbps)** (`bitrateKbps`) | integer | `128` | For MP3, M4A and Opus (32–320) |
| **Sample rate** (`sampleRate`) | string | `original` | `original`, `16000`, `22050`, `44100` or `48000` |
| **Channels** (`channels`) | string | `original` | `original`, `mono` or `stereo` |
| **Normalize loudness** (`normalizeLoudness`) | boolean | `false` | EBU R128 normalization to -16 LUFS |
| **Max minutes per file** (`maxDurationMinutes`) | integer | `180` | Only the first N minutes are extracted and billed |

Example:

```json
{
  "mediaUrls": [
    "https://example.com/webinar.mp4",
    "https://www.dropbox.com/s/abc123/interview.mov?dl=0"
  ],
  "format": "mp3",
  "bitrateKbps": 96,
  "channels": "mono",
  "normalizeLoudness": true
}
```

**Supported inputs:** MP4, MOV, WebM, MKV, AVI, FLV, MPEG-TS, and audio files such as M4A, OGG, WAV, FLAC and MP3. The first audio track of each file is used.

### Output

One dataset row per file. The audio file is stored in the run's key-value store and linked in `audioUrl`.

```json
{
  "sourceUrl": "https://example.com/webinar.mp4",
  "audioUrl": "https://api.apify.com/v2/key-value-stores/.../records/audio-0001.mp3",
  "format": "mp3",
  "durationSeconds": 3605.2,
  "sizeMB": 55.1,
  "sourceAudioCodec": "aac",
  "truncated": false,
  "billedMinutes": 61
}
```

| Field | Meaning |
|---|---|
| `audioUrl` | Download link of the audio file |
| `format`, `durationSeconds`, `sizeMB` | What you got |
| `sourceAudioCodec` | Audio codec of the input (aac, opus, mp3, pcm…) |
| `truncated` | `true` when only the first *Max minutes per file* were extracted |
| `billedMinutes` | Started minutes billed for this file |
| `error` | Present only when a file failed. Files without an audio track return a clear error and are not billed |

### Pricing

Pay per **started minute of audio** extracted, plus a tiny per-run start fee. No subscription. Failed files are free. Apify plan discounts apply automatically (see the **Pricing** tab).

| Example | Minutes billed |
|---|---|
| 30-second clip | 1 |
| 1-hour webinar | 60 |
| 100 short videos of 2 minutes | 200 |

Use **Max minutes per file** or the run's *maximum cost* setting to cap spending; the Actor stops cleanly when the limit is reached.

### Use it from code or an AI agent

**Python: extract, then transcribe**

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("adorable_partial/audio-extractor").call(
    run_input={"mediaUrls": ["https://example.com/webinar.mp4"], "speechReady": True}
)
audio_urls = [i["audioUrl"] for i in client.dataset(run["defaultDatasetId"]).iterate_items() if "audioUrl" in i]

## Optional: send the audio to the Whisper transcriber Actor
t = client.actor("adorable_partial/whisper-audio-video-transcriber").call(run_input={"mediaUrls": audio_urls})
for item in client.dataset(t["defaultDatasetId"]).iterate_items():
    print(item.get("text", "")[:200])
```

**JavaScript**

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('adorable_partial/audio-extractor').call({
    mediaUrls: ['https://example.com/episode-12.mp4'],
    format: 'mp3',
    normalizeLoudness: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].audioUrl);
```

**HTTP**

```bash
curl -X POST "https://api.apify.com/v2/acts/adorable_partial~audio-extractor/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"mediaUrls": ["https://example.com/webinar.mp4"], "format": "opus", "bitrateKbps": 48}'
```

**No-code and agents:** use the Apify modules in **Make**, **Zapier** or **n8n**, trigger runs from **webhooks** or **schedules**, or let an AI agent call it through the **Apify MCP server**.

### FAQ

**Which format should I pick for transcription?** Turn on **Speech-to-text ready** (WAV, 16 kHz, mono). If your transcription service limits upload size, use **Opus 32–64 kbps mono** instead: about 4–6× smaller with practically the same accuracy for speech.

**Does normalization change the content?** No. It only adjusts the volume so the whole file sits around -16 LUFS without clipping peaks.

**Can it extract from YouTube or social media links?** It needs a direct link to the file (or a public Google Drive / Dropbox link). Page URLs of streaming sites are not supported.

**What about videos with several audio tracks?** The first audio track is extracted.

**How long can files be?** Up to 4 GB per file; up to *Max minutes per file* (default 180) are extracted.

**Is my data kept?** Files are processed inside your run and stored only in your own Apify storage, under your account's data-retention settings.

### Related Actors

- [Audio & Video to Text (Whisper)](https://apify.com/adorable_partial/whisper-audio-video-transcriber) — transcripts, SRT and VTT subtitles
- [Video Optimizer](https://apify.com/adorable_partial/video-optimizer) — reduce FPS, resize and compress videos
- [Video Frame Extractor for AI](https://apify.com/adorable_partial/video-frame-extractor-ai) — turn videos into clean image datasets
- [Image & Video Anonymizer](https://apify.com/adorable_partial/image-anonymizer) — blur faces and license plates (GDPR / LGPD)

### Support

Need another format or setting? Open an issue in the **Issues** tab.

# Actor input Schema

## `mediaUrls` (type: `array`):

Direct links to videos (MP4, MOV, WebM, MKV...) or audio files to convert. Public Google Drive and Dropbox share links work too.

## `speechReady` (type: `boolean`):

Overrides the settings below with the format speech recognition engines such as Whisper work best with.

## `format` (type: `string`):

File format of the extracted audio.

## `bitrateKbps` (type: `integer`):

For MP3, M4A and Opus. 64 = voice, 128 = good, 192–320 = music.

## `sampleRate` (type: `string`):

Audio samples per second. Keep the original unless a tool asks for a specific rate.

## `channels` (type: `string`):

Mono halves the size and is enough for speech.

## `normalizeLoudness` (type: `boolean`):

Even out the volume to a podcast-standard level (-16 LUFS).

## `maxDurationMinutes` (type: `integer`):

Only the first N minutes of each file are extracted and billed.

## Actor input object example

```json
{
  "mediaUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"
  ],
  "speechReady": false,
  "format": "mp3",
  "bitrateKbps": 128,
  "sampleRate": "original",
  "channels": "original",
  "normalizeLoudness": false,
  "maxDurationMinutes": 180
}
```

# Actor output Schema

## `results` (type: `string`):

One row per file: link to the audio file, format, duration and size.

## `audio` (type: `string`):

The extracted audio files.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mediaUrls": [
        "https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("adorable_partial/audio-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "mediaUrls": ["https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"] }

# Run the Actor and wait for it to finish
run = client.actor("adorable_partial/audio-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mediaUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"
  ]
}' |
apify call adorable_partial/audio-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,adorable_partial/audio-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5BZngyJNmr0tFFenZ/builds/6JKxZXgftnE3Kg0vp/openapi.json
