# Bulk Transcription: Audio & Video to Text from CSV or Sheet (`nerolabs/bulk-transcription`) Actor

Transcribes every audio or video link in an Apify dataset, CSV, Excel or Google Sheet, keeping your columns. Inputs: datasetId or fileUrl or mediaUrls, urlField, mode (plain text, or speakers with timestamps and SRT). Charged per started minute of audio. Agent-ready: x402, MCP.

- **URL**: https://apify.com/nerolabs/bulk-transcription.md
- **Developed by:** [Adam Pearce](https://apify.com/nerolabs) (community)
- **Categories:** AI, Automation, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $7.00 / 1,000 audio minute, plain texts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Transcription: Audio & Video to Text from CSV, Google Sheet or Dataset

Got a spreadsheet full of call recordings, podcast episodes, voicemails or meeting videos and need the words, not the audio? Point this Actor at the sheet and it transcribes every link in it, then hands the sheet back with a transcript next to each row. Your own columns (call ID, date, client, episode title) stay exactly where they were.

- **Bulk, from a list you already have.** An Apify dataset, a CSV, an Excel file, a Google Sheet link, or a plain list of URLs. No copying links into a form one at a time.
- **Keeps your columns.** Every row comes back with its original fields plus the transcript, so results line up with your CRM, call log or content calendar.
- **Speakers, timestamps and subtitles.** Switch on "Speakers, timestamps and subtitles" and each transcript reads like a conversation (`Speaker A: ...`, `Speaker B: ...`), with start and end times and a ready-made SRT subtitle file per recording.
- **Any audio or video.** MP3, WAV, M4A, MP4, MOV, WebM, OGG, FLAC, AAC and more. Video files are fine: only the audio is used. Google Drive and Dropbox share links are converted to direct downloads automatically.
- **Long recordings just work.** Files are cut into pieces and transcribed in parallel; a 13-minute recording comes back in about 15 seconds in plain text mode.
- **You only pay for what was transcribed.** Broken links, web pages, silent files and failed downloads are reported and never charged.

### What it is good for

- **Call recordings and voicemails** from phone systems, AI receptionists and sales dialers (GoHighLevel, Vapi, Retell, Twilio, Aircall, CallRail exports).
- **Podcast and YouTube back catalogues** for show notes, search and repurposing into posts.
- **Interviews, meetings and webinars** for notes and quotes.
- **AI agent pipelines**: feed transcripts straight into an LLM, a vector database or another Actor.

Want each call scored as well (booked or not, missed questions, lead quality, a 0 to 10 score)? Use the sister Actor **[Call Score](https://apify.com/nerolabs/call-score)**.

### Input

Give it one of:

| Input | What to put in |
|---|---|
| **Dataset** | Any Apify dataset with a column of audio or video links, for example a scraper's output. |
| **File or Google Sheet URL** | A link to a CSV, Excel, JSON or JSON Lines file, or a normal Google Sheets link shared as "Anyone with the link can view". |
| **Audio or video URLs** | A plain list of links for a quick run. |

The column holding the links is found automatically (it prefers values ending in `.mp3`, `.wav`, `.m4a`, `.mp4` and names like `recordingUrl` or `audio`); set **Recording URL field** to choose it yourself.

**Transcript type**

| Type | You get | Price per started minute |
|---|---|---|
| Plain text | One clean block of text per file. Accepts **Spelling hints** for names and jargon. | $0.01 |
| Speakers, timestamps and subtitles | One line per speaker turn, timed segments (optional) and SRT subtitles. | $0.02 |

Other useful settings: **Language** (two-letter code, or leave empty to detect), **Max minutes per file** (your cost ceiling per file, default 120), **Export files** (a downloadable CSV or Excel of the results), **Webhook URL** (POSTs the run summary to Zapier, Make, n8n, Slack or your API), and **Append to named dataset** (builds one growing archive across runs).

### Output

One row per recording. Example (speakers mode, shortened):

```json
{
    "callId": "C-1001",
    "client": "Demo Electrical",
    "recordingUrl": "https://nerolabs-samples.nerolabs.workers.dev/sample-call-electrician.mp3",
    "transcriptStatus": "ok",
    "mode": "detailed",
    "mediaType": "audio",
    "durationSec": 64.8,
    "durationMinutes": 1.08,
    "minutesCharged": 2,
    "speakerCount": 2,
    "wordCount": 208,
    "text": "Speaker A: Hello, you've reached Demo Electrical. This is the virtual assistant. How can I help you today?\nSpeaker B: Hi, yeah, I've got no power to half the house...",
    "srt": "1\n00:00:00,000 --> 00:00:04,350\nHello, you've reached Demo Electrical. This is the virtual assistant.\n..."
}
```

`transcriptStatus` is one of `ok`, `no_speech` (read end to end, no words: silence or music), `no_audio` (a video with no sound track), `not_supported`, `http_error`, `unreachable`, `blocked`, `too_large`, `invalid_url`, `no_url`, `transcription_failed` or `skipped_budget`. Only `ok` is charged. `statusDetail` explains every other status in plain words.

The run's **OUTPUT** record holds the totals: files transcribed, minutes of audio, minutes charged, words, status counts and export links.

### Pricing

Pay per event, no subscription:

- **$0.01 per started minute**, plain text.
- **$0.02 per started minute**, speakers, timestamps and subtitles.
- $0.01 per CSV or Excel export file, $0.02 per delivered webhook. A tiny start fee of $0.00005 per run.
- Apify subscribers on Bronze, Silver and Gold plans get 10%, 20% and 30% off.

Worked examples: 100 sales calls of 4 minutes each in plain text is 400 minutes, about **$4**. A 50-episode podcast back catalogue at 45 minutes an episode, with speakers and subtitles, is 2,250 minutes, about **$45**. Each file is rounded up to the next whole minute, so a 30-second voicemail costs one minute.

Set **Max minutes per file** and Apify's own **maximum charge per run** to cap any run exactly.

### How it works

Each file is downloaded inside the run (streamed to disk, so large videos are fine), converted to mono audio with ffmpeg, cut into pieces and sent to OpenAI's speech models: `gpt-4o-mini-transcribe` for plain text and `gpt-4o-transcribe-diarize` for speakers and timestamps. Only the audio is sent and nothing is kept after the run.

### FAQ

**Can it transcribe a YouTube or TikTok page link?** No. It needs a direct link to the audio or video file. Use a downloader Actor first and feed its dataset in here.

**How accurate are the speaker labels?** Very good on two-person calls and interviews. The letters (A, B, C) are the model's own and stay consistent within each 5-minute section of a long recording; on a long multi-speaker recording a person can get a different letter in a later section.

**Which languages?** About 57, including English, Spanish, French, German, Portuguese, Italian, Dutch, Polish, Japanese and Chinese. Language is detected automatically; setting it helps on short or noisy clips.

**Is my audio stored?** No. Files are processed inside your run and deleted at the end of it; the speech provider receives only the audio of each piece.

**Can an AI agent use it?** Yes. It charges per event, so agents can call it through the Apify MCP server or pay per run with x402.

If this saved you hours of typing up recordings, a review on the Store helps a lot. Questions or a format that will not read? Open an issue on the Issues tab and I will look at it the same day.

# Actor input Schema

## `datasetId` (type: `string`):

An Apify dataset holding one row per recording, for example call logs, podcast episodes or a scraper's video links. Every original column is kept and the transcript, duration and speaker count are added alongside. Use the picker rather than typing an ID.

## `fileUrl` (type: `string`):

A public link to a CSV, TSV, Excel, JSON or JSON Lines file holding one row per recording (call recordings, voicemails, podcast episodes, meeting videos). A normal Google Sheets link works: share it as 'Anyone with the link can view'. Used when no dataset is given.

## `fileFormat` (type: `string`):

Leave on 'Detect automatically' unless the link has no file extension and the server reports the wrong content type.

## `sheetName` (type: `string`):

Which sheet to read from an Excel workbook. Defaults to the first sheet.

## `mediaUrls` (type: `array`):

A plain list of direct links to audio or video files (MP3, WAV, M4A, MP4, MOV, WebM, OGG, FLAC and more; Google Drive and Dropbox share links work), for a quick run with no spreadsheet. Use the dataset, file or Google Sheet inputs above to keep your own columns alongside each transcript.

## `data` (type: `array`):

Rows as inline JSON, an alternative to a dataset or file. Each object needs a field holding the audio or video link.

## `urlField` (type: `string`):

The column holding the audio or video link. Left empty it is detected automatically, preferring a column whose values end in .mp3, .wav, .m4a or .mp4 over one merely named 'url'.

## `mode` (type: `string`):

'Plain text' is the cheapest: one block of text per file. 'Speakers and timestamps' labels who spoke (Speaker A, Speaker B), gives start and end times, and writes SRT subtitles, for calls, interviews, meetings and podcasts. Each is charged per started minute of audio at its own rate.

## `language` (type: `string`):

Two-letter code of the spoken language (en, es, fr, de, pt, it, nl, pl, ja, zh and about 50 more). Leave empty to detect it automatically; setting it helps accuracy on short or noisy clips.

## `vocabulary` (type: `string`):

Plain text mode only. Names, brands or jargon the recording uses, for example 'NICEIC, consumer unit, Part P, Brightwell'. Helps the model spell them right.

## `includeSrt` (type: `boolean`):

Speakers mode only. Adds an 'srt' column with ready-to-use subtitles for the recording.

## `includeSegments` (type: `boolean`):

Speakers mode only. Adds a 'segments' array: one entry per speaker turn with speaker, start, end and text. Useful for feeding an AI agent or building your own player.

## `maxMinutesPerFile` (type: `integer`):

Only the first this-many minutes of each file are transcribed and charged. Your cost ceiling per file: a 6-hour livestream cannot run up a bill you did not expect.

## `keep` (type: `string`):

Filtering happens after the file has been processed, so it does not make a run cheaper. 'Problems only' is the quick way to find broken links and unsupported files in a large list.

## `keepOriginalFields` (type: `boolean`):

Keep every column from the input row next to the transcript, so results line up with your own data (call ID, date, client, episode title). Turn off for transcripts and counts only.

## `concurrency` (type: `integer`):

How many files to work on in parallel. Transcription runs on a remote speech service, so this is mostly limited by download speed; 3 suits the default memory.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for one file to download before giving up on it. A timed-out file is never charged.

## `maxFileMb` (type: `integer`):

Files bigger than this are skipped and not charged. Video files are large; only their audio is used.

## `maxItems` (type: `integer`):

A hard ceiling on how many rows are read from the input, as a safety net on a large dataset.

## `exportFormats` (type: `array`):

Also write the results as a real downloadable CSV or Excel file, linked from the run's output. Timed segments are JSON-encoded into a single cell so they fit a spreadsheet.

## `outputDatasetName` (type: `string`):

Also append every kept row to a named dataset that persists across runs, building one growing transcript archive. Not charged again.

## `webhookUrl` (type: `string`):

POST the run summary to this URL when the run finishes, for Slack, Zapier, Make, n8n or your own API. Charged only on a confirmed 2xx response.

## Actor input object example

```json
{
  "fileFormat": "auto",
  "mediaUrls": [
    "https://nerolabs-samples.nerolabs.workers.dev/sample-call-electrician.mp3",
    "https://nerolabs-samples.nerolabs.workers.dev/sample-call-plumber.mp3"
  ],
  "mode": "detailed",
  "includeSrt": true,
  "includeSegments": false,
  "maxMinutesPerFile": 120,
  "keep": "all",
  "keepOriginalFields": true,
  "concurrency": 3,
  "requestTimeoutSecs": 120,
  "maxFileMb": 500,
  "exportFormats": []
}
```

# Actor output Schema

## `results` (type: `string`):

Every original row with the transcript, duration, speaker count and status added.

## `summary` (type: `string`):

Statuses, minutes transcribed and charged, word totals, export links and warnings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "fileUrl": "",
    "mediaUrls": [
        "https://nerolabs-samples.nerolabs.workers.dev/sample-call-electrician.mp3",
        "https://nerolabs-samples.nerolabs.workers.dev/sample-call-plumber.mp3"
    ],
    "mode": "detailed"
};

// Run the Actor and wait for it to finish
const run = await client.actor("nerolabs/bulk-transcription").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "fileUrl": "",
    "mediaUrls": [
        "https://nerolabs-samples.nerolabs.workers.dev/sample-call-electrician.mp3",
        "https://nerolabs-samples.nerolabs.workers.dev/sample-call-plumber.mp3",
    ],
    "mode": "detailed",
}

# Run the Actor and wait for it to finish
run = client.actor("nerolabs/bulk-transcription").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "fileUrl": "",
  "mediaUrls": [
    "https://nerolabs-samples.nerolabs.workers.dev/sample-call-electrician.mp3",
    "https://nerolabs-samples.nerolabs.workers.dev/sample-call-plumber.mp3"
  ],
  "mode": "detailed"
}' |
apify call nerolabs/bulk-transcription --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nerolabs/bulk-transcription"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HsrY5zhd26m1Ytxhu/builds/TN6hUqfMtyNRjd11p/openapi.json
