Bulk Transcription: Audio & Video to Text from CSV or Sheet avatar

Bulk Transcription: Audio & Video to Text from CSV or Sheet

Pricing

from $7.00 / 1,000 audio minute, plain texts

Go to Apify Store
Bulk Transcription: Audio & Video to Text from CSV or Sheet

Bulk Transcription: Audio & Video to Text from CSV or Sheet

Transcribes every audio or video link in an Apify dataset, CSV, Excel or Google Sheet, keeping your columns. Inputs: datasetId or fileUrl or mediaUrls, urlField, mode (plain text, or speakers with timestamps and SRT). Charged per started minute of audio. Agent-ready: x402, MCP.

Pricing

from $7.00 / 1,000 audio minute, plain texts

Rating

0.0

(0)

Developer

Adam Pearce

Adam Pearce

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Bulk Transcription: Audio & Video to Text from CSV, Google Sheet or Dataset

Got a spreadsheet full of call recordings, podcast episodes, voicemails or meeting videos and need the words, not the audio? Point this Actor at the sheet and it transcribes every link in it, then hands the sheet back with a transcript next to each row. Your own columns (call ID, date, client, episode title) stay exactly where they were.

  • Bulk, from a list you already have. An Apify dataset, a CSV, an Excel file, a Google Sheet link, or a plain list of URLs. No copying links into a form one at a time.
  • Keeps your columns. Every row comes back with its original fields plus the transcript, so results line up with your CRM, call log or content calendar.
  • Speakers, timestamps and subtitles. Switch on "Speakers, timestamps and subtitles" and each transcript reads like a conversation (Speaker A: ..., Speaker B: ...), with start and end times and a ready-made SRT subtitle file per recording.
  • Any audio or video. MP3, WAV, M4A, MP4, MOV, WebM, OGG, FLAC, AAC and more. Video files are fine: only the audio is used. Google Drive and Dropbox share links are converted to direct downloads automatically.
  • Long recordings just work. Files are cut into pieces and transcribed in parallel; a 13-minute recording comes back in about 15 seconds in plain text mode.
  • You only pay for what was transcribed. Broken links, web pages, silent files and failed downloads are reported and never charged.

What it is good for

  • Call recordings and voicemails from phone systems, AI receptionists and sales dialers (GoHighLevel, Vapi, Retell, Twilio, Aircall, CallRail exports).
  • Podcast and YouTube back catalogues for show notes, search and repurposing into posts.
  • Interviews, meetings and webinars for notes and quotes.
  • AI agent pipelines: feed transcripts straight into an LLM, a vector database or another Actor.

Want each call scored as well (booked or not, missed questions, lead quality, a 0 to 10 score)? Use the sister Actor Call Score.

Input

Give it one of:

InputWhat to put in
DatasetAny Apify dataset with a column of audio or video links, for example a scraper's output.
File or Google Sheet URLA link to a CSV, Excel, JSON or JSON Lines file, or a normal Google Sheets link shared as "Anyone with the link can view".
Audio or video URLsA plain list of links for a quick run.

The column holding the links is found automatically (it prefers values ending in .mp3, .wav, .m4a, .mp4 and names like recordingUrl or audio); set Recording URL field to choose it yourself.

Transcript type

TypeYou getPrice per started minute
Plain textOne clean block of text per file. Accepts Spelling hints for names and jargon.$0.01
Speakers, timestamps and subtitlesOne line per speaker turn, timed segments (optional) and SRT subtitles.$0.02

Other useful settings: Language (two-letter code, or leave empty to detect), Max minutes per file (your cost ceiling per file, default 120), Export files (a downloadable CSV or Excel of the results), Webhook URL (POSTs the run summary to Zapier, Make, n8n, Slack or your API), and Append to named dataset (builds one growing archive across runs).

Output

One row per recording. Example (speakers mode, shortened):

{
"callId": "C-1001",
"client": "Demo Electrical",
"recordingUrl": "https://nerolabs-samples.nerolabs.workers.dev/sample-call-electrician.mp3",
"transcriptStatus": "ok",
"mode": "detailed",
"mediaType": "audio",
"durationSec": 64.8,
"durationMinutes": 1.08,
"minutesCharged": 2,
"speakerCount": 2,
"wordCount": 208,
"text": "Speaker A: Hello, you've reached Demo Electrical. This is the virtual assistant. How can I help you today?\nSpeaker B: Hi, yeah, I've got no power to half the house...",
"srt": "1\n00:00:00,000 --> 00:00:04,350\nHello, you've reached Demo Electrical. This is the virtual assistant.\n..."
}

transcriptStatus is one of ok, no_speech (read end to end, no words: silence or music), no_audio (a video with no sound track), not_supported, http_error, unreachable, blocked, too_large, invalid_url, no_url, transcription_failed or skipped_budget. Only ok is charged. statusDetail explains every other status in plain words.

The run's OUTPUT record holds the totals: files transcribed, minutes of audio, minutes charged, words, status counts and export links.

Pricing

Pay per event, no subscription:

  • $0.01 per started minute, plain text.
  • $0.02 per started minute, speakers, timestamps and subtitles.
  • $0.01 per CSV or Excel export file, $0.02 per delivered webhook. A tiny start fee of $0.00005 per run.
  • Apify subscribers on Bronze, Silver and Gold plans get 10%, 20% and 30% off.

Worked examples: 100 sales calls of 4 minutes each in plain text is 400 minutes, about $4. A 50-episode podcast back catalogue at 45 minutes an episode, with speakers and subtitles, is 2,250 minutes, about $45. Each file is rounded up to the next whole minute, so a 30-second voicemail costs one minute.

Set Max minutes per file and Apify's own maximum charge per run to cap any run exactly.

How it works

Each file is downloaded inside the run (streamed to disk, so large videos are fine), converted to mono audio with ffmpeg, cut into pieces and sent to OpenAI's speech models: gpt-4o-mini-transcribe for plain text and gpt-4o-transcribe-diarize for speakers and timestamps. Only the audio is sent and nothing is kept after the run.

FAQ

Can it transcribe a YouTube or TikTok page link? No. It needs a direct link to the audio or video file. Use a downloader Actor first and feed its dataset in here.

How accurate are the speaker labels? Very good on two-person calls and interviews. The letters (A, B, C) are the model's own and stay consistent within each 5-minute section of a long recording; on a long multi-speaker recording a person can get a different letter in a later section.

Which languages? About 57, including English, Spanish, French, German, Portuguese, Italian, Dutch, Polish, Japanese and Chinese. Language is detected automatically; setting it helps on short or noisy clips.

Is my audio stored? No. Files are processed inside your run and deleted at the end of it; the speech provider receives only the audio of each piece.

Can an AI agent use it? Yes. It charges per event, so agents can call it through the Apify MCP server or pay per run with x402.

If this saved you hours of typing up recordings, a review on the Store helps a lot. Questions or a format that will not read? Open an issue on the Issues tab and I will look at it the same day.