Audio & Video Transcriber: OpenAI and Gemini Speech to Text avatar

Audio & Video Transcriber: OpenAI and Gemini Speech to Text

Pricing

from $1.60 / 1,000 transcripts

Go to Apify Store
Audio & Video Transcriber: OpenAI and Gemini Speech to Text

Audio & Video Transcriber: OpenAI and Gemini Speech to Text

Transcribe audio and video from any link: files, media pages, podcast feeds, Google Drive and Dropbox. Gemini 3.5 Transcribe is built in, so no API key is needed; OpenAI GPT Transcribe and Whisper run with your own key. Speaker labels, timestamps, SRT and VTT subtitles and translation.

Pricing

from $1.60 / 1,000 transcripts

Rating

0.0

(0)

Developer

Stan Van Rooy

Stan Van Rooy

Maintained by Community

Actor stats

4

Bookmarked

49

Total users

0

Monthly active users

a day ago

Last modified

Share

Audio & Video Transcriber: OpenAI and Gemini Speech to Text 🎙️

Free-plan file sample: Apify Free accounts can process up to 5 files per run, including files expanded from podcast feeds, whether using your own key or built-in access. Built-in access also retains its 5 audio minutes total per run limit. Upgrade your Apify plan for larger runs. Normal Actor charges apply. The sample resets each run; it is not a daily or monthly quota.

Free-plan keyless sample: Apify Free accounts can transcribe up to 5 audio minutes total per run using the built-in Gemini access. Each file is rounded up to a whole minute. Files exceeding the remaining allowance are skipped before transcription, not truncated. Upgrade your Apify plan, or provide your own API key, for longer files and larger runs. Normal Actor charges still apply. The allowance resets each run; it is not a daily or monthly quota.

Turn any audio or video link into text. Paste links to audio and video files, media pages, podcast RSS feeds or Google Drive and Dropbox files, and get a clean transcript per file, with optional speaker labels, timestamps, SRT and VTT subtitles and a translation.

No API key needed: Google Gemini (Gemini 3.5 Transcribe and general Gemini models) is built in, and you pay per audio minute. OpenAI's speech models (GPT Transcribe, GPT-4o Transcribe, GPT-4o mini Transcribe, GPT-4o Transcribe Diarize, Whisper) run with your own OpenAI key, and you can bring your own Gemini key too. With your own key you pay only a small fee per transcript here.

🚀 How to transcribe audio and video

  1. Paste one or more links in URLs.
  2. Keep the recommended model (Gemini 3.5 Transcribe, no key needed), or pick another one. Turn on Speaker labels, Timestamps or Translate to if you need them.
  3. Optional: add your own Gemini API key, or your OpenAI API key to use the OpenAI models. Without a key, Gemini is used and charged per audio minute.
  4. Click Start. Each file becomes one row with the full text, and TXT, SRT, VTT or JSON files are saved for download.

Long recordings are handled for you: the audio is converted to compact mono speech audio and split at natural pauses, the parts are transcribed in parallel, and the timestamps are joined back into one timeline.

  • Audio and video files: MP3, M4A, WAV, FLAC, OGG, OPUS, AAC, MP4, MOV, MKV, WEBM, AVI and most other formats, up to 2 GB per file.
  • Media pages: TikTok, X (Twitter), Instagram, Facebook, SoundCloud, Loom, Twitch VODs, Vimeo player links, Apple Podcasts, archive.org and many more sites.
  • YouTube works when YouTube allows it. YouTube often asks cloud servers to sign in to prove they are not a bot. The actor does not get around that check, so such a video fails with a clear message instead.
  • Podcast RSS feeds: the latest episodes (set how many with Episodes per podcast feed).
  • Google Drive and Dropbox: share links of single files that are shared with "Anyone with the link".

Private, login-only, paywalled or DRM protected media (for example Spotify or Netflix) and live streams are not supported.

🤖 Models

ModelProviderBest forTimestampsSpeaker labelsWithout a key
Gemini 3.5 Transcribe (default)GoogleSpeakers plus word timing, smart formattingSegments and wordsYes (up to 8)Yes, standard
Gemini 3.5 Flash-LiteGoogleTranscribe and translate in one pass, cheapApproximate segmentsNoYes, standard
Gemini 3.8 FlashGoogleTranscribe and translate in one passApproximate segmentsNoYes, premium
Gemini 3.1 Pro PreviewGoogleHard audio, heavy accentsApproximate segmentsNoYes, premium
GPT Transcribe (OpenAI default)OpenAIOpenAI's recommended model, many languagesNoNoOwn OpenAI key
GPT-4o mini TranscribeOpenAIFast and cheapNoNoOwn OpenAI key
GPT-4o TranscribeOpenAIAccurate transcriptsNoNoOwn OpenAI key
GPT-4o Transcribe DiarizeOpenAIWho said what, named voices from clipsSegmentsYesOwn OpenAI key
WhisperOpenAISubtitles, word timing, English translationSegments and wordsNoOwn OpenAI key

You do not have to remember this table: when you turn on Speaker labels or Timestamps, or ask for SRT/VTT files, the actor switches to a model that supports it and says so in the log.

🔑 No key needed, or bring your own key

Without an API key (Gemini built in). Leave both key fields empty and keep a Gemini model. The actor uses its built-in Gemini access and charges per audio minute: $0.01 per minute for Gemini 3.5 Transcribe and Gemini 3.5 Flash-Lite and $0.02 per minute for Gemini 3.8 Flash and Gemini 3.1 Pro Preview (Silver plan 10% off, Gold and above 20% off). Minutes are rounded up per file and charged only after the file was transcribed; a failed file is never charged per minute. Limits: 4 hours per file and 600 audio minutes per run by default (Max audio minutes per run).

OpenAI models need your own OpenAI key. GPT Transcribe, GPT-4o Transcribe, GPT-4o mini Transcribe, GPT-4o Transcribe Diarize and Whisper run only with your key in openaiApiKey. If you pick one of them without a key, the run stops right away with a clear message and nothing is charged for minutes or results.

With your own API key. Add your OpenAI key (openaiApiKey) or Gemini key (geminiApiKey). You pay OpenAI or Google directly at their list prices (for example about $0.0045 per minute for GPT Transcribe, $0.003 for GPT-4o mini Transcribe, about $0.005 for Gemini 3.5 Transcribe), and this actor only charges its per-transcript fee. No per-file length limit beyond the 2 GB download limit. Keys are stored encrypted and never written to the log.

📥 Input

FieldTypeNotes
urlsarrayLinks to files, media pages, podcast feeds, Google Drive or Dropbox files. Required.
providerstringopenai or gemini. Leave empty to choose automatically: Gemini without a key, or the provider of the key you enter.
modelstringSee the table above. Empty: gemini-3.5-transcribe (no key needed), or gpt-transcribe when you only enter an OpenAI key. Picking a model also picks its provider.
openaiApiKeystringSecret. Your OpenAI key, needed for the OpenAI models.
geminiApiKeystringOptional, secret. Your Gemini (Google AI Studio) key. Empty: the built-in Gemini access, charged per minute.
languagestringOptional spoken language, like en, de or pt-BR. Empty: automatic detection.
promptstringOptional context, like the topic or the names of the speakers.
keywordsarrayOptional names and terms to spell exactly.
smartFormattingbooleanGemini 3.5 Transcribe only: removes filler words and formats numbers and lists.
diarizebooleanSpeaker labels. Default false.
knownSpeakerNamesarrayNames instead of Speaker 1, Speaker 2.
knownSpeakerReferencesarrayOpenAI only: one 2 to 10 second clip per name, same order, up to 4.
timestampsstringnone (default), segment or word.
translateTostringOptional target language, like en, de or Spanish.
outputFormatsarrayFiles to save: txt (default), srt, vtt, json.
maxEpisodesPerFeedintegerLatest episodes per podcast feed. Default 1.
maxMinutesPerRunintegerWithout your own key: audio minutes per run. Default 600.
maxConcurrencyintegerFiles processed in parallel. Default 3.
maxRetriesintegerRetries for temporary download or API errors. Default 2.

Minimal input:

{
"urls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"]
}

Inputs saved with the first version of this actor (video_urls, openai_api_key, openai_model and the other openai_* fields) keep working, and their rows still include download_url and transcription.

📤 Output

One row per file. Example with speaker labels and segment timestamps:

{
"url": "https://example.com/podcast/episode-42.mp3",
"input_url": "https://example.com/podcast/feed.xml",
"source_title": "Episode 42: Building in public",
"duration_sec": 2412.6,
"provider": "gemini",
"model": "gemini-3.5-transcribe",
"language": "en",
"text": "Welcome back to the show. Today my guest is ...",
"segments": [
{"start": 0.1, "end": 3.9, "text": "Welcome back to the show.", "speaker": "Host"},
{"start": 4.2, "end": 7.8, "text": "Thanks for having me.", "speaker": "Guest"}
],
"words": [],
"speakers": ["Host", "Guest"],
"translation": null,
"translation_language": null,
"files": {
"txt": "https://api.apify.com/v2/key-value-stores/.../records/001-Episode-42-Building-in-public.txt",
"srt": "https://api.apify.com/v2/key-value-stores/.../records/001-Episode-42-Building-in-public.srt"
},
"api_access": "built_in",
"billed_minutes": 41,
"cost_estimate_usd": 0.41,
"status": "succeeded",
"error": null,
"processed_at": "2026-09-25T10:15:02+00:00"
}

Field descriptions

  • url: The file or page that was transcribed (for podcast feeds, the episode's audio file).
  • input_url: The URL from your input that led to this file.
  • source_title: Page or episode title, when known.
  • duration_sec: Audio length in seconds.
  • provider and model: Who made the transcript.
  • language: Detected or given language code, when the model reports one.
  • text: The full transcript.
  • segments: Sentences or phrases with start and end in seconds, and speaker with speaker labels.
  • words: Word timestamps, with timestamps set to word.
  • speakers: Speaker labels or names in order of appearance.
  • translation and translation_language: The translation, with translateTo.
  • files: Download links of the saved TXT, SRT, VTT and JSON files.
  • api_access: own_key or built_in.
  • billed_minutes: Audio minutes charged by this actor (built-in access only).
  • cost_estimate_usd: Estimated transcription cost of this file: the provider's list price with your own key, or this actor's per-minute charge with the built-in access. The per-transcript fee is not included.
  • status: succeeded, failed or skipped, and error says why.

The Segments view of the dataset lists every timed segment as its own row, handy for spreadsheets.

💡 Use cases

  • Podcasts: transcripts and show notes material for the latest episodes, straight from the RSS feed.
  • Meetings and interviews: who said what, with speaker names, from a Drive or Dropbox recording.
  • Subtitles: SRT and VTT files for social videos, courses and webinars.
  • Research and monitoring: make spoken content from social media and news clips searchable.
  • Translation: an English (or any other language) version of foreign language audio.
  • AI pipelines: clean text and timestamps for summaries, search and LLM workflows.

💰 Pricing

This actor uses pay per event: no monthly fee, you pay for what you use.

  • Transcript: $0.002 per result row (Silver $0.0018, Gold and above $0.0016). Every URL gets a row, also when it fails, so you can see why.
  • Audio minutes without your own key (Gemini): $0.01 per minute on Gemini 3.5 Transcribe and Gemini 3.5 Flash-Lite, $0.02 per minute on Gemini 3.8 Flash and Gemini 3.1 Pro Preview (Silver 10% off, Gold and above 20% off). Not charged when you use your own key.
  • Actor start: $0.005 per GB of memory (Silver $0.0045, Gold and above $0.004), about $0.01 for a default 2 GB run.

Examples: a 1 hour podcast without any key on Gemini 3.5 Transcribe costs $0.60 for the minutes plus $0.002 per transcript and the start fee. With your own OpenAI key it costs $0.002 plus the start fee here, plus about $0.27 at OpenAI for GPT Transcribe. New Apify accounts get monthly free usage credit to try it.

🎯 Example inputs

A short clip, no API key needed:

{
"urls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"],
"model": "gemini-3.5-transcribe"
}

Latest 3 episodes of a podcast with speaker labels and subtitles:

{
"urls": ["https://feeds.example.com/my-podcast.xml"],
"maxEpisodesPerFeed": 3,
"diarize": true,
"outputFormats": ["txt", "srt"]
}

A meeting on Google Drive with named speakers on Gemini:

{
"urls": ["https://drive.google.com/file/d/FILE_ID/view?usp=sharing"],
"provider": "gemini",
"diarize": true,
"knownSpeakerNames": ["Anna", "Ben"],
"timestamps": "word",
"outputFormats": ["txt", "vtt", "json"]
}

A Spanish video with an English translation, on OpenAI with your own key:

{
"urls": ["https://www.tiktok.com/@creator/video/1234567890"],
"language": "es",
"translateTo": "en",
"openaiApiKey": "sk-..."
}

🔌 Run it from code

Every Apify Actor is also an API. With the Apify Python client:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("stanvanrooy6/audio-video-transcriber").call(run_input={
"urls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"],
"model": "gemini-3.5-transcribe", # no API key needed
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row["status"], row["text"])

The same works from JavaScript, cURL, Make, Zapier or n8n. See the API tab for ready-made snippets.

❓ FAQ

Do I need an OpenAI or Gemini API key?

No. Gemini models work without any key and are charged per audio minute. The OpenAI models (GPT Transcribe, GPT-4o Transcribe, Whisper and the others) need your own OpenAI key. With your own key you pay the provider directly and only the per-transcript fee here, which is cheaper for large volumes.

Which model should I choose?

Gemini 3.5 Transcribe (the default, no key needed) for most jobs, including speakers and word timestamps. With your own OpenAI key: GPT Transcribe (OpenAI's recommended model) for plain transcripts, GPT-4o mini Transcribe when cost matters most, Whisper for subtitles. The general Gemini models can transcribe and translate in one pass; their timestamps are approximate.

YouTube often shows cloud servers a "Sign in to confirm you're not a bot" page. The actor does not log in or get around that check, so the video fails with a clear message. Try again later, or use a direct link to the file. Other sites can block automated downloads the same way.

How long can a file be?

With your own key there is no length limit beyond the 2 GB download limit; long files are split into parts and joined back. Without a key the limit is 4 hours per file and Max audio minutes per run in total (600 by default).

Are speaker labels consistent in very long recordings?

Speakers are recognised per part of about 20 minutes (OpenAI) or 29 minutes (Gemini). In longer recordings the same person can get a different label in another part, unless you give OpenAI reference clips of the voices.

Which languages are supported?

Most spoken languages on OpenAI, and more than 85 locales on Gemini 3.5 Transcribe. Setting Audio language improves accuracy and helps with short clips.

What happens to my audio and my key?

The audio is sent only to the provider you chose. Files uploaded to Gemini are deleted right after each request, and requests are sent with storage turned off. Temporary files are deleted after each file. Your API key is a secret input: stored encrypted and never shown in the log.

Can you add a site, a format or a feature?

Yes. Open an issue on the Issues tab and describe what you need. I actively maintain this actor.

🤝 Feedback and support

Found a bug or missing a feature? Open an issue on the Issues tab and I will get back to you.

  • TikTok Transcript Scraper: the spoken text of TikTok videos, whole profiles or keyword searches from TikTok's own captions, with timestamps. No login and no API key.

Built with ❤️ for the Apify community