Audio & Video Transcriber (Whisper) with Podcast RSS avatar

Audio & Video Transcriber (Whisper) with Podcast RSS

Pricing

Pay per event

Go to Apify Store
Audio & Video Transcriber (Whisper) with Podcast RSS

Audio & Video Transcriber (Whisper) with Podcast RSS

Transcribe audio, video and podcast RSS episodes with Whisper on Groq or OpenAI using your own API key. Get text, SRT and VTT subtitles and timestamped JSON segments.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Rod Services

Rod Services

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

What does Audio & Video Transcriber (Whisper) do?

Audio & Video Transcriber turns podcasts, meeting recordings, interviews, lectures, webinars and videos into text with OpenAI Whisper. You get a clean transcript, SRT and WebVTT subtitles, and JSON segments with timestamps for every file.

Paste direct file links, upload a file, or give it a podcast RSS feed and it transcribes the latest episodes. Long recordings are split into chunks automatically, so a 3 hour podcast works the same as a 3 minute voice memo.

It runs on your own Groq or OpenAI API key (bring your own key). Groq runs Whisper large v3 turbo at about $0.04 per audio hour, so a 1 hour podcast costs you about 4 cents at Groq plus $0.003 per minute here.

As an Apify Actor you also get an API, scheduling, webhooks, integrations with Make, Zapier, n8n and LangChain, and run monitoring. AI agents can call it through the Apify MCP server.

Why use this Whisper transcription tool?

  • Podcast transcription from RSS. Point it at any podcast feed. It picks the newest N episodes and adds episode title and publish date.
  • Subtitles in SRT and VTT. Ready for YouTube Studio uploads, video editors, HTML5 players and accessibility.
  • Meeting recordings and interviews. MP4, MOV, MKV and WEBM files work. Only the audio track is used.
  • Timestamps for search and RAG. Every segment has start and end seconds. Link quotes back to the exact moment.
  • AI agents and LLM pipelines. Feed transcripts into summarizers, show notes generators, chatbots or vector databases.
  • Long files handled for you. ffmpeg converts audio to mono 16 kHz and splits it into chunks under the 25 MB provider limit, with overlap so no words are lost at the cuts.
  • Translate to English. One switch gives an English transcript of speech in about 100 languages.
  • Cheap and fast. You pay the provider directly at their rates. There is no markup on the model.

How to transcribe audio or a podcast

  1. Get an API key. Groq keys are free to create at console.groq.com/keys. OpenAI keys are at platform.openai.com/api-keys.
  2. Open the Input tab and paste the key into API key. It is stored encrypted.
  3. Add links in Audio or video URLs, upload a file, or paste a Podcast RSS feed URL.
  4. Optional: set a Language hint such as en, de or lt, and pick Output formats.
  5. Click Start. Transcripts appear in the Output tab. Subtitle files are in the key-value store.
  6. Download the dataset as JSON, CSV or Excel, or call the API from your own code.

Want to check your links first? Turn on Dry run. It downloads, converts and splits the audio without calling the provider and without the per minute charge.

Running without an API key

If you start the Actor without a key, it does not download anything. It finishes successfully and writes one dataset item that explains a key is required. This is also what happens with the example input, so you can try the Actor safely.

Input

All fields are on the Input tab. The main ones:

FieldWhat it does
audioUrlsDirect links to MP3, M4A, WAV, FLAC, OGG, OPUS, AAC, MP4, MOV, MKV, WEBM and more. Google Drive and Dropbox share links work.
rssFeedUrl, maxEpisodesPodcast RSS or Atom feed, and how many of the newest episodes to transcribe.
uploadedFile, keyValueStoreRecordsFiles uploaded in Console, or records in a key-value store.
provider, apiKey, modelgroq or openai, your key, and the model. auto picks the best fit.
languageISO 639-1 language hint. Empty means auto detect.
translateToEnglishReturn an English translation.
outputFormatstext, srt, vtt, json files saved to the key-value store.
maxDurationMinutesTranscribe at most this many minutes per file. Protects you from surprise costs.
dryRunDownload and split only. No key needed.
promptNames, brands and jargon that help the model spell them right.

Example input:

{
"provider": "groq",
"apiKey": "YOUR_GROQ_KEY",
"rssFeedUrl": "https://librivox.org/rss/389",
"maxEpisodes": 3,
"language": "en",
"outputFormats": ["text", "srt", "vtt", "json"]
}

Models

ProviderModelTimestampsTranslationProvider price
Groqwhisper-large-v3-turbo (default)YesNo, switches to v3~$0.04 per hour
Groqwhisper-large-v3YesYes~$0.111 per hour
OpenAIwhisper-1YesYes$0.006 per minute
OpenAIgpt-4o-mini-transcribeNo, estimatedNo$0.003 per minute
OpenAIgpt-4o-transcribeNo, estimatedNo$0.006 per minute

With auto on OpenAI, the Actor uses whisper-1 when you ask for SRT, VTT or JSON, because gpt-4o transcribe models return no timestamps. Check current prices on the provider sites.

Output

One dataset item per file. The example below is shortened. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

{
"sourceUrl": "https://www.archive.org/download/gettysburg_shurtagal_librivox/Gettysburg_Address_Lincoln_64kb.mp3",
"sourceType": "rss",
"title": "Gettysburg Address",
"feedTitle": "Gettysburg Address, The by Abraham Lincoln (1809 - 1865)",
"publishedAt": null,
"durationSeconds": 100.34,
"language": "en",
"text": "Four score and seven years ago our fathers brought forth on this continent a new nation...",
"segments": [
{ "id": 0, "start": 0.0, "end": 6.2, "text": "Four score and seven years ago" },
{ "id": 1, "start": 6.2, "end": 11.8, "text": "our fathers brought forth on this continent a new nation," }
],
"wordCount": 272,
"srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/srt-0000-Gettysburg-Address.srt",
"vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/vtt-0000-Gettysburg-Address.vtt",
"provider": "groq",
"model": "whisper-large-v3-turbo",
"billedMinutes": 2,
"warnings": [],
"error": null
}

Data fields

FieldDescription
sourceUrlMedia URL or RSS enclosure URL.
title, feedTitle, publishedAtEpisode title and date from RSS, or the media title tag, or the file name.
durationSeconds, transcribedSecondsLength of the file and of the part that was transcribed.
languageISO 639-1 code, detected or from your hint.
textFull transcript.
segments{id, start, end, text} with times in seconds.
srtUrl, vttUrl, txtUrl, jsonUrlFiles in the key-value store.
wordCountWords in the transcript.
provider, modelWhat transcribed the file.
timestampsApproximatetrue when the model gives no timestamps and times are estimated.
billedMinutesAudio minutes charged for this file.
warnings, errorWhat went wrong or was changed, in plain words.

The Overview view shows one row per file. The Segments view shows one row per timestamped segment.

How much does it cost to transcribe audio?

This Actor uses pay per event pricing:

  • $0.003 per audio minute transcribed, rounded up per file.
  • A small start fee per run.
  • Dry runs, rejected links and failed files are free of the per minute charge.

The provider bills its own part to your key. Examples with Groq whisper-large-v3-turbo:

AudioThis ActorGroq (approx.)Total
1 hour podcast$0.18$0.04about $0.22
10 episodes of 45 min$0.90$0.30about $1.20
100 hours of meetings$12.00$4.00about $16

Set Maximum cost per run in the run options. The Actor stops taking new minutes when it is reached, and cuts a long file short with a warning.

Tips and advanced options

  • Give a language hint. It avoids wrong language detection on short clips and music intros.
  • Use the vocabulary prompt for names, product names and acronyms.
  • Rate limits. Groq free keys have hourly and daily audio limits. Check them on your Groq limits page. The Actor retries 429 responses with backoff and honours retry-after. Lower Parallel API requests if you see many retries, or use a paid Groq tier.
  • Exact subtitles. Use Groq models or OpenAI whisper-1. gpt-4o transcribe models return text only, so cue times are spread by text length.
  • Chunk length. 10 minutes is a good default. Chunks overlap by 5 seconds and the transcript is stitched at the middle of the overlap.
  • Memory. 1 GB is enough. Audio is processed on disk, not in memory.

FAQ, disclaimers and support

No. Their terms of service forbid downloading media with third-party tools, so these page links are rejected with a message. Use a direct link to a file you own or may use, upload the file, or use the podcast RSS feed.

Is my API key safe?

The key is a secret input. Apify stores it encrypted. The Actor sends it only to the provider API you picked, never logs it, and never writes it to the dataset. We recommend a separate key for this Actor with a spending limit, and revoking it when you are done. You are responsible for using your key within your provider's terms.

Do you store my audio?

Audio is downloaded into the run container, converted, sent to the provider and deleted when the file is done. Transcripts are saved in your own run storage. The provider processes the audio under its own data policy.

Known limitations

  • Speaker labels (diarization) are not included yet.
  • Word counts for Chinese and Japanese count characters.
  • Files behind logins or DRM cannot be downloaded.

Feedback

Found a bug or need a feature? Open an issue on the Issues tab. Custom pipelines such as speaker labels, summaries or delivery to your storage are available on request.