Speech to Text & Audio Transcription: $0.0025/min, No Login avatar

Speech to Text & Audio Transcription: $0.0025/min, No Login

Pricing

from $2.50 / 1,000 audio minute (english, standard)s

Go to Apify Store
Speech to Text & Audio Transcription: $0.0025/min, No Login

Speech to Text & Audio Transcription: $0.0025/min, No Login

$0.0025 per audio minute in English. Transcribe audio and video files and podcast RSS episodes to text with timestamps and SRT or VTT subtitles. High-accuracy English $0.003 and about 100 other languages $0.01 per minute. No API key or login. Failed or silent files are free.

Pricing

from $2.50 / 1,000 audio minute (english, standard)s

Rating

0.0

(0)

Developer

Don Mangu

Don Mangu

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Audio and Podcast Transcription

Audio and Podcast Transcription turns audio and video files into text. Give it direct links to your own recordings, or podcast RSS feeds, and it returns one row per file with the full transcript, the spoken language, the length, and timestamped segments. It also saves SRT and VTT subtitle files for each transcript. It costs $0.0025 per audio minute in English ($0.003 with high accuracy) and $0.01 per minute in about 100 other languages, with no API key and no login.

The speech recognition runs inside the Actor with open-weights models. Your audio is not sent to an outside transcription service, and nothing is kept after the run apart from the results in your own Apify storage.

What the audio to text transcription returns

  • Transcript: text, the full transcript, and wordCount.
  • Timestamped segments: segments, a list of start and end (in seconds) and text, one per spoken phrase. Turn off Include timestamped segments to leave them out.
  • Subtitles: srtUrl and vttUrl link to SubRip (.srt) and WebVTT (.vtt) files in the run's key-value store, ready for video players and editors.
  • File details: language, durationSeconds, transcribedSeconds, speechSeconds, audioCodec, fileSizeBytes and the final fileUrl after redirects.
  • Podcast details for feed episodes: feedTitle, episodeTitle and episodePublishedAt, as the feed publishes them.
  • Status for every file: status is ok, or says why there is no transcript (for example not_found, not_media, no_speech, file_too_large), with a plain-language error.

How to transcribe a podcast or audio file, step by step

  1. Open the Actor and go to the Input tab.
  2. In Audio and video file links, paste direct links to files, one per line. MP3, M4A, AAC, WAV, FLAC, OGG, OPUS, MP4, MOV, WEBM and MKV work. A link to a web page that plays audio is not a file link: open the page, find the download or file link, and use that.
  3. To transcribe a podcast, paste its RSS feed link in Podcast RSS feeds and set Episodes per feed (the newest episodes are taken first).
  4. Pick the Language. For English, pick English accuracy: Standard (lowest price) or High (fewer errors on real-world recordings). For any other language, pick it, or pick Detect automatically.
  5. Choose the Subtitle files you want (SRT, VTT or both).
  6. Click Start. The example input transcribes an 11-second clip in a few seconds.
  7. Open the Output tab. The Overview view shows each transcript with its subtitle links; the Segments view shows one row per timestamped segment; the Podcast episodes view shows feed and episode details. Download as JSON, CSV or Excel, or read the results through the Apify API.

How much does audio transcription cost?

You pay per minute of audio transcribed, counted for each file and rounded up to the next whole minute:

SettingEventPrice per audio minute
English, Standard accuracyaudio-minute$0.0025
English, High accuracyaudio-minute-high-accuracy$0.003
Any other language, or Detect automaticallyaudio-minute-multilingual$0.01

Files that could not be downloaded or decoded, links that are not audio or video, and files with no speech return a row with the reason and are not charged.

Worked example: a podcast feed with 4 English episodes of 32, 45, 51 and 38 minutes, plus one broken link, costs (32 + 45 + 51 + 38) x $0.0025 = 166 x $0.0025 = $0.42 at Standard accuracy, or 166 x $0.003 = $0.50 at High accuracy. The broken link costs nothing. The same 166 minutes in Spanish cost 166 x $0.01 = $1.66.

Set Maximum minutes per file to transcribe only the start of long files, and set a spending limit on the run: the Actor transcribes only the minutes that fit, and the row says the file was cut (truncated is true).

Input example

{
"audioUrls": ["https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac"],
"rssFeedUrls": [],
"episodesPerFeed": 1,
"maxFiles": 1,
"language": "en",
"englishAccuracy": "high",
"subtitleFormats": ["srt", "vtt"],
"includeSegments": true,
"maxMinutesPerFile": 240,
"maxFileSizeMb": 1024
}
FieldMeaning
audioUrlsDirect links to audio or video files.
rssFeedUrlsPodcast RSS feed links.
episodesPerFeedNewest episodes to transcribe from each feed (1 to 100).
maxFilesMost files to transcribe in the run (1 to 1,000). File links come first, then feed episodes.
languageen for English, another language code, or auto to detect it.
englishAccuracystandard (lowest price) or high (English only).
subtitleFormatssrt, vtt, both, or none.
includeSegmentsAdd timestamped segments to each row.
maxMinutesPerFileTranscribe at most this many minutes of each file (1 to 720).
maxFileSizeMbLargest file to download, in MB (1 to 4,096).

Output example

A row from a test run with the example input, High accuracy (the 11-second clip is the "Ask not" passage of the 1961 US presidential inaugural address):

{
"inputUrl": "https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac",
"fileUrl": "https://raw.githubusercontent.com/openai/whisper/main/tests/jfk.flac",
"source": "url",
"status": "ok",
"language": "en",
"durationSeconds": 11.0,
"transcribedSeconds": 11.0,
"speechSeconds": 8.02,
"billedMinutes": 1,
"truncated": false,
"text": "And so my fellow Americans. Ask not. What your country can do for you? Ask what you can do for your country.",
"wordCount": 22,
"segments": [
{ "start": 0.29, "end": 2.29, "text": "And so my fellow Americans." },
{ "start": 3.27, "end": 4.5, "text": "Ask not." },
{ "start": 5.38, "end": 7.8, "text": "What your country can do for you?" },
{ "start": 8.17, "end": 10.55, "text": "Ask what you can do for your country." }
],
"srtUrl": "https://api.apify.com/v2/key-value-stores/<store id>/records/transcript-0001.srt",
"vttUrl": "https://api.apify.com/v2/key-value-stores/<store id>/records/transcript-0001.vtt",
"model": "English, high accuracy",
"audioCodec": "flac",
"charged": true,
"error": null
}

The SRT file for the same clip:

1
00:00:00,290 --> 00:00:02,290
And so my fellow Americans.
2
00:00:03,270 --> 00:00:04,500
Ask not.

Use cases

  • Podcasters: show notes, blog posts and searchable archives from your episodes; subtitles for video versions.
  • Content and marketing teams: quotes and clips from webinars, interviews and recorded talks you own.
  • Researchers: text from interview recordings and public-domain speech collections.
  • Developers: a transcription step in a pipeline, called through the Apify API or scheduled to pick up new podcast episodes.

Other Actors from the same developer turn public sources into clean datasets: company job boards, SEC filings and Form D funding rounds, public tenders, weather observations and website technology lookups. Search the Apify Store for them if your pipeline needs more data next to your transcripts.

FAQ

Is it legal to transcribe audio with this Actor? Transcribe files you own or have the right to use: your own recordings, public-domain audio, and podcast episodes for your own listening, research or accessibility needs. The Actor downloads only the files and feeds you list, reads robots.txt on every host first and skips files it disallows, and never logs in. Republishing someone else's podcast transcript can need the owner's permission; check the podcast's terms.

Does it work with YouTube, Spotify, Apple Podcasts or social media links? No. It needs a direct link to an audio or video file, or a podcast RSS feed. Page links from video and social platforms are not supported. For a podcast, use its RSS feed: the Actor reads the episode file links from it.

Which languages are supported? English has its own models (Standard and High accuracy). Other languages (about 100, including Spanish, French, German, Portuguese, Italian, Dutch, Polish, Russian, Hebrew, Arabic, Hindi, Japanese, Korean and Chinese) use the multilingual model. Pick the language when you know it; Detect automatically finds it from the audio.

How accurate is it, and Standard or High? On a clean 11-minute English recording, both English settings got about 95 to 97 words in 100 right, not counting punctuation and capital letters. On short real-world recordings, High accuracy was clearly better: on 5 test clips it made no errors that changed the meaning, while Standard misheard words in 2 of them (for example "As not" for "Ask not"). Use Standard for clear, single-speaker audio and when price matters most; use High for interviews, phone-quality audio and anything you will publish. Accuracy drops with background music, crosstalk, heavy accents and poor microphones. There is no speaker labelling (who said what).

How long does it take? At the default 4 GB memory, an hour of English audio takes about 3 minutes of processing at Standard accuracy and about 5 to 6 minutes at High accuracy; other languages take about 17 to 20 minutes per hour. Downloading adds a little. More memory gives more CPU and a faster run; the price per minute stays the same.

Why does a file show robots_disallowed or blocked? The host's robots.txt does not allow that address, or the server refused the download (HTTP 401 or 403). The Actor does not retry or work around it, and the row is free. After two refusals from the same host, the rest of that host's files in the run are skipped.

What are the limits? Up to 1,000 files per run, files up to 4,096 MB, and up to 720 minutes per file. Very long files are best split across runs with Maximum minutes per file.

Is my audio stored? Downloaded files are deleted as soon as each one is transcribed. Only the transcript rows, subtitle files and run statistics stay, in your own Apify storage.