OpenAI whisper-1 Alternative — Speech-to-Text API
Pricing
from $1.50 / 1,000 15 seconds of audios
OpenAI whisper-1 Alternative — Speech-to-Text API
whisper-1 alternative and speech-to-text API for apps that need SRT, VTT, or word timestamps: point the official OpenAI SDK at this host. OpenAI removes whisper-1 on 26 February 2027. $0.006 per minute, billed per started 15 seconds.
Pricing
from $1.50 / 1,000 15 seconds of audios
Rating
0.0
(0)
Developer
drop-in apis
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
3 hours ago
Last modified
Categories
Share
whisper-1 alternative and speech-to-text API that serves POST /v1/audio/transcriptions and /v1/audio/translations with srt, vtt and verbose_json word timestamps, using the official OpenAI SDKs. OpenAI removes whisper-1 on 26 February 2027 (deprecations).
At a glance: $0.006 per minute of audio ($0.0015 per started 15 seconds; failed requests are not charged) · about 4 seconds of audio per second of compute (a 1-minute file returns in about 12 seconds) · srt, vtt and verbose_json output. Try it with no code: click Start with the prefilled sample URL (batch mode, billed at least 1 minute per file).
Last updated: 2026-10-02 · Full migration guide: https://alidaram99.github.io/api-alternatives/whisper-1-alternative/
You change base_url and api_key, and keep the same multipart request and the same json, text, srt, vtt and verbose_json responses, including OpenAI's error format.
- ✅ Same endpoints, form fields and response formats as whisper-1
- ✅ Word and segment timestamps (
timestamp_granularities[]), SRT and VTT subtitles,usagein seconds - ✅ Translation to English (
/v1/audio/translations) - ✅ 99 languages with auto-detection, the same set as Whisper
- ✅ flac, m4a, mp3, mp4, mpeg, mpga, oga, ogg, wav, webm, up to 25 MB and 10 minutes per file
- ✅ $0.006 per minute, billed per started 15 seconds. That is OpenAI's whisper-1 per-minute rate. OpenAI rounds to the second and we round up to 15 s, so very short clips cost slightly more here (see Pricing).
- 🔒 Audio is processed in memory and never stored. No third-party API is called.
Migrate in two lines
Your endpoint is the Actor's Standby URL (on the Actor's Standby tab), for example https://<username>--whisper-compat.apify.actor. Use your Apify API token as the API key. The OpenAI SDKs send it as Authorization: Bearer …, which is how Apify authenticates Standby requests.
Pass the token only as
api_key. Don't append?token=…tobase_url: the SDK adds the route after it, and the request ends up at the wrong URL.
Python (openai)
import osfrom openai import OpenAIclient = OpenAI(base_url="https://<username>--whisper-compat.apify.actor/v1", # was: default api.openai.comapi_key=os.environ["APIFY_TOKEN"], # was: your OpenAI key)with open("meeting.mp3", "rb") as f:result = client.audio.transcriptions.create(model="whisper-1", file=f, response_format="verbose_json",timestamp_granularities=["word", "segment"],)print(result.text)print(result.words[0], len(result.segments), "segments")
Node.js (openai)
import fs from 'node:fs';import OpenAI from 'openai';const client = new OpenAI({baseURL: 'https://<username>--whisper-compat.apify.actor/v1',apiKey: process.env.APIFY_TOKEN,});const srt = await client.audio.transcriptions.create({model: 'whisper-1',file: fs.createReadStream('talk.mp4'),response_format: 'srt',});console.log(srt);
curl
curl -s https://<username>--whisper-compat.apify.actor/v1/audio/transcriptions \-H "Authorization: Bearer $APIFY_TOKEN" \-F model=whisper-1 \-F response_format=vtt \-F file=@interview.m4a
Response examples (real output)
response_format=json:
{ "text": "Mr. Quilter is the Apostle of the Middle Classes, and we are glad to welcome his Gospel. Nor is Mr. Quilter's manner less interesting than his matter.","usage": { "type": "duration", "seconds": 12 } }
response_format=srt (NASA's public-domain Apollo 11 recording):
100:00:00,000 --> 00:00:05,000I'm going to step off the limb now.200:00:15,000 --> 00:00:18,000That's one small step for man.300:00:20,000 --> 00:00:24,000One giant leap for mankind.
response_format=verbose_json returns task, language (e.g. "english"), duration, text and segments[]. Each segment has id, seek, start, end, text, tokens, temperature, avg_logprob, compression_ratio and no_speech_prob. The response also includes words[] (word, start, end) when you ask for word timestamps, and usage. That is the same shape OpenAI documents.
Compatibility
| Feature | Status |
|---|---|
POST /v1/audio/transcriptions, POST /v1/audio/translations, GET /v1/models | ✅ Same paths (also served without the /v1 prefix) |
Fields file, model (whisper-1), language, prompt, response_format, temperature, timestamp_granularities[] | ✅ Same names, defaults and validation |
json (+ usage), text, srt, vtt, verbose_json (segments/words as requested) | ✅ Same shapes |
Error envelope {"error":{"message","type","param","code"}}, 400/413/429 | ✅ Same format. Unsupported file formats return OpenAI's exact message |
| Transcription quality | ⚠️ Runs Whisper small (multilingual) with CTranslate2 int8 and beam search 5. OpenAI's whisper-1 is the much larger Whisper V2. On clean English speech we measured a 3.8% word error rate (LibriSpeech). Expect more errors than whisper-1 on noisy audio, heavy accents and rare languages |
| Files longer than 10 minutes | ❌ Rejected with 400 audio_too_long. Split long recordings into chunks (OpenAI has a 25 MB cap anyway) |
stream, include, chunking_strategy, diarized_json, gpt-4o / gpt-transcribe models | ❌ Not supported (whisper-1 didn't support them either) |
How it compares
| Option | Same whisper-1 request and formats? | Price | Notes |
|---|---|---|---|
| This Actor | Yes: official OpenAI SDK, change base_url and api_key | $0.006/min, billed per started 15 s | Whisper small (about 4% word error on clean English); up to 10 min per file |
| OpenAI whisper-1 | It is the original | $0.006/min | Removed 26 Feb 2027 |
| OpenAI gpt-4o-transcribe / gpt-4o-mini-transcribe | No: OpenAI's API spec lists json as their only response format | OpenAI pricing | No SRT/VTT or word timestamps |
| File-to-text transcriber Actors on Apify | No OpenAI SDK compatibility | Varies by Actor | Fine if you only need text from a file |
Pricing
$0.0015 per started 15 seconds of audio, which is $0.006 per minute (pay-per-event audio-15s). OpenAI billed whisper-1 at $0.006 per minute rounded to the second, so the two prices match on whole 15-second blocks. A clip shorter than 15 s costs $0.0015 here instead of about $0.001 at OpenAI. Examples:
| Audio | Price |
|---|---|
| 10-second voice note | $0.0015 |
| 1-minute clip | $0.006 |
| 10-minute recording | $0.06 |
Failed requests (bad format, too long, can't decode) are not charged. Batch mode (a list of audioUrls, see below) bills at least 1 minute ($0.006) per file, because each batch run starts a new container. Standby API requests have no minimum. Any additional platform charges, such as Apify's per-run start event, are listed on this Actor's Pricing tab and follow Apify's pay-per-event terms.
Spending limit. If you set a maximum total charge on the run, a request whose audio would cost more than the remaining budget is refused before it is transcribed, with 429 insufficient_quota, so no compute is spent on audio your budget can't cover. In batch mode the run stops cleanly at the limit and keeps the files transcribed so far.
Speed and limits
Each instance has 1 CPU core and processes about 4 seconds of audio per second of compute. A 1-minute file returns in about 12 seconds and a 10-minute file in about 2.5 minutes. Requests to one instance are processed one at a time. When an instance is busy, it returns 429 rate_limit_exceeded, which the OpenAI SDKs retry automatically, and Apify can start more instances. After about 5 minutes without requests the instance sleeps, and the next request waits a few extra seconds while it starts.
Batch mode (no code)
Run the Actor normally with a list of audioUrls to get one dataset row per file, containing the transcript, language, duration and full response. You can export it as CSV, Excel or JSON. Each file is billed at least 1 minute. For large batches of audio or video files with SRT/VTT output and a cheaper fast model, see the media URL transcriber.
{ "audioUrls": ["https://upload.wikimedia.org/wikipedia/commons/d/dd/Armstrong_Small_Step.ogg"], "response_format": "srt" }
FAQ
What is a whisper-1 alternative that still returns SRT and VTT?
This Actor. response_format=srt or vtt works with the official OpenAI SDK after you change base_url and api_key.
When is whisper-1 removed?
26 February 2027, per OpenAI's deprecations page: https://developers.openai.com/api/docs/deprecations
Do I still get word-level timestamps?
Yes. Send timestamp_granularities[]=word (and/or segment) with response_format=verbose_json, exactly as with whisper-1.
How accurate is it compared with OpenAI whisper-1?
It runs the open Whisper small model. We measured about 4% word error rate on clean English (LibriSpeech). Noisy audio, strong accents and rare languages do worse than OpenAI's larger model.
How long can the audio be?
Up to 10 minutes and 25 MB per request. Longer files are rejected with 400 audio_too_long; split them first.
What is the price?
$0.0015 per started 15 seconds of audio, which is $0.006 per minute. A clip under 15 seconds costs $0.0015. Failed requests are not charged.
Which languages are supported?
All 99 Whisper languages, with auto-detection. Passing language (ISO-639-1) improves speed and accuracy.
Do you keep my audio?
No. Uploaded audio is decoded in memory, transcribed and discarded. Nothing is written to disk or logged.
Is this affiliated with OpenAI?
No. It is an independent service that implements the same public API shape and runs OpenAI's open-source Whisper model (MIT license) through faster-whisper.