YouTube Transcripts + Speech AI
Pricing
from $1.20 / 1,000 transcripts
YouTube Transcripts + Speech AI
Never charged for a video we can't transcribe. YouTube videos, Shorts and live VODs to text — captions first, built-in speech-to-text when a video has none, so caption-less videos still return real text. JSON, plain text, SRT or VTT. No API key, no cookies.
Pricing
from $1.20 / 1,000 transcripts
Rating
0.0
(0)
Developer
Steadyfetch Team
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
37 minutes ago
Last modified
Share
Never charged for a video we can't transcribe. Captions first, built-in speech-to-text when a video has none — so caption-less videos still return real text. Videos, Shorts and live VODs, as JSON, plain text, SRT or VTT.
Issues answered in about 3 hours. Unofficial: steadyfetch is not affiliated with, endorsed by, or sponsored by YouTube or Google. "YouTube" is a trademark of Google LLC, used here only to say what this actor reads.
Every row carries charged and statusReason, so you can reconcile the invoice from the
dataset itself without opening the console. Only rows with charged: true were billed.
What a row looks like
Real output, unedited apart from trimming the segment list:
{"status": "ok","charged": true,"statusReason": null,"videoId": "jNQXAC9IVRw","url": "https://www.youtube.com/watch?v=jNQXAC9IVRw","inputUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw","title": "Me at the zoo","channelName": "jawed","channelId": "UC4QobU6STFB0P71PMvOGN5A","durationSeconds": 19,"viewCount": 406463143,"isLiveContent": false,"language": "en","source": "captions","captionKind": "manual","text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say","segments": [{ "start": 1.2, "end": 3.36, "text": "All right, so here we are, in front of the elephants" },{ "start": 5.318, "end": 7.974, "text": "the cool thing about these guys is that they have really..." }],"srt": null,"vtt": null,"chargeEvents": { "transcript": 1, "speechMinutes": 0 }}
| field | notes |
|---|---|
text · segments[{start,end,text}] | the transcript, and the same text with timestamps in seconds |
srt · vtt | filled only when you ask for that format; otherwise null |
source | captions or speech_ai — how this transcript was produced |
captionKind | manual (uploaded by the channel) or auto (YouTube's own auto-captions); null on the speech route |
language | ISO code, same vocabulary on both routes |
title · channelName · channelId · durationSeconds · viewCount · isLiveContent | as YouTube reports them |
url · inputUrl · sourceIndex | the canonical watch URL, the exact link you passed, and its position in your input |
charged · statusReason · chargeEvents | the reconciliation trio |
retryable | on a row that did not deliver: true means YouTube refused us this time and the same input is worth running again, false means the answer will not change. It always agrees with the note in statusReason, so code can branch on the column instead of parsing the sentence |
Sample dataset — six real transcripts from one run, no sign-in: open the JSON
Agent / API paste-block
Actor: steadyfetch/youtube-transcript-scraperRequired: videoUrls (array of YouTube video links or 11-character video IDs)Optional: language (string, e.g. "en", "es", "pt-BR" — preferred caption track)format (json | text | srt | vtt, default json)enableSpeechFallback (boolean, default true — off = captions only, no speech minutes)maxSpeechMinutes (integer, default 60 — hard cap for the whole run)maxItems (integer, default 100 — hard cap on videos)Charges: transcript once per delivered transcript, captions or speech-to-textspeech_minute per started minute, only when speech-to-text actually ranBuild spec: https://apify.com/steadyfetch/youtube-transcript-scraper/apiToken: https://console.apify.com/settings/integrations
curl -X POST "https://api.apify.com/v2/acts/steadyfetch~youtube-transcript-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H 'Content-Type: application/json' \-d '{"videoUrls":["https://www.youtube.com/watch?v=jNQXAC9IVRw","https://youtu.be/dQw4w9WgXcQ"],"format":"srt"}'
Calling from an agent or MCP client: omit an option you do not want rather than sending
null — the platform rejects an explicit null before the run is even created.
Links you can paste
Watch links, youtu.be links, /shorts/, /live/, /embed/, m.youtube.com,
music.youtube.com, youtube-nocookie.com, and bare 11-character video IDs. Casing in the
link does not matter — chat apps and link shorteners rewrite it constantly — but the video
ID itself is case-sensitive, because YouTube treats it that way.
A channel link gets an uncharged row pointing you at steadyfetch/youtube-channel-transcripts, which does whole channels in one run. A playlist link gets an uncharged row too — paste the individual video links here instead.
What you are charged for
Pricing: from $1.20/1,000 transcripts. Two events, and nothing else:
- Transcript — once per video that returns real text, from captions or from speech-to-text.
- Speech-to-text minute — per started minute, and only when a video had no usable
captions so speech-to-text had to run. A captioned video never triggers it, and a
caption-less video too long for the speech route (see
audio_too_long_for_speechbelow) is reported uncharged instead of being charged for a partial transcript.
What can fail, and what it costs you: nothing. The two most common non-deliveries are
no_speech (the audio is music or silence, so there is no transcript to sell) and
blocked_retry (YouTube challenged or throttled the fetch, or was still processing the video).
Both come back as a row with charged: false naming the reason — as do no_captions,
private_or_members_only, removed_or_unavailable, age_restricted, region_blocked and
live_no_transcript_yet. You pay for delivered transcripts and nothing else.
Switch "Use speech-to-text when a video has no captions" off and you will never be
charged a speech-to-text minute; caption-less videos come back as uncharged rows instead.
Max speech-to-text minutes and Max videos are hard stops, not suggestions: the run
finishes successfully and each skipped row names the limit that stopped it.
When a video is not charged
status | what happened | temporary? |
|---|---|---|
no_speech | the audio is music or silence — there is no transcript to sell | no |
no_audio_stream | the video exposes no audio track at all | no |
no_captions | no captions, and you switched speech-to-text off for this run | no |
asr_unavailable | our speech-to-text service refused this actor's access mid-run — that is on us, not you. Captioned videos still delivered; caption-less ones came back uncharged | yes — try again later |
audio_too_long_for_speech | no captions, and the video is too long for the speech-to-text route — YouTube does not release enough of its audio for a complete transcript, and we never sell a partial one as whole. Captioned videos are unaffected at any length | no |
private_or_members_only | private or members-only | no |
removed_or_unavailable | YouTube says the video is gone; its own words are in the row | no |
age_restricted | YouTube requires a signed-in, age-verified account | no |
region_blocked | the uploader has not published it in the country we fetched from | no |
live_no_transcript_yet | a live stream — no transcript exists until it ends | re-run after it ends |
blocked_retry | YouTube challenged, throttled or was still processing | yes — re-run |
skipped_too_large · skipped_budget · skipped_speech_cap | a limit stopped it; the row names which | yes |
input_error | the link was not a YouTube video link | fix and re-run |
A temporary problem is never reported as a permanent one. Bare video IDs are the one thing
we cannot sanity-check: a typo in an 11-character ID is indistinguishable from a real ID
that has been deleted, so it comes back as removed_or_unavailable, uncharged.
This actor may fail when the platform changes things — failed items are never charged.
FAQ
How do I get a YouTube transcript without an API key? Paste the video links and run it. There is no YouTube API key, no cookies and no sign-in.
Can I download YouTube subtitles as SRT?
Set format to srt (or vtt) and each row carries a ready-to-save subtitle string.
What if a video has no captions?
Speech-to-text runs on the audio and you get real text, marked source: "speech_ai". If
the audio has no speech at all, the row comes back no_speech and uncharged. The speech
route works on short videos only: past a few minutes YouTube stops releasing the audio to
anything but its own player, so a long caption-less video comes back
audio_too_long_for_speech and uncharged rather than half-transcribed. Videos that have
captions — the large majority, including nearly every spoken upload — are unaffected at
any length.
Can I pick the caption language?
Set language. If that language is not published for the video, the default track is used
and the row's language field tells you what you actually got.
Can I transcribe a whole channel?
Use steadyfetch/youtube-channel-transcripts — paste a channel URL, @handle or channel ID
and it returns every video's transcript in one run. Playlists are not supported by either actor
yet; for a playlist, paste its video links here.
Can I use this through an MCP server? Yes. It is a standard Apify actor, so any MCP client that can call Apify actors can call it.
Why does a run cost more than the transcripts?
Apify bills platform usage (compute and proxy) for what a run actually consumes, separately
from these events. Maximum cost per run is the ceiling that covers both.
Steadyfetch YouTube suite
Same transcript engine, different way in. All-inclusive pay per event, no start fee, charged only on delivery.
| What you paste | Actor |
|---|---|
| Video URLs or IDs | this actor |
A channel URL, @handle or channel ID | YouTube Channel: All Transcripts |
The rest of the steadyfetch shelf — same contract everywhere: all-inclusive pay per event, no start fee, charged only on delivery.
| Family | Actors |
|---|---|
| Ad creative intelligence | Facebook · Google Ads video · TikTok · LinkedIn · Google Ads text & OCR |
| Trends & keywords | Google Trends · Trends Now · Breakout keywords · Autocomplete keywords · Keyword volume & CPC · Social trends |
| YouTube transcripts | YouTube videos · YouTube channels |
| Reel transcripts · Profile posts | |
| Jobs | Indeed · Career sites by domain · Glassdoor · Multi-board |
| Amazon | Products · Search · Bestsellers · Sellers |
| Any media file | Speech to Text · any link or file |
Unlinked names are publishing shortly on the same account — search steadyfetch on Apify Store.
Free n8n templates for the suite: github.com/steadyfetch/n8n-templates — no community nodes needed.
Feedback & support
Found an issue? Open it on the Issues tab — issues are answered in about 3 hours, and always within one business day.