YouTube Transcript Scraper — SRT/VTT + Whisper AI Fallback avatar

YouTube Transcript Scraper — SRT/VTT + Whisper AI Fallback

Pricing

from $5.00 / 1,000 transcripts

Go to Apify Store
YouTube Transcript Scraper — SRT/VTT + Whisper AI Fallback

YouTube Transcript Scraper — SRT/VTT + Whisper AI Fallback

Extract YouTube transcripts as plain text, timestamped JSON, SRT or VTT from videos, Shorts, live replays and whole channels. Opt-in Whisper AI transcribes caption-less videos. Lists every caption track, any language, and bills only new uploads on a schedule. Pay only for delivered transcripts.

Pricing

from $5.00 / 1,000 transcripts

Rating

5.0

(3)

Developer

Muhamed Didovic

Muhamed Didovic

Maintained by Community

Actor stats

0

Bookmarked

34

Total users

6

Monthly active users

15 hours ago

Last modified

Categories

Share

YouTube Transcript Scraper — SRT, VTT, JSON & Plain Text with Whisper AI Fallback

Turn any YouTube video or Short into a clean transcript in one run. Every row carries the full text, timestamped segments, ready-made SRT and VTT subtitle files, and the video's core metadata (title, channel, views, duration, description). When a video has no captions at all, the optional Whisper AI fallback downloads the audio and transcribes it with speech-to-text — something caption-only scrapers simply return empty for.

How it works

How the YouTube Transcript Scraper works

✨ Why use this scraper?

  • Fallback ladder — YouTube's caption track (desktop, then mobile), then two transcript libraries, then yt-dlp subtitle download, then Whisper AI speech-to-text. One source being blocked or missing doesn't kill your run.
  • Whisper AI for caption-less videos — the differentiator: videos with no captions still come back transcribed (opt-in, billed as a separate premium event so you never pay it unknowingly).
  • Every format in one row — plain text for LLM pipelines, timestamped JSON segments for analysis, SRT and VTT for subtitle workflows. No post-processing.
  • Language selection with auto-translate — request en, de, es… and where the exact track is missing, YouTube's translated track is used when available.
  • Metadata included free — title, channel and @handle, publish date, view count, duration, keywords, description and thumbnails ride along on every transcript row.
  • Every caption track listed — availableCaptionTracks names each language YouTube has for the video and whether a person wrote it or it is auto-generated, so you can pick the track yourself.
  • Live stream replays work — paste the youtube.com/live/… link as-is. A stream that is still live, or ended so recently that YouTube has not published its captions, is reported as live_not_ready and not charged.
  • Bulk-friendly — paste hundreds of URLs; failed videos are isolated and never billed.
  • The hook on every row — hook3s is the first 3 seconds of speech and hookStartSeconds says when the talking starts, so openers compare side by side without reading transcripts.
  • Chain any YouTube scraper — pass a finished run's datasetId (or paste rows in datasetItems) and every video or channel link in those rows is transcribed; the row shape does not matter.
  • Scheduled runs bill only what is new — name a watchlistId and switch on newItemsOnly; a video transcribed on an earlier run is left out, and a channel walk transcribes only its new uploads.
  • Every miss on the record — asked = delivered + reported. A video that could not be transcribed is in the run's ERRORS record with a reason and what to change, not billed as a row, and the status line reconciles the count.

🎯 Use cases

WhoWhat they do with it
AI & LLM buildersFeed clean plain-text transcripts into RAG pipelines, summarizers, and agents (MCP-friendly output).
Content & SEO teamsRepurpose videos into articles, show notes, and quote pulls; mine competitor channels for topics.
Researchers & analystsBuild searchable corpora from talks, interviews and news coverage, with timestamps intact.
Subtitle & localization teamsGet SRT/VTT straight from the source, plus translated tracks where YouTube offers them.
Media monitoringTrack what's being said about brands and people across YouTube at scale.

📥 Supported inputs

InputExample
Standard video URLshttps://www.youtube.com/watch?v=dQw4w9WgXcQ
Short linkshttps://youtu.be/dQw4w9WgXcQ
Shortshttps://www.youtube.com/shorts/{id}
Live stream replayshttps://www.youtube.com/live/{id}
Channelhttps://www.youtube.com/@handle/videos — walked newest first up to maxItems
Chained datasetdatasetId: "nGkT…" — the default dataset of a finished YouTube scraper run; rows are deep-scanned for video and channel links
Pasted rowsdatasetItems: [ { "url": "https://youtu.be/…" } ]

Not supported: private, members-only or age-gated videos. A stream still in progress has no captions yet; it is reported, not charged. For a playlist, list its videos with the YouTube Channel Videos scraper and pass that run's datasetId.

🔄 How a run works

  1. Each URL is resolved to its video ID and fetched with browser-grade TLS.
  2. The caption track in your requested language is located (auto-translate applied when needed).
  3. If captions are missing or blocked, the ladder steps down: mobile caption track → transcript libraries → yt-dlp subtitles → (opt-in) Whisper AI speech-to-text.
  4. Segments are normalised to timestamped JSON and rendered to SRT and VTT.
  5. One row per video is pushed — transcript, formats, and metadata together.

⚙️ Input parameters

FieldTypeDefaultNotes
startUrlsarray—Video/Shorts/channel URLs, any standard form
datasetIdstring—Default dataset of a finished YouTube scraper run; every video or channel link in its rows is transcribed
datasetItemsarray—Rows pasted from a scraper instead of chaining by id
watchlistIdstring—Named memory for scheduled runs; videos transcribed under this id are remembered
newItemsOnlybooleanfalseLeave out videos already on the watchlist. Needs watchlistId
languagestringdefaultCaption language code (en, de, …); default = the video's original track
whisperFallbackbooleanfalseWhisper AI speech-to-text for caption-less videos — billed per transcribed video as a premium event
maxItemsinteger100Hard cap on billed transcript rows
maxConcurrencyinteger10Parallel video fetches
proxyobjectAutomaticPaid-plan runs use the actor's built-in premium residential pool automatically

📊 Output overview

One row per transcribed video. The transcript appears three ways — transcript (timestamped segments), transcript_only_text (plain text), and transcript_srt / transcript_vtt (ready-to-save subtitle files) — alongside the video's metadata and the spoken hook. A video that could not be transcribed produces no row and no charge; it is listed in the run's ERRORS record with the reason (see below).

🧾 What was not delivered, and why

Every video the run set out to transcribe ends as a dataset row with a transcript or as an entry in the run's ERRORS record (key-value store), never neither, and never as a billed row without a transcript. The record is written on every run, with an empty items list when everything was delivered, so an integration can read it unconditionally. Each entry carries a reason, a plain-language detail with what to change, and a retryable flag:

reasonWhat happenedRe-run as-is?
no_captionsNo caption track and whisperFallback is offSwitch on whisperFallback
no_speechNo captions, and Whisper heard only music, silence or noiseNo
too_longNo captions, and the video is longer than Whisper's 30-minute capNo
live_not_readyA live stream still in progress, not started yet, or ended in the last day before YouTube published captionsLater: captions usually appear a few hours after the stream ends
unavailablePrivate, removed, members-only or age-gatedNo
login_requiredYouTube asked the fetch to sign inYes
rate_limitedYouTube throttled the fetchYes, in a few minutes
fetch_failedAnother transport failureYes

The status line carries the same arithmetic, for example "Delivered 48 of 50 videos. 2 not delivered and not charged: 2 no captions and no Whisper fallback — see the ERRORS record." Videos the watchlist left out are counted separately and never charged. A run that delivers nothing because every fetch failed on YouTube's side is marked FAILED so it stands out in your run list.

🔁 Coming from another YouTube transcript scraper?

What users of other transcript actors report on their issue boards, and what this one does about it:

Reported elsewhereHere
Live video links fail with "impossible to retrieve video ID"youtube.com/live/… links are read as-is; a stream still in progress is reported as live_not_ready, not billed
"No captions found" and an empty resultOpt-in Whisper transcribes caption-less videos up to 30 minutes; a video left untranscribed is named in ERRORS with a reason and never billed
Charged for empty resultsA row is written, and billed, only when it carries a transcript
Want only a channel's new videos each daywatchlistId + newItemsOnly: a scheduled run transcribes and bills only uploads it has not seen
Want the publish date in the outputpublishDate and uploadDate on every row
Want the whole transcript as one string, no timestampstranscript_only_text
Want to see which caption tracks exist before choosingavailableCaptionTracks lists every language and whether it is auto-generated
Requested language ignored, always Englishlanguage returns the transcript in the language you ask for (checked with de and fr on an English video)

📦 Output sample

Real trimmed row from a live run:

{
"videoId": "dQw4w9WgXcQ",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"author": "Rick Astley",
"channelId": "UCuAXFkgsw1L7xaCfnd5JJOw",
"lengthSeconds": "213",
"viewCount": "1699540216",
"transcript": [
{ "text": "[♪♪♪]", "startMs": "1360", "endMs": "3040", "startTimeText": "0:01" },
{ "text": "♪ We're no strangers to love ♪", "startMs": "18800", "endMs": "22140", "startTimeText": "0:18" }
],
"transcript_only_text": "[♪♪♪] ♪ We're no strangers to love ♪ ♪ You know the rules and so do I ♪ …",
"transcript_srt": "1\n00:00:01,360 --> 00:00:03,040\n[♪♪♪]\n…",
"transcript_vtt": "WEBVTT\n\n00:00:01.360 --> 00:00:03.040\n[♪♪♪]\n…",
"keywords": ["rick astley", "Never Gonna Give You Up", "nggyu"],
"thumbnail": { "thumbnails": [{ "url": "https://i.ytimg.com/vi/dQw4w9WgXcQ/…", "width": 168, "height": 94 }] }
}

🗂 Key output fields

FieldMeaning
transcript[]Timestamped segments: text, startMs, endMs, startTimeText
transcript_only_textThe whole transcript as one plain string — LLM-ready
transcript_srt / transcript_vttComplete subtitle files as strings — save and use directly
transcriptSourceWhich ladder step produced it (timedtext / yt-dlp / whisper)
availableCaptionTracks[]Every caption track on the video: languageCode, name, kind (manual or auto-generated), isTranslatable
publishDate / uploadDateWhen the video was published
channelHandleThe channel's @handle, from its profile link
liveBroadcastDetailsStart and end time for a video that was a live stream
hook3s / hookStartSecondsThe first 3 seconds of speech, and when speech starts; a music-only opening reads as the caption marker (for example [♪♪♪])
videoId, title, author, channelIdVideo identity
viewCount, lengthSeconds, keywords, shortDescription, thumbnailMetadata that rides along free

❓ FAQ

What happens with videos that have no captions? Without whisperFallback they produce no row and no charge, and the ERRORS record lists them as no_captions. With whisperFallback: true, the audio is downloaded and transcribed by Whisper speech-to-text — billed as a separate premium event per video, only when it actually produces a transcript. Over music or silence Whisper tends to invent filler ("Thank you.", "♪"); a result made only of that counts as no transcript, so the video is listed as no_speech and neither the row nor Whisper is charged. Whisper handles videos up to 30 minutes long.

Which languages are supported? Any language YouTube has a caption track for. Set language to a code like de or es; when that exact track is missing, YouTube's auto-translated track is used where available.

Can I transcribe a whole channel or playlist? A channel URL (youtube.com/@handle/videos) is walked newest first up to maxItems. For playlists or search results, run the matching scraper and pass its dataset id in datasetId; every video link in the rows is transcribed. On a schedule, add a watchlistId with newItemsOnly so only new uploads are transcribed and billed.

Do I need to configure proxies? No. Runs on a paid Apify plan go through the actor's built-in premium residential pool automatically — YouTube throttles datacenter IPs aggressively, and this keeps success rates high at volume with zero setup. Free-plan runs use Apify's automatic proxy, and the proxy input lets them supply their own.

Do failed videos cost me anything? No. A video without a transcript produces no row, so nothing is billed for it; it is named in the ERRORS record instead. Whisper is only charged when it delivers text.

💬 Support

Found a bug or missing a field? Open an issue on the actor's Issues tab in Apify Console — issues are answered within 1–2 business days.

🛠 Additional services

Need scheduled transcript archives, a merged multi-platform transcript feed (YouTube + TikTok + Instagram + Loom), or delivery straight to your database? Custom builds and SLAs available — contact me through the actor page.

🔎 Explore more scrapers

Same developer, same stack: Video & Audio Transcriber (Whisper), Instagram Transcript Scraper, YouTube Comments, YouTube Search — and the full portfolio at memo23 on Apify Store.

🤖 For AI Agents & LLM Apps

Built for machine consumption: transcript_only_text drops straight into a context window; timestamped segments support citation and chaptering; stable field names across every row. Pair with the Video Transcripts MCP Server to expose transcripts as a tool in agent frameworks. Keep maxItems low per call for cost control; every row is self-contained.


⚠️ Disclaimer

This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by YouTube or Google LLC. All trademarks mentioned are the property of their respective owners.

The scraper accesses only publicly available video pages and caption data — no login, no age-gated, members-only or private content. Users are responsible for ensuring their use complies with YouTube's Terms of Service, copyright law as it applies to transcript content, applicable data-protection law (GDPR, CCPA, etc.), and any contractual obligations of their own organisation.


SEO Keywords

youtube transcript scraper, youtube transcript api, extract youtube transcript, youtube captions scraper, youtube subtitles downloader, srt from youtube, vtt from youtube, youtube video to text, youtube transcription tool, whisper youtube transcription, transcribe youtube videos without captions, youtube transcript for llm, youtube rag pipeline, bulk youtube transcripts, youtube caption extractor, video to text api, youtube shorts transcript, youtube transcript json, apify youtube transcript, pintostudio alternative, youtube-transcript-scraper alternative