YouTube Shorts Transcript Scraper - Bulk Shorts to Text avatar

YouTube Shorts Transcript Scraper - Bulk Shorts to Text

Pricing

from $1.70 / 1,000 transcript from native captions

Go to Apify Store
YouTube Shorts Transcript Scraper - Bulk Shorts to Text

YouTube Shorts Transcript Scraper - Bulk Shorts to Text

Transcribe YouTube Shorts to text, timestamped segments, SRT and VTT in bulk. Uses YouTube's own caption track when there is one and falls back to Whisper speech-to-text when there is not. No API key, no login, no cookies.

Pricing

from $1.70 / 1,000 transcript from native captions

Rating

0.0

(0)

Developer

Saulius Saulenas

Saulius Saulenas

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

YouTube Shorts Transcript Scraper — bulk Shorts to text

Turn YouTube Shorts into transcripts, timestamps, SRT and VTT — in bulk.

Paste Shorts links, get the words back. Every Short returns plain text, timestamped segments, a ready-to-use .srt subtitle file, a .vtt file, and the source metadata: channel, title, hashtags, duration, language, view count.

YouTube captions most Shorts itself, and this Actor reads that track first — which is fast and costs $0.002 a video. A Short with no usable caption track is transcribed from its audio with Whisper instead, so a caption-less Short is not simply skipped.

  • No API key, no login, no cookies.
  • 90+ languages. The transcript comes back in the video's own language by default.
  • Bulk-safe. One dead link never breaks the rest of the run.
  • You only pay for delivered transcripts. Failures, private videos and over-length videos are free.

What it handles

InputSupported
youtube.com/shorts/VIDEOID
youtu.be/VIDEOID
youtube.com/watch?v=VIDEOID✅ — the same video, whichever URL you copied
youtube.com/embed/VIDEOID, m.youtube.com/…, youtube-nocookie.com/…
Regular (non-Short) YouTube videos✅ from captions, up to maxDurationSeconds
Channel, playlist or search URLs❌ — not single videos; returned as INVALID_URL
Private, deleted, age-restricted or members-only videos❌ — returned as an error row, uncharged
Live streams❌ — LIVE_VIDEO, uncharged

Other platforms return a clear error pointing you at Short Video Transcriber, which handles TikTok and YouTube Shorts in one run.


Captions vs Whisper — and why it matters to your bill

  1. YouTube's caption track first. Creator-written subtitles when they exist, YouTube's own auto-captions otherwise. Charged $0.002.
  2. Whisper when there is none. The audio is downloaded and transcribed locally with faster-whisper. Charged $0.040, because it costs roughly a second of CPU per 1.5 seconds of audio.

YouTube auto-captions the great majority of spoken-word Shorts, so most runs sit near the $0.002 end. Exactly one of those two events is charged per video, after the row is written — never both, never before. A video that fails is not charged at all.


What you get

One dataset row per input URL:

{
"url": "https://www.youtube.com/shorts/BDf5mOfSkkw",
"platform": "youtube_shorts",
"videoId": "BDf5mOfSkkw",
"author": "UCUvUKGDBCqJzV4OwJkQWCHg",
"authorName": "King Arthur Baking Company",
"title": "The Only Sourdough Recipe You'll Ever Need",
"hashtags": ["baking", "sourdough"],
"durationSeconds": 15.0,
"language": "en",
"source": "native_captions",
"sourceDetail": "youtube_shorts_auto_captions",
"isAutoGenerated": true,
"text": "I'm going to show you how to make the only sourdough bread that you need to know…",
"segments": [
{ "start": 0.0, "end": 1.14, "text": "I'm going to show you how to make the" },
{ "start": 1.14, "end": 2.46, "text": "only sourdough bread that you need to" }
],
"srt": "1\n00:00:00,000 --> 00:00:01,140\nI'm going to show you how to make the\n\n2\n…",
"vtt": "WEBVTT\n\n00:00:00.000 --> 00:00:01.140\nI'm going to show you how to make the\n\n…",
"wordCount": 62,
"segmentCount": 9,
"viewCount": 184231,
"status": "ok",
"error": null
}

source is always exactly one of:

  • native_captions — the words came from YouTube's caption track.
  • whisper_asr — the video had no usable captions, so the audio was transcribed with Whisper.

The srt and vtt fields are complete, valid subtitle files. The scrolling overlaps that YouTube's auto-caption track produces are trimmed, so subtitles do not double up on screen.


Pricing

EventPriceWhen it fires
caption_transcript$0.002A transcript came from YouTube's caption track.
asr_transcript$0.040A transcript came from Whisper because there were no captions.
apify-actor-start$0.00005Apify's standard start event, charged once per gigabyte of the run's memory. This Actor runs at 2 GB, so $0.0001 per run.

Those are the Free-plan prices. Apify's paid plans get the standard Store discount off every event — Bronze 5 %, Silver 10 %, Gold and above 15 % — so asr_transcript costs $0.038, $0.036 or $0.034 on those plans and caption_transcript costs $0.0019, $0.0018 or $0.0017.

What a run actually costs

1,000 ShortsCost
All captioned$2.00
90 % captions / 10 % Whisper$5.80
70 % captions / 30 % Whisper$13.40
None captioned$40.00

(Free-plan prices; a Gold plan pays 15 % less than each of these.)

Turn "Transcribe with Whisper when a Short has no captions" off to cap the price at $0.002 per video: caption-less Shorts then return a CAPTIONS_UNAVAILABLE error, free of charge. On YouTube that is a reasonable default for large batches, because the caption coverage is good.

Nothing else is charged. In particular you are not charged for:

  • videos that are deleted, private, age-restricted or members-only;
  • live streams;
  • videos longer than your maxDurationSeconds cap;
  • caption-less videos longer than maxAsrDurationSeconds (they return ASR_DURATION_EXCEEDED);
  • videos where Whisper finds no speech at all (music-only clips);
  • channel URLs, malformed URLs or links from other platforms;
  • any run that fails.

Input

FieldTypeDefaultWhat it does
videoUrlsarrayRequired. YouTube Shorts URLs, one per line. youtu.be and watch?v= links work.
languagestring(blank)Two-letter code (en, es, pt…). Blank keeps the video's own language. Setting it selects that language's caption track when the channel published one, and tells Whisper what to expect.
allowWhisperFallbackbooleantrueOff = captions only, and caption-less Shorts cost nothing.
maxDurationSecondsinteger600Videos longer than this are skipped, free. Max 1200.
maxAsrDurationSecondsinteger180Length limit for the Whisper path only — exactly the length of the longest Short. Caption-less videos longer than this return ASR_DURATION_EXCEEDED, free. Captioned videos are unaffected.
maxVideosinteger1000Safety cap on the number of URLs processed.
forceWhisperbooleanfalseAdvanced. Transcribe with Whisper even when captions exist. Charged at the asr_transcript rate — 20× the caption price.
whisperModelstringbasetiny, base or small. Only affects the Whisper fallback.
beamSizeinteger1Whisper decoding beam width.
concurrencyinteger5Videos resolved in parallel. Whisper always runs one at a time.
includeFailedItemsbooleantrueKeep a row (with an error code) for every failure, so inputs reconcile to outputs.
useResidentialFallbackbooleantrueLeave this on. YouTube refuses datacenter IPs outright; this retries those requests through a residential IP.
proxyConfigurationobjectApify ProxyRecommended. Used for the first attempt on every request.

Minimal input

{
"videoUrls": [
"https://www.youtube.com/shorts/BDf5mOfSkkw",
"https://youtu.be/pM62yIJNuao"
]
}

Run it

In the Apify Console

Open the Actor, paste your links into YouTube Shorts URLs, one per line, and click Start. Results appear in the dataset and export as JSON, CSV, XLSX or Excel — or download the srt column straight into subtitle files.

API

curl -X POST "https://api.apify.com/v2/acts/sauliusautomatesit~youtube-shorts-transcript-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"videoUrls": [
"https://www.youtube.com/shorts/BDf5mOfSkkw"
],
"allowWhisperFallback": true
}'

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("sauliusautomatesit/youtube-shorts-transcript-scraper").call(input={
"videoUrls": [
"https://www.youtube.com/shorts/BDf5mOfSkkw",
"https://www.youtube.com/shorts/pM62yIJNuao",
],
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
if item["status"] != "ok":
print(f"{item['inputUrl']}: {item['error']}{item['errorMessage']}")
continue
print(item["authorName"], item["language"], item["wordCount"], "words", f"({item['source']})")
with open(f"{item['videoId']}.srt", "w", encoding="utf-8") as handle:
handle.write(item["srt"])

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('sauliusautomatesit/youtube-shorts-transcript-scraper').call({
videoUrls: ['https://www.youtube.com/shorts/BDf5mOfSkkw'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) {
if (item.status !== 'ok') {
console.warn(`${item.inputUrl}: ${item.error}`);
continue;
}
console.log(`${item.authorName} (${item.source}): ${item.text}`);
// await fs.writeFile(`${item.videoId}.srt`, item.srt);
}

n8n, Make and Zapier

Works with the standard Apify → Run Actor node in n8n and the Apify → Run an Actor module in Make. Point the node at this Actor, pass videoUrls as the input JSON, then read the dataset items in the next step — text for the transcript, srt for a subtitle file, error to branch on failures.

A common shape: a sheet or webhook supplies Shorts links → this Actor returns transcripts → an LLM node summarises, tags or repurposes them → the result is written back. Because failures come back as rows with an error field rather than as a broken run, a filter on status == "ok" is all the error handling most workflows need.


Errors

Failed videos get a dataset row with status: "error", an error code and a plain-English errorMessage. They are never charged.

CodeMeaning
VIDEO_UNAVAILABLEDeleted, never existed, or region-locked.
VIDEO_PRIVATEPrivate or members-only.
LOGIN_REQUIREDYouTube demanded a signed-in session. Turn useResidentialFallback on.
AGE_RESTRICTEDAge-restricted or behind a content warning.
LIVE_VIDEOA live stream, which cannot be transcribed.
DURATION_EXCEEDEDLonger than maxDurationSeconds.
ASR_DURATION_EXCEEDEDNo captions, and longer than maxAsrDurationSeconds. Raise that limit to transcribe it anyway.
CAPTIONS_UNAVAILABLENo captions and the Whisper fallback is turned off.
NO_SPEECH_DETECTEDWhisper found no speech — a music-only or silent video.
INVALID_URLNot a link to a single video — a channel or playlist URL, for example.
PLATFORM_NOT_ENABLEDA TikTok or other non-YouTube URL.
MEDIA_UNAVAILABLEThe audio could not be downloaded for transcription.
MEMORY_LIMITThe video does not fit in the run's memory for the chosen model. Raise the memory or pick a smaller model.

Limits and good to know

  • Leave the residential retry on. YouTube answers datacenter addresses — Apify Proxy's and the platform's own alike — with "Sign in to confirm you're not a bot", for videos that are perfectly public. This is not optional cost-shaving: with useResidentialFallback off, nearly every Short returns LOGIN_REQUIRED. Only the requests the cheap route already refused go over residential, and caption downloads stay on the cheap route, so the extra traffic is a few hundredths of a cent per video.
  • Public videos only. No login, no cookies, no members-only or age-restricted content.
  • Auto-captions are auto-captions. When a transcript comes from YouTube's own ASR track (isAutoGenerated: true), it carries that track's mistakes — including missing punctuation on older tracks. For maximum accuracy on a specific video, switch forceWhisper on and accept the asr_transcript price, or raise whisperModel to small.
  • Caption availability changes. YouTube backfills captions hours after upload, so the same video can take the Whisper path today and the caption path tomorrow.
  • Whisper is length-limited; captions are not. A caption-less video is transcribed up to maxAsrDurationSeconds (180 s by default — exactly the maximum length of a Short). Past that, Whisper's cost outgrows a per-video price, so the video returns ASR_DURATION_EXCEEDED and costs you nothing. Videos with captions run to the full maxDurationSeconds at the $0.002 rate, whatever their length.
  • Memory. Run with at least 2 GB. That covers the whole duration range with the base model; the Actor refuses an impossible combination up front instead of being killed mid-run.
  • One URL per video. This Actor does not crawl channels or playlists to find Shorts for you.

What this Actor does not do

It transcribes videos. It does not scrape comments, channel statistics or search results, does not monitor channels, and does not summarise, translate or score sentiment — pipe the transcript into an LLM step for that.


YouTube Shorts transcript · Shorts to text · transcribe YouTube Shorts · YouTube Shorts subtitles · Shorts SRT · YouTube Shorts caption extractor · bulk Shorts transcripts · YouTube subtitle downloader · Shorts speech to text · YouTube Shorts VTT