YouTube Transcript Scraper - Timestamps, No API Key or Login avatar

YouTube Transcript Scraper - Timestamps, No API Key or Login

Pricing

Pay per event

Go to Apify Store
YouTube Transcript Scraper - Timestamps, No API Key or Login

YouTube Transcript Scraper - Timestamps, No API Key or Login

YouTube transcript API without a YouTube API key: transcripts with timestamps from videos, channels or playlists, manual or auto captions in any language, plus title, channel, views and date. Optional speech-to-text. Failed videos free. $3/1,000.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Yukai Lin

Yukai Lin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Categories

Share

What does YouTube Transcript Scraper do?

YouTube Transcript Scraper is a YouTube transcript API that needs no YouTube API key: paste video, channel or playlist links and get the full transcript, timestamped segments and video details, one row per video, for $3 per 1,000 transcripts. Videos without a transcript and failed videos are free.

How to get YouTube transcripts without an API key

  1. Paste video links into Video URLs, or @handles and playlist URLs into Channels and playlists.
  2. Optional: set Preferred languages (for example en, es).
  3. Click Start, or call the Actor from the Apify API with your Apify token: no YouTube API key, cookies or login.
  4. Each video comes back with transcript and timestamped segments; export as JSON, CSV or Excel. 100 videos cost $0.30.

Why this one

  • $3 per 1,000 transcripts, no start fee. The two most used YouTube transcript Actors charge $5 (starvibe) and $10 (pintostudio) per 1,000 (comparison, checked 3 October 2026). Not the cheapest for captions: a few charge $1 or less.
  • Failed videos are free, and each one says why (failureReason), plus which language you got and whether it came from manual captions, auto captions or speech-to-text.
  • Videos without captions: optional speech-to-text (Whisper large-v3-turbo) at $0.008 per audio minute, off by default. The five most used transcript Actors have no priced speech-to-text event.
  • Notes, RAG chunks and a channel digest included: a Markdown page per video, chunks that link to the second in the video (&t=), and only-new-videos runs for channels, at no extra charge.
  • No YouTube API key, no cookies, no login. Public videos only.

Input (copy and paste)

{
"videos": ["https://www.youtube.com/watch?v=jNQXAC9IVRw"]
}

For a channel's latest uploads: {"channels": ["@mitocw"], "maxVideosPerSource": 10}.

Output (real result, shortened)

The video above (human captions), $0.003:

{
"success": true,
"videoId": "jNQXAC9IVRw",
"title": "Me at the zoo",
"channel": "jawed",
"durationSec": 19,
"language": "en",
"source": "manual_captions",
"transcript": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
"segments": [
{ "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" }
],
"wordCount": 39,
"failureReason": null
}

All fields, speech-to-text, RAG chunks and video notes: Output example.

How much does it cost?

EventPriceWhen
Video transcript$3.00 / 1,000 ($0.003 each)a video whose transcript was returned (captions)
Speech-to-text minute$8.00 / 1,000 minutes ($0.008 per minute)only when you turn speech-to-text on, per audio minute (rounded up per video) of a transcript returned by speech-to-text, instead of the video charge

No start fee, no proxy costs, no platform usage on top. Bigger Apify plans get volume discounts (Scale 10%, Business and higher 20%).

Examples: 1,000 videos with captions = $3.00. A 20-minute video without captions, transcribed = 20 × $0.008 = $0.16. A 19-second video transcribed with speech-to-text is billed one minute = $0.008.

Control your cost

  • Charged: only videos returned with success: true.
  • Free: every failure (failureReason set, success: false): no captions (with speech-to-text off), captions only in other languages, no speech (music, ambience), over your speech-to-text limits, private, members-only, age-restricted or removed videos, blocked or busy service, errors.
  • Cap it: set your maximum charge per run. The Actor stops cleanly (SUMMARY.status is LIMIT_REACHED) and reserves the money for captions first, then speech-to-text minutes.
  • Speech-to-text is off by default, so a normal run can never cost more than $0.003 per video.

Features

Captions are the default (manual first, then auto-generated, in the language you ask for). For videos that have no captions at all you can switch on speech-to-text (Whisper large-v3-turbo, $0.008 per audio minute, off by default).

  • Videos, channels and playlists: paste video links, @handles, channel URLs or playlist URLs. Channels give their latest uploads first.
  • Clear language results: each item shows the languages you asked for (requestedLanguages), the language returned (language), whether it matched exactly or only the base language (languageMatch, e.g. zh-Hans returned for zh-TW), and where the text came from (source: manual_captions, auto_captions or asr).
  • Timestamps: segments (start, duration, text) next to the full text. Styled "karaoke" captions are cleaned (no <font> tags, no repeated lines).
  • Ready-made outputs: Markdown video notes, RAG chunks that link back to the moment in the video (&t=), and a "new videos" digest for scheduled channel runs. No extra charge.
  • Failed videos are free, and each one has a failureReason: captions_unavailable, language_mismatch, empty_text, blocked, video_unavailable, age_restricted, service_busy or error.
  • No YouTube API key, no cookies, no login. Only public videos.

Use cases

  • Video notes: turn a talk, lecture or tutorial into a Markdown page for Notion, Obsidian or a doc. Each 3-minute section is headed by a link to that point of the video.
  • Channel digest: schedule a channel daily and get only the new uploads, with links and the opening lines of each transcript. A run with nothing new costs nothing.
  • RAG and search data: transcripts cut into time windows with a url that opens the video at that second, so an answer can cite "see 3:00". Ready for embeddings.
  • Research and content work: quote search, subtitles in several languages, competitor or creator analysis, training data from videos you are allowed to use.

Input

FieldWhat it does
Video URLsWatch, youtu.be (also with ?t=), Shorts, live-replay or mobile links, attribution_link links, or 11-character video ids. One per line or several in one cell; slightly broken pasted links (/watch/?v=, a trailing space or quote) still work. A channel or playlist link pasted here is expanded like one in the next field
Channels and playlists@handle, https://www.youtube.com/@handle, /channel/UC... or playlist URLs
Max videos per channel or playlistLatest N videos from each (default 20)
Max videos in totalLimit for the whole run (default 100, up to 500)
Preferred languagesCaption languages in order of preference (en, es, zh-TW ...)
Include timestamped segmentsAdds segments next to the plain text (default on)
Transcribe videos without captionsSpeech-to-text for videos that have no caption track (default off)
Max minutes per video / per runSpeech-to-text limits (default 30 / 60). Your maximum charge per run also limits them
Add RAG chunks with timestamp linksAdds chunks (about 60 s each by default) with a link to that moment
Add video notes (Markdown)Adds notesMarkdown: details plus the transcript in 3-minute sections with timestamp links
Only new channel videosFor schedules: skips channel/playlist videos processed in earlier runs (free) and saves a DIGEST page

The default input is one 19-second video: it finishes in seconds and costs $0.003.

How to use it

  1. Paste video links (or channels, playlists) into Video URLs / Channels and playlists.
  2. Optionally set Preferred languages, turn on Add video notes or Add RAG chunks.
  3. Click Start, then export the results as JSON, CSV or Excel, or call the Actor from the API.

Output example

Real result (3 October 2026) for https://www.youtube.com/watch?v=jNQXAC9IVRw (human captions). The description and keywords fields are left out here.

{
"success": true,
"status": "ok",
"videoId": "jNQXAC9IVRw",
"url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
"title": "Me at the zoo",
"channel": "jawed",
"channelId": "UC4QobU6STFB0P71PMvOGN5A",
"publishDate": "2005-04-23T20:31:52-07:00",
"durationSec": 19,
"viewCount": 439408167,
"requestedLanguages": [],
"language": "en",
"languageMatch": null,
"source": "manual_captions",
"isGenerated": false,
"availableLanguages": [{ "code": "en", "name": "English", "generated": false }, { "code": "de", "name": "German", "generated": false }],
"transcript": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
"segments": [
{ "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" },
{ "start": 5.318, "duration": 2.656, "text": "the cool thing about these guys is that they have really..." }
],
"wordCount": 39,
"charCount": 217,
"failureReason": null
}

Two more real results from the same run: a 3-minute cooking video with auto-generated English captions (source: "auto_captions", 79 segments, 598 words) and a 3-minute French news video with auto-generated French captions (language: "fr", 84 segments, 469 words). Auto-generated captions keep YouTube's own wording and may lack punctuation and have recognition errors.

Speech-to-text (video without captions)

Real result for an 89-second BBC News video with no caption track, with speech-to-text on. It was billed 2 minutes (89 seconds rounded up per video) = $0.016:

{
"success": true,
"videoId": "CKRxRE7qRzc",
"durationSec": 89,
"language": "en",
"source": "asr",
"transcript": "Presidents Trump and Xi are about to meet here in the US. So what's at stake? We're in New York City. And Beijing to break it down. ...",
"segments": [{ "start": 0, "duration": 3.64, "text": "Presidents Trump and Xi are about to meet here in the US." }],
"wordCount": 256,
"asr": { "status": "ok", "model": "large-v3-turbo", "minutes": 2, "audioSeconds": 89.4, "detectedLanguage": "en", "languageProbability": 0.999, "transcribeSeconds": 19 },
"chargedMinutes": 2,
"failureReason": null
}

The same video requested with languages: ["zh-TW"] came back as status: "language_mismatch", success: false, transcript: null: the speech was English, so nothing was returned and nothing was charged.

RAG chunks and video notes

With RAG chunks on, each item also has chunks (here with a 60-second target; the 3-minute cooking video gave 4):

"chunks": [{ "chunkIndex": 0, "start": 0, "end": 58.3, "url": "https://www.youtube.com/watch?v=oxbFSZADg-Y&t=0s", "text": "[Music] hello this is Chef John from foodwishes.com with Sherry brazed beef short ribs ..." }]

With video notes on, notesMarkdown looks like this (shortened):

# Me at the zoo
Video: https://www.youtube.com/watch?v=jNQXAC9IVRw
Channel: jawed
Published: 2005-04-23
Length: 0:19
Language: en (captions)
## [0:01](https://www.youtube.com/watch?v=jNQXAC9IVRw&t=1s)
All right, so here we are, in front of the elephants the cool thing about these guys ...

Three ready-to-use recipes

  1. Video notes: {"videos": ["<url>"], "notes": true} gives notesMarkdown, a Markdown page with the video details and the transcript in 3-minute sections; each heading links to that point of the video.
  2. New-videos digest for a channel: {"channels": ["@mitocw"], "maxVideosPerSource": 10, "onlyNewVideos": true, "monitorName": "mit"} on a daily schedule. The first run takes the latest videos; later runs return only new uploads, and the run's DIGEST record (key-value store) is a Markdown list with links and the opening lines of each transcript. A run with nothing new costs nothing.
  3. RAG data with citations: {"videos": [...], "ragChunks": true, "chunkSeconds": 60, "includeTimestamps": false} gives chunks with start, end and a url that opens the video at that moment.

How it compares

Prices of the Apify Store's "YouTube transcript" Actors with the most users, read from the public Store API on 3 October 2026 (FREE-plan price per transcript; per-1,000 in brackets). Most also charge a small start event; we have none. Prices change, so check the listing before you decide.

ActorUsers (30 days)Price per transcriptSpeech-to-text for videos without captions
TidyTools YouTube Transcript Scraper (this Actor)new$0.003 ($3)yes, $0.008 / minute
starvibe/youtube-video-transcript2,763$0.005 ($5)no priced event
pintostudio/youtube-transcript-scraper2,324$0.01 ($10)no priced event
johnvc/YoutubeTranscripts1,994listed at about $0.00001no priced event
supreme_coder/youtube-transcript-scraper862$0.001 ($1), $0.0007 on higher plansno priced event
karamelo/youtube-transcripts817$0.007 ($7), $0.005 on higher plansno priced event
codepoetry/youtube-transcript-ai-scraper444$0.001 ($1) + $0.0025 startyes, $0.012 / minute
apple_yang/youtube-transcripts-scraper8$0.001 ($1)yes, $0.0045 / minute

What this means, honestly: we are not the cheapest for captions. Several Actors charge $0.001 or less per transcript. We compare on what you get for that price: transcripts that tell you which language and source you got, never charge a failed video, include notes, RAG chunks with timestamp links and a channel digest at no extra charge, and offer speech-to-text for videos without captions as an option. The table compares prices and priced events only; we did not run the other Actors, so we make no claim about their reliability or output quality.

Measured results (our own test runs, October 2026)

  • Captions: 300 public videos in 10+ languages; every video got a definite answer (222 transcripts, 73 without captions, 5 unavailable), no errors or blocks. Language requests: Spanish and Japanese returned exactly on 4/4 multilingual videos; zh-TW returned zh / zh-Hans tracks (marked languageMatch: "base") because those videos have no Traditional track; 3/3 videos with captions only in other languages were reported as language_mismatch.
  • Speech-to-text on 20 videos without captions (11 s to 25 min): 12 transcripts (6 English, 5 Chinese written in Traditional characters, 1 Korean), 7 with no speech (music, intros, ambience: not charged), 1 language mismatch (English speech when Chinese was requested: not charged), 0 errors. 73 audio minutes billed in total.
  • Accuracy against human captions (same videos transcribed both ways): English word error rate 2.3%, 3.5% and 4.5% on three 2-5 minute talks (13% on a 19-second clip). Chinese is harder to score: against edited human subtitles of a 17-minute talk the character difference was 41%, partly because subtitles are not word-for-word; names are the most common real errors.
  • Speed: about 5x faster than real time when the queue is empty (a 15-minute video in 2.5 minutes). The 89-second sample above took 19 seconds of transcription.

Limitations

  • Captions are not speech recognition. Manual captions are the creator's; auto-generated captions are YouTube's and can be wrong. Speech-to-text is only used for videos with no caption track at all: a video that has captions, but not in your languages, is language_mismatch and is not sent to speech-to-text.
  • Many videos have no captions: in our 300-video sample, 25% of playable videos had none. News channels in Chinese often burn the subtitles into the picture, so there is no caption track: turn on speech-to-text for those.
  • Public videos only. The Actor never signs in: members-only, private and age-restricted videos are skipped (free). Live streams that are still running are not transcribed; a finished live stream (replay) is treated like any video.
  • No machine translation. Asking for a language the video has no caption track in gives language_mismatch (free) with availableLanguages; pick one of those, or leave the language empty.
  • Links that are not videos are skipped, not charged. A link that is neither a video, channel nor playlist (another site, a YouTube home or search page) is listed in SUMMARY.notes; if nothing usable is left the run stops at once with a message saying what to paste.
  • Capacity is limited, on purpose. Captions are fetched at about 20 videos per minute and up to 3,000 videos per day across all users. Speech-to-text runs one video at a time at about 5x real time, with a daily budget across all users (600 audio minutes for the speech-to-text of this Actor). When capacity is used up or the service is busy, videos come back as service_busy or blocked (free): run them again later. Very large runs (hundreds of videos) take time: 100 videos is about 5-6 minutes.
  • Speech-to-text limits: at most 240 minutes per video, 30 by default; your per-run limit and maximum charge apply.
  • YouTube can change. If YouTube blocks our access, results come back as blocked (free) until we fix it.
  • No AI summary is included in notes: they are the transcript, organized with links.

Responsible use

Transcripts belong to the video creators. Use them within YouTube's terms and copyright law: for personal notes, research, accessibility, search and analysis, or content you have the right to use. Do not republish other people's transcripts as your own.

Use it from code

curl -X POST "https://api.apify.com/v2/acts/tidytools~youtube-transcript/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"videos":["https://www.youtube.com/watch?v=jNQXAC9IVRw"],"notes":true}'

With AI agents (MCP): connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/youtube-transcript) and ask, for example, "Get the transcript of this YouTube video with timestamps".

FAQ

Does the Actor collect usage data? The Actor sends one anonymous usage counter per run (input type, item count range, outcome) to improve the tool; no URLs, content or account data.

Which caption is returned when there are several? With languages set: the first language that exists, manual before auto-generated; languageMatch says whether it was an exact or only a base-language match. Without languages: the video's original language (manual if available).

Why did a video with captions come back as language_mismatch? It has captions, but not in your languages. availableLanguages lists what exists. Speech-to-text is only used for videos with no caption track at all.

Is speech-to-text accurate? It is Whisper large-v3-turbo (the same open model many transcription tools use). On our English samples the word error rate against human captions was 2-5%; names and rare terms are the usual mistakes. Chinese requested as zh-TW is written in Traditional characters.

Why was a short video charged a full minute? Speech-to-text is billed per started audio minute of the video, and only when text was returned.

Which links work? watch?v=, youtu.be/ID (with or without ?t=), /shorts/, /live/, /embed/, m. and music. links, attribution_link links and bare 11-character ids; channels (@handle, /channel/, /c/, /user/) and playlists (/playlist?list=) are expanded to their latest videos. A watch link that also has &list= is that one video.

Do I need a YouTube API key or an account? No.

What are the limits? Up to 500 videos per run (Max videos in total, default 100) and up to 500 per channel or playlist (default 20). Speech-to-text: up to 240 minutes per video (default 30) and 3,000 per run (default 60). Capacity is shared by all users (about 3,000 caption videos a day); videos over it come back as free service_busy rows to run again later. See Limitations.

How fast is it? In our test runs captions came at about 20 videos per minute (100 videos in about 5-6 minutes), and speech-to-text ran at about 5x real time (a 15-minute video in 2.5 minutes). See Measured results.

Is it legal to scrape YouTube transcripts? The Actor reads only the captions of public videos, without signing in, cookies or an API key; it skips private, members-only and age-restricted videos. Transcripts belong to the creators, so use them within YouTube's terms and copyright law (Responsible use). This is not legal advice.

Does it work with the API, webhooks, Make, n8n or Zapier? Yes. One HTTP request returns the transcripts (example); runs can also be started from a schedule, a webhook, the Apify apps for Make, n8n and Zapier, and from AI agents through MCP.

Support

Open an issue in the Issues tab with the video URL and the failureReason. Issues are checked regularly.