YouTube Transcript API avatar

YouTube Transcript API

Pricing

from $0.40 / 1,000 video transcripts

Go to Apify Store
YouTube Transcript API

YouTube Transcript API

Extract available captions from public YouTube videos using URLs or video IDs. Get full transcript text, timestamped segments, optional word timings and paragraphs, caption details, and SRT, VTT, Markdown, and PDF download links.

Pricing

from $0.40 / 1,000 video transcripts

Rating

0.0

(0)

Developer

Maxime Dupré

Maxime Dupré

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

🎬 Turn public YouTube videos into searchable text

YouTube Transcript API is for developers, researchers, creators, and teams that need text from public YouTube videos. Give it video URLs or IDs and get the full available transcript, timestamped segments, caption language and source, optional word timings and paragraphs, plus download links for SRT, VTT, Markdown, and PDF files. Use the results in your own search, review, or content workflow.

Use it as a YouTube transcript extractor or YouTube to text API for available captions. It is not a speech-to-text YouTube transcript generator, so it does not create new transcripts when a video has no usable captions.

📚 Full YouTube transcripts with timing

Each saved dataset row contains the full available transcript text, source video identity, caption language, caption source, timestamped segments, and links to SRT, VTT, Markdown, and PDF files. Some rows also contain word-level timing and paragraph grouping. The transcriptResults output link opens the dataset.

▶️ Run a transcript job from URLs or IDs

  1. Add one or more public YouTube video URLs or raw video IDs.
  2. Set a preferred caption language if you have one.
  3. Choose creator-provided captions only, or allow auto-generated captions too.
  4. Start the run and open the transcript dataset when it finishes.

The Actor reads the caption source made available for each public video. If the preferred language is not available, another available caption language may be used and the returned language is included in the row. A video without a usable transcript does not stop the rest of the submitted list.

⚙️ Input

Add one or more public video references. The schema has no fixed maximum, so use a list sized for the run and the source limits you expect.

Input fields

FieldTypeWhat it does
videoUrlsOrIdsarray of stringsRequired. Adds public YouTube video URLs or raw video IDs. URLs and IDs can be mixed in the same list.
preferredCaptionLanguagestringOptional language code such as en or en-US. If it is unavailable, another available caption language is used and reported.
captionSourcestring enumChoose creator for creator-provided captions only, or creator_or_auto to allow auto-generated captions too.

This example is copied from the public input of a successful current-beta default-input run:

{
"videoUrlsOrIds": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ"
],
"preferredCaptionLanguage": "en",
"captionSource": "creator_or_auto"
}

🧾 Output

Run output

FieldTypeWhat it does
transcriptResultsstring URLOpens the default dataset that contains the transcript rows.

Transcript row

Each saved row has this complete field shape. Optional fields can be absent when the source does not provide them.

FieldTypeWhat it does
videoIdstringYouTube video ID for the source video.
titlestring, when availableTitle of the YouTube video.
channelNamestring, when availableName of the YouTube channel.
thumbnailUrlURL string, when availableURL for the video thumbnail.
captionLanguagestringLanguage of the captions returned for the video.
captionSourcestring enumcreator for creator-provided captions or autoGenerated for auto-generated captions.
transcriptTextstringComplete available transcript as plain text.
segmentsarray of objectsTimestamped transcript segments in playback order.
segments[].textstringText spoken in the segment.
segments[].startSecondsnumberSegment start time from the beginning of the video, in seconds.
segments[].durationSecondsnumberSegment duration in seconds.
segments[].wordsarray of objects, optionalWord-level timing data when it is available.
segments[].words[].textstringText of the word.
segments[].words[].startSecondsnumberWord start time from the beginning of the video, in seconds.
segments[].words[].durationSecondsnumberWord duration in seconds.
paragraphsarray of strings, optionalParagraph grouping of the transcript text.
downloadsobjectDownload URLs for the generated transcript files.
downloads.srtUrlURL stringSRT subtitle file URL.
downloads.vttUrlURL stringWebVTT subtitle file URL.
downloads.markdownUrlURL stringMarkdown transcript file URL.
downloads.pdfUrlURL stringPDF transcript report URL.

This complete row is from a successful current-beta run. It shows a creator-provided transcript without optional word timings:

{
"videoId": "jNQXAC9IVRw",
"title": "Me at the zoo",
"channelName": "jawed",
"thumbnailUrl": "https://i.ytimg.com/vi/jNQXAC9IVRw/hqdefault.jpg?sqp=-oaymwEmCOADEOgC8quKqQMa8AEB-AG-AoAC8AGKAgwIABABGFUgWShlMA8=&rs=AOn4CLA9eLBatYv9WbkD4BbZ2Im-biSPTw",
"captionLanguage": "en",
"captionSource": "creator",
"transcriptText": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say",
"segments": [
{
"text": "All right, so here we are, in front of the elephants",
"startSeconds": 1.2,
"durationSeconds": 2.16
},
{
"text": "the cool thing about these guys is that they have really...",
"startSeconds": 5.318,
"durationSeconds": 2.656
},
{
"text": "really really long trunks",
"startSeconds": 7.974,
"durationSeconds": 4.642
},
{
"text": "and that's cool",
"startSeconds": 12.616,
"durationSeconds": 1.751
},
{
"text": "(baaaaaaaaaaahhh!!)",
"startSeconds": 14.421,
"durationSeconds": 1.312
},
{
"text": "and that's pretty much all there is to say",
"startSeconds": 16.881,
"durationSeconds": 2
}
],
"paragraphs": [
"All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say"
],
"downloads": {
"srtUrl": "https://api.apify.com/v2/key-value-stores/e7Q8cG3AEmd4BrwmP/records/transcript-jNQXAC9IVRw.srt?signature=15FbEZZTiuuI9oqYAWDhm",
"vttUrl": "https://api.apify.com/v2/key-value-stores/e7Q8cG3AEmd4BrwmP/records/transcript-jNQXAC9IVRw.vtt?signature=LvDxwUJum07BoX4JQtDP",
"markdownUrl": "https://api.apify.com/v2/key-value-stores/e7Q8cG3AEmd4BrwmP/records/transcript-jNQXAC9IVRw.md?signature=2S3dTG4ppsWeZUaCnPX1",
"pdfUrl": "https://api.apify.com/v2/key-value-stores/e7Q8cG3AEmd4BrwmP/records/transcript-jNQXAC9IVRw.pdf?signature=1r6ljRrwL9xAh277cpvdD"
}
}

This shortened row is also from a successful current-beta run and shows optional word timings. The string "..." marks omitted transcript, segment, word, and paragraph data.

{
"videoId": "EJED0LNsdPc",
"title": "How to Use YouTube Transcript Auto Generate Feature !",
"channelName": "Tekwiq",
"thumbnailUrl": "https://i.ytimg.com/vi/EJED0LNsdPc/sddefault.jpg",
"captionLanguage": "en",
"captionSource": "autoGenerated",
"transcriptText": "...",
"segments": [
{
"text": "Hey there guys, welcome back to another",
"startSeconds": 0.08,
"durationSeconds": 4.239,
"words": [
{
"text": "Hey",
"startSeconds": 0.08,
"durationSeconds": 0.32
},
{
"text": "there",
"startSeconds": 0.4,
"durationSeconds": 0.319
},
"..."
]
},
"..."
],
"paragraphs": [
"..."
],
"downloads": {
"srtUrl": "https://api.apify.com/v2/key-value-stores/ZCdEt7x52xcUFvtDW/records/transcript-EJED0LNsdPc.srt?signature=A2A6MvYr5ROiVrWvyTyB",
"vttUrl": "https://api.apify.com/v2/key-value-stores/ZCdEt7x52xcUFvtDW/records/transcript-EJED0LNsdPc.vtt?signature=10mQyHnWw8LpgykKzi2pf",
"markdownUrl": "https://api.apify.com/v2/key-value-stores/ZCdEt7x52xcUFvtDW/records/transcript-EJED0LNsdPc.md?signature=1JEfw4V2wgVPn9PExbx6s",
"pdfUrl": "https://api.apify.com/v2/key-value-stores/ZCdEt7x52xcUFvtDW/records/transcript-EJED0LNsdPc.pdf?signature=100trB3p4uaRiOwhmbztp"
}
}

💳 Pricing

This Actor uses pay-per-event pricing. The primary event is youtube-video-transcript, titled Video transcript, and it covers one complete available transcript saved for one public YouTube video. The current tier price is shown in the Pricing tab and can vary by plan. This README does not promise a fixed rate.

🔌 Integrations

Use the Apify Actor and API access to open the transcriptResults dataset link and read transcript rows in your own workflow. The output keeps the video ID, caption details, timing, transcript text, and file links together.

❓ FAQ

What happens when the preferred caption language is missing?

The Actor can use another available caption language and reports the language it returned in captionLanguage. If you need creator-provided captions only, set captionSource to creator.

Can I choose creator-provided captions only?

Yes. Set captionSource to creator. Use creator_or_auto when auto-generated captions are also acceptable.

Does this create a transcript from video audio?

No. It extracts available creator-provided or auto-generated captions. It does not download video or audio and does not create new speech-to-text transcripts when captions are missing.

Does it search YouTube transcripts by keyword?

No. Submit the public video URLs or IDs you want to read. The Actor does not discover videos from search terms, channels, or playlists. You can use the returned text in your own YouTube transcript search workflow.

What is included in a YouTube transcript download?

Each saved row includes links for SRT, VTT, Markdown, and PDF files. It also includes plain transcript text, timestamped segments, caption language, and caption source.

What happens when the same video is submitted more than once?

The first eligible occurrence is saved. Later copies of that same source video are ignored, so one saved row describes the first match rather than collecting every later submitted value.

Can I use the transcript in an AI or search workflow?

Yes. The Actor returns structured text and timing through the dataset and API. It does not summarize, translate, create chapters, or generate embeddings, so apply those transformations in your own workflow.

What if one video has no usable transcript?

That video does not stop the rest of the submitted list. Results depend on the public captions available for each video and the caption options you choose.

📝 Changelog

v0.0 (19-09-2026)

  • Initial release.

🆘 Support

For issues, questions, or feature requests, file a ticket and I'll fix or implement it in less than 24h 🫡

Made with ❤️ by Maxime Dupré