YouTube Transcript Scraper avatar

YouTube Transcript Scraper

Pricing

$4.00 / 1,000 results

Go to Apify Store
YouTube Transcript Scraper

YouTube Transcript Scraper

Extract timestamped captions and video metadata from public YouTube videos or channels. Export transcript segments, language options, publication details, and engagement counts.

Pricing

$4.00 / 1,000 results

Rating

0.0

(0)

Developer

MLG Data

MLG Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Scrape YouTube transcripts from public videos, Shorts, and channel video lists. Export timestamped captions and video details to JSON, CSV, or Excel for analysis, accessibility review, search indexing, and content workflows. This YouTube transcript extractor accepts one video, a list of videos, or a public channel; no account is required for the input.

Each successful dataset item represents one video with an available public caption track. The transcript array preserves the timing of every caption segment. If you also need a continuous paragraph, enable include_transcript_text to add transcriptText to the same item. The output includes the language chosen, other listed caption languages, video duration, public view count, and the video and channel identifiers needed to join results with other data.

What data can you extract from YouTube?

FieldDescriptionExample
videoIdStable video identifiergN07gbipMoY
urlCanonical watch URLhttps://www.youtube.com/watch?v=gN07gbipMoY
titlePublic video titleA Practical Guide to Taking Control of Your Life
descriptionPublic description textA short introduction to the talk
channelNameDisplayed channel namePublic channel name
channelIdStable channel identifierAn identifier beginning with UC
channelUsernameHandle when present on the page@example
channelThumbnailPublic channel image URL when presentImage URL
subscriberCountDisplayed subscriber count, rounded when abbreviated27800000
publishedAtUTC publication time when available2025-08-28T15:01:16Z
timestampUnix publication time when available1756393276
viewCountPublic view count571888
likeCountDisplayed like count, rounded when abbreviated13000
commentCountPublic comment count when exposednull
durationSecondsVideo length in seconds520
thumbnailVideo image URLImage URL
transcriptTimed segment objects with text, start, and end[ { "text": "Five years ago,", "start": 4.035, "end": 6.003 } ]
transcriptTextJoined text when requested; otherwise nullFive years ago, I was...
transcriptSegmentCountNumber of caption segments160
languageSelected caption language codeen
selectedLanguageSelected caption label or codeen
availableLanguagesLanguage codes listed for the video["en", "es", "fr"]
isAutoGeneratedWhether the selected track is automaticfalse
geoRestrictGeographic restriction when exposednull
statusSuccessful result statesuccess
messageShort result summaryTranscript retrieved

A timing value is measured in seconds from the beginning of the video. A transcript segment may contain one line, several words, or a full sentence; the uploader and caption format determine those boundaries. The joined text is convenient for reading, while the segment array is better for playback alignment, clipping, quoting, or finding the exact part of a video that contains a phrase.

Some metadata is optional. A missing value is represented by null; it is not an estimate. The public player may expose a view count while the page does not expose a precise comment count. Subscriber and like counts can be abbreviated on the page, so those integers reflect the displayed rounded number rather than a precise count.

How to scrape YouTube transcripts

  1. Paste a public watch URL into youtube_url, add several links to video_urls, or supply a public channel URL in channel_url. Use one mode per run.
  2. Choose a preferred caption language. A two-letter code such as en, es, or fr works for common cases; regional codes may also be used.
  3. For channel runs, set max_videos and optional publication dates. The maximum is a count of videos examined, including videos without usable public captions.
  4. Enable include_transcript_text if you need a single text field as well as timed segments, then start the run.
  5. Open the dataset to download JSON, CSV, or Excel. One dataset row is one video with a retrieved transcript.

A video URL can use the standard watch format, a short link, an embedded-video link, a live replay path, or a Shorts path. A bare eleven-character video identifier also works in video_urls. Public channel handles and channel ID URLs are accepted for channel collection. Private, deleted, age-restricted, or otherwise inaccessible videos may not yield a record.

Input

ParameterTypeDefaultDescription
youtube_urlstringEmptyOne public video URL.
video_urlsstring array[]Several public video URLs or IDs.
channel_urlstringEmptyOne public channel whose video list should be scanned.
languagestringenPreferred caption language code. Another listed language is used if necessary.
max_videosinteger10Maximum channel videos to examine, from 1 to 200.
start_datestringEmptyInclusive publication date for channel videos, YYYY-MM-DD.
end_datestringEmptyInclusive final publication date for channel videos, YYYY-MM-DD.
include_transcript_textbooleanfalseAdd the joined plain-text transcript.
maxItemsinteger0Maximum successful transcript items across the run; zero means no total cap.
proxyConfigurationobjectEnabledConnection configuration for public player and caption requests.

For a single video, use:

{
"youtube_url": "https://www.youtube.com/watch?v=gN07gbipMoY",
"language": "en",
"include_transcript_text": true
}

For several videos, put their URLs in video_urls and leave youtube_url and channel_url empty. Repeated video IDs are removed within the run, so a duplicate input does not create a duplicate result. For a channel, provide channel_url, set max_videos, and optionally set one or both dates. A date filter is applied to the video's reported publication date. An end date earlier than the start date is invalid. Channel order comes from the public video list, generally newest first, and a continuation is requested when the first batch is insufficient.

maxItems limits successful rows, while max_videos limits channel candidates. For example, a channel run that examines 50 videos could return 43 items if seven videos have no usable captions. Set max_videos above the number of results you need to allow for those gaps. A single-video run naturally returns at most one item, because one item contains the whole transcript of that video.

Output example

The following values are from a successful single-video run. The caption array and joined text are shortened here for readability; the dataset item contained 160 timed segments and the full joined transcript. Fields with names or long prose are omitted from this excerpt, and every output field is listed in the table above.

{
"videoId": "gN07gbipMoY",
"url": "https://www.youtube.com/watch?v=gN07gbipMoY",
"viewCount": 571888,
"durationSeconds": 520,
"transcript": [
{"text": "Five years ago,", "start": 4.035, "end": 6.003},
{"text": "I was a prisoner in my own life.", "start": 6.037, "end": 8.172},
{"text": "I was hopelessly addicted to drugs.", "start": 9.44, "end": 11.876}
],
"transcriptText": "Five years ago, I was a prisoner in my own life. I was hopelessly addicted to drugs...",
"transcriptSegmentCount": 160,
"language": "en",
"isAutoGenerated": false,
"status": "success"
}

The actual transcript array contains all segments for the selected track. The transcriptText field is null by default and is populated only when requested. For each segment, start is the first timestamp and end is the final timestamp in seconds. This makes the dataset suitable for subtitle inspection without requiring a separate subtitle file.

Use cases

  • Search within a video library: Collect transcripts from a set of public videos, index the text, and link each search result back to its watch URL and caption time.
  • Accessibility checks: See whether public captions exist, which languages are listed, and whether the selected captions were generated automatically.
  • Research and analysis: Gather spoken content from a defined public channel or a curated list of video URLs, then compare language, duration, and publication dates.
  • Editorial review: Locate a passage through its text and timestamp, inspect the surrounding segments, and prepare a time-specific reference for a review workflow.
  • Monitoring: Schedule a channel run with a date range to collect transcripts from recent uploads and compare them with earlier datasets.
  • Archiving public captions: Preserve the text and timing of available caption tracks in a consistent JSON shape for internal records.

The actor does not create speech-to-text content for videos that have no accessible caption track. Its purpose is to retrieve public captions and attach useful video details to them. That distinction matters when planning coverage for music videos, newly uploaded videos whose captions are still processing, or videos where captions have been disabled.

How much does it cost to scrape YouTube?

The current event price is $4.00 per 1,000 successful transcript items, or $0.004 for one successful video. An item is charged when it is saved to the dataset. Videos without a retrieved transcript do not produce an item. The platform's usage charge is part of this result price.

Successful video transcriptsResult charge
10$0.04
100$0.40
1,000$4.00
5,000$20.00

The table counts returned videos, not caption segments. A 160-segment transcript is one result, just as a 10-segment transcript is one result. Scanning a channel may make requests for videos without captions, but those videos do not add charged dataset items. The maxItems input provides a direct ceiling on successful results. Run time also depends on how many channel candidates must be examined and how many caption tracks must be tried before one yields content.

Tips for best results

Use the exact video URL when you need one known transcript. A single-video run is quick to inspect and provides a clear answer about that video's public caption availability. Use video_urls for a curated list rather than starting many separate runs. Duplicates in that list are ignored within the run.

For a channel, choose a larger max_videos than your target result count. Some uploads are music, live events, very new videos, or videos without accessible captions. max_videos is an examination limit, while maxItems caps the number of successful rows. If you need 30 transcripts and expect some caption gaps, a limit of 40 or 50 candidate videos gives the run room to reach the target.

Choose a language code that matches the material you want to read. The run prefers an exact language code, then a regional variant, then another available track. When both manually created and automatic captions are available in the same language, the manually created track is preferred. Always inspect language, selectedLanguage, and isAutoGenerated before treating the text as a verified human transcription.

Set a publication range for channel monitoring. Dates are inclusive and are evaluated against the public publication time when that time is exposed. If a video has no usable publication date and a date filter is active, it is skipped rather than guessed into the range. Use separate runs for distinct time windows if you need reproducible historical batches.

For a smaller dataset, leave include_transcript_text off. Timed segments already contain all retrieved text. Turn it on when a downstream system expects one plain string per video, accepting that it repeats the text already present in transcript.

Limits

Only public, accessible caption tracks can be retrieved. A video's page may load while its player is restricted, captions are absent, or a listed caption URL returns no segment content. In those cases the actor logs the video and does not add a successful transcript item. There is no promise that every public video has captions.

Video and channel pages can change. The actor reads a public channel listing and a mobile player response, then follows the caption URL supplied by that response. Connection restrictions can vary by video and region. The actor uses a residential connection for those requests because a datacenter connection returned a sign-in interstitial during validation. This makes runs more reliable for accessible public videos, but it does not override private access controls or geographic rights.

A channel's public listing is paginated. The run requests additional pages when needed, up to its 200-video input cap. The video count is a scan limit, not a guarantee of that many transcripts. Dates and metadata depend on what the public player and page expose. In particular, the comment count and geographic restriction can be null; a missing value should not be interpreted as zero or unrestricted access.

The selected caption text is returned as the site provides it. Automatic captions can contain errors, repeated words, or unusual segment boundaries. Translated caption tracks, when listed, may differ from the spoken language. The actor does not correct, summarize, translate, or certify the accuracy of text.

Automation examples

Use a recurring channel run to collect captions for uploads published within a selected time period. Store the dataset and compare videoId values across runs to identify new transcripts. Another workflow can pass a list of video URLs, enable include_transcript_text, and send each successful item's text and watch URL to a search index. For timestamped excerpts, retain the transcript array and build links using the segment's start value.

For repeatable ingestion, keep the same language preference and record the output language for each item. Caption availability can change after an upload, so a later run may produce a transcript for a video that did not produce an item earlier. The stable videoId supports joining those later results to an existing catalog.

FAQ

The actor works with publicly accessible material. You are responsible for following the site's terms, copyright rules, and applicable privacy law. Do not use transcript data to collect or misuse personal information. Review the rights for a particular video before republishing its text.

Do I need to configure a proxy?

The default connection is set up for the public player and caption route. You can inspect the proxy setting in the input, but most runs do not need a custom value. A restricted video will still be restricted even if its watch page opens.

How fast is a run?

A single video needs a page request, a player request, and at least one caption request. A channel run adds listing pages and processes each candidate video. Time therefore varies with the number of videos examined, caption availability, and temporary connection delays. Use a small candidate limit for quick checks.

Can I schedule or monitor a channel?

Yes. Save a channel input with a publication date range and schedule it to run again. Compare videoId values in successive datasets. A fixed date range helps avoid scanning the whole public listing when you only need recent uploads.

Can I export results to a spreadsheet?

Yes. The dataset can be exported as CSV or Excel as well as JSON. The transcript array remains nested in JSON; if you need one spreadsheet row per caption segment, expand that array after export and retain the parent videoId on each segment row.

Why is transcriptText empty?

That field is optional and is null unless include_transcript_text is enabled. The transcript array still contains the full retrieved text with timing. If no caption content is available, the video does not produce a successful item.

Why are some counts empty or rounded?

The public page does not always expose every count. An absent comment count is null, not zero. When a like or subscriber count is shown in abbreviated form, the saved integer represents that displayed rounded value. Use the count as a broad indicator rather than an exact historical measurement.

Integrations

Use the default dataset with the platform API, webhooks, scheduling, or workflow services. JSON works well for systems that preserve nested transcript segments. CSV and Excel are useful for video-level reporting. A webhook can notify an ingestion process when a channel run finishes, and a scheduled run can keep a catalog of public caption text current without manually copying subtitles from each video.

Support

Open an issue on the Issues tab; we reply within 24h and add fields on request.