Youtube Text & Metadata Scraper avatar

Youtube Text & Metadata Scraper

Pricing

from $2.90 / 1,000 results

Go to Apify Store
Youtube Text & Metadata Scraper

Youtube Text & Metadata Scraper

Extract YouTube transcripts, AI summary, subtitles, video metadata, hashtags, thumbnails, views, duration, and release dates from YouTube videos using either a search query or a list of direct YouTube URLs. It helps you turn YouTube search results and specific video links into structured JSON data

Pricing

from $2.90 / 1,000 results

Rating

0.0

(0)

Developer

Fabio Borsotti

Fabio Borsotti

Maintained by Community

Actor stats

1

Bookmarked

41

Total users

8

Monthly active users

4.6 hours

Issues response

a day ago

Last modified

Share

YouTube Transcript Scraper Actor

Extract YouTube transcripts, AI summary, subtitles, video metadata, hashtags, thumbnails, views, duration, and release dates from YouTube videos using either a search query or a list of direct YouTube URLs. This YouTube transcript scraper helps you turn YouTube search results and specific video links into structured JSON data for research, SEO, lead generation, monitoring, and content analysis.

What does YouTube Transcript Scraper do?

YouTube Transcript Scraper is an APIFY Actor that can search YouTube videos by keyword, apply a time filter, collect video metadata, and retrieve the first available transcript based on your preferred language order. It can also process direct YouTube video URLs and include those videos in the final output.

This Actor is ideal if you want to:

  • train RAG algorithms for AI
  • scrape YouTube transcripts
  • extract YouTube subtitles
  • collect YouTube video metadata
  • monitor YouTube search results
  • build datasets from YouTube videos
  • automate YouTube content research
  • extract transcripts from specific YouTube videos

Why use this YouTube scraper?

If you work with YouTube data, transcripts are one of the most valuable sources of structured content. Video transcripts help you analyze what creators actually say, not just what appears in titles and descriptions.

This Actor is useful because it supports two collection modes in the same run:

  • search by keyword using query
  • direct processing of specific YouTube videos using direct_url

This makes it practical both for broad monitoring and for targeted transcript extraction from known videos.

Common use cases

  • SEO research for YouTube videos and keywords
  • competitor monitoring on YouTube
  • AI training and text dataset preparation
  • content repurposing from video to text
  • lead generation from niche YouTube channels
  • trend monitoring by date range
  • transcript-based topic clustering
  • YouTube video catalog enrichment
  • transcript extraction from manually selected videos

What data does the Actor extract?

For each processed YouTube video, the Actor can return:

  • YouTube video ID
  • video title
  • video URL
  • channel name
  • channel URL
  • transcript text with timestamps
  • available subtitles metadata
  • video view count
  • video duration in seconds
  • release date
  • thumbnail URL
  • keywords / hashtags
  • likes count
  • comments count
  • video description
  • video categories
  • live stream indicator
  • geographical availability
  • video language
  • chapters (if available)
  • series, season, and episode information

Note: By default, the Actor extracts transcript only and returns empty values for metadata fields. To include video metadata and subtitles, set enableMetadata: true in the input. See the "Metadata Retrieval Strategy" section for details.

Output Fields Reference

FieldTypeDescriptionAvailability
idstringYouTube video IDAlways
titlestringVideo title (short form)Always
urlstringYouTube video watch URLAlways
autorstringChannel/creator name (from search)Always
textstringVideo transcript with timestamps in format [MM:SS]Always if available. It uses before list languages specified. If not available get the first language available in subtitle list
channel_namestringChannel/uploader nameWith enableMetadata = true
channel_urlstringURL to the channel pageWith enableMetadata = true
video_titlestringFull video titleWith enableMetadata = true
video_urlstringYouTube video URLWith enableMetadata = true
views_countintegerTotal number of video viewsWith enableMetadata = true
duration_secondsintegerVideo duration in secondsWith enableMetadata = true
release_datestringVideo publish date in format YYYY-MM-DDWith enableMetadata = true
thumbnail_urlstringURL to the video thumbnail imageWith enableMetadata = true
hashtagsarrayHashtags associated with the videoWith enableMetadata = true
like_countintegerNumber of likes on the videoWith enableMetadata = true
comment_countintegerNumber of comments on the videoWith enableMetadata = true
descriptionstringFull video description textWith enableMetadata = true
keywordsarrayKeywords associated with the videoWith enableMetadata = true
categoriesarrayVideo categories/topicsWith enableMetadata = true
is_livebooleanWhether the video is a live streamWith enableMetadata = true
availabilitystringGeographical availability statusWith enableMetadata = true
languagestringPrimary language code of the videoWith enableMetadata = true
chaptersarrayChapter markers with timestamps (if available)With enableMetadata = true
seriesstringSeries name if video is part of a seriesWith enableMetadata = true
seasonintegerSeason number if video is part of a seriesWith enableMetadata = true
episodeintegerEpisode number if video is part of a seriesWith enableMetadata = true
subtitlesarrayAvailable subtitle/transcript languages with metadataWith enableMetadata = true
summarystringAI-generated summary of the video transcript using Gemini flash lite (always the latest version).With summarize = true

Why this Actor is useful

This YouTube Transcript Scraper is useful when you need structured YouTube data without building and maintaining your own scraping workflow. It is designed for users who want fast access to transcript-rich video data in a reusable JSON format.

Compared with a basic YouTube metadata scraper, this Actor focuses on transcript extraction and subtitle discovery, which makes it especially helpful for:

  • content intelligence
  • SEO workflows
  • machine learning pipelines
  • research automation
  • enrichment of video datasets

Input

Metadata Retrieval Strategy (enableMetadata flag)

By default, the Actor extracts transcripts only for optimal performance and minimal proxy credit consumption.

However, you can enable comprehensive metadata retrieval by setting enableMetadata: true:

  • enableMetadata: false (default): Fast, low-credit mode

    • Extracts: transcript text + video URL
    • Output fields for metadata are present but empty (None, [])
    • Recommended for large-scale transcript extraction
  • enableMetadata: true: Full metadata mode

    • Extracts: transcript + video metadata + available subtitles
    • Metadata fields populated: channel name/URL, view count, duration, release date, thumbnail, hashtags
    • Subtitles field includes full list of available subtitle languages
    • Useful for comprehensive video data collection

When enableMetadata: true, the Actor emits a single metadata-enabled event per run to notify external systems.

Example: Enabling metadata

{
"query": "apify tutorial",
"limit": 5,
"langs": ["it", "en"],
"enableMetadata": true,
"file_output": "output.json"
}

Supported input fields

  • direct_url (array of strings, optional): list of direct YouTube video URLs to process. Supported formats include youtube.com/watch?v=..., youtu.be/..., and youtube.com/shorts/...
  • query (string, optional): YouTube search query. Optional if direct_url is provided
  • range (array of strings): time filter for YouTube search, one of hour, today, this_week, this_month, this_year
  • limit (integer): maximum number of videos to process from search results
  • langs (array of strings): preferred language order used when selecting transcripts (see table before)
  • file_output (string): name of the JSON file saved in the key-value store
  • enableMetadata (boolean, optional): if true, fetches video metadata and subtitles; if false or omitted, extracts transcript only (default: false)
  • summarize (boolean, optional): if true, summarizes video transcript using Gemini Flash Lite (always the latest) into output field summary (default: false)
  • max_words (integer, optional): maximum word count of the generated AI summary (default: 300)

AI Summarization Strategy (summarize and max_words flags)

When summarize: true, the Actor uses the google.genai library to call Google Gemini service to summarize each video's transcript.

  • Environment Variables:

    • GEMINI_KEY: API key for accessing Gemini AI service
    • MODEL: Gemini model name (default: gemini-flash-lite-latest)
    • AI_SUMMARIZE_PROMPT: Custom prompt template for Gemini summarization (default: "Summarize the following YouTube video transcript in maximum {max_words} words:\n\n{text}")
    • PARALLEL: Concurrency limit for parallel summarization of video transcripts
  • Event Emission: When summarize: true, after saving the dataset and key-value store output, the Actor iterates through the output items and emits an ai-enabled Apify event for each returned video.

Example: Enabling AI Summarization

{
"query": "apify tutorial",
"limit": 5,
"langs": ["it", "en"],
"summarize": true,
"max_words": 300,
"file_output": "output.json"
}

At least one between query and direct_url must be provided.

Supported languages (iso ISO 639-1)

CodiceLinguaCodiceLinguaCodiceLinguaCodiceLinguaCodiceLinguaCodiceLingua
aaAfarabAbkhazaeAvestanafAfrikaansakAkanamAmharic
arArabicasAssameseavAvaricayAymaraazAzerbaijanibaBashkir
beBelarusianbgBulgarianbiBislamabmBambarabnBengaliboTibetan
brBretonbsBosniancaCatalanceChechenchChamorrocoCorsican
crCreecsCzechcuChurch SlaviccvChuvashcyWelshdaDanish
deGermandvDivehidzDzongkhaeeEweelGreekenEnglish
eoEsperantoesSpanishetEstonianeuBasquefaPersianffFulah
fiFinnishfjFijianfoFaroesefrFrenchfyWestern FrisiangaIrish
gdScottish GaelicglGaliciangnGuaraniguGujaratigvManxhaHausa
heHebrewhiHindihoHiri MotuhrCroatianhtHaitianhuHungarian
hyArmenianhzHereroiaInterlinguaidIndonesianieInterlingueigIgbo
iiSichuan YiikInupiaqioIdoisIcelandicitItalianiuInuktitut
jaJapanesejvJavanesekaGeorgiankgKongokiKikuyukjKuanyama
kkKazakhklKalaallisutkmCentral KhmerknKannadakoKoreankrKanuri
ksKashmirikuKurdishkvKomikwCornishkyKirghizlaLatin
lbLuxembourgishlgGandaliLimburganlnLingalaloLaoltLithuanian
luLuba-KatangalvLatvianmgMalagasymhMarshallesemiMaorimkMacedonian
mlMalayalammnMongolianmrMarathimsMalaymtMaltesemyBurmese
naNaurunbNorwegian BokmålndNorth NdebeleneNepalingNdonganlDutch
nnNorwegian NynorsknoNorwegiannrSouth NdebelenvNavajonyChichewaocOccitan
ojOjibwaomOromoorOriyaosOssetianpaPanjabipiPali
plPolishpsPashtoptPortuguesequQuechuarmRomanshrnRundi
roRomanianruRussianrwKinyarwandasaSanskritscSardiniansdSindhi
seNorthern SamisgSangoshSerbo-CroatiansiSinhalaskSlovakslSlovenian
smSamoansnShonasoSomalisqAlbaniansrSerbianssSwati
stSouthern SothosuSundanesesvSwedishswSwahilitaTamilteTelugu
tgTajikthThaitiTigrinyatkTurkmentlTagalogtnTswana
toTongatrTurkishtsTsongattTatartwTwityTahitian
ugUighurukUkrainianurUrduuzUzbekveVendaviVietnamese
voVolapükwaWalloonwoWolofxhXhosayiYiddishyoYoruba
zaZhuangzhChinesezuZulu

Example input with direct URLs only

{
"direct_url": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"https://youtu.be/9bZkp7q19f0"
],
"langs": ["it", "en"],
"file_output": "output.json"
}

Example input with query and direct URLs

{
"direct_url": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ"
],
"query": "apify tutorial",
"range": ["this_week"],
"limit": 5,
"langs": ["it", "en"],
"file_output": "output.json"
}

Example input with direct URLs only and enableMetadata = true

{
"direct_url": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"https://youtu.be/9bZkp7q19f0"
],
"langs": ["it", "en"],
"enableMetadata": true,
"file_output": "output.json"
}

Example input with query, enableMetadata = true & summarize = true

{
"enableMetadata": true,
"file_output": "output.json",
"langs": [
"en"
],
"limit": 1,
"max_words": 300,
"query": "weather forecast new york",
"range": [
"today"
],
"summarize": true
}

When both direct_url and query are provided, all videos listed in direct_url are processed and added to the output together with the search results. When enableMetadata is set to true, the Actor will fetch comprehensive metadata and available subtitles for each video and emit an apify-actor-metadata event upon completion.

Output

Each dataset item contains structured YouTube transcript, metadata & AI summary fields like these:

[
{
"id": "7mn-Slhcwkg",
"title": "Nor'easter impacting Tri-State Area | 12:15 p.m. Sunday update",
"url": "https://www.youtube.com/watch?v=7mn-Slhcwkg",
"autor": null,
"text": "[00:01] This is a First Alert Weather Special\n[00:04] Report.\n[00:05] >> Hey, good afternoon. It's uh 12:18. Hope\n[00:07] you enjoyed the uh point with Marsha\n[00:09] Kramer. I love that show. Um uh weather\n[00:11] show, I I don't know if anybody loves\n[00:15] this. It's so dreary. I love this shot,\n[00:17] but we could because we're talking about\n[00:18] like the a low [music] cloud deck. There\n[00:20] it is. Uh it's like we're playing with\n[00:22] half a deck. It's just dreary out there.\n[00:25] And then we're going to I'm just going\n[00:26] to show you here's another picture. So,\n[00:27] here's a camera that we positioned in\n[00:29] Where is the city? Where is the city? Uh\n[00:32] murky out there. And\n[00:34] >> [laughter]\n[00:34] >> it's still 62°. It's been 62°\n[00:38] for about 4 hours. Winds are gusting at\n[00:40] 32. That's a change. We'll talk about\n[00:42] that in just a minute. By the way,\n[00:44] thanks for spending some time with us\n[00:45] here on CBS News New York. Uh\n[00:48] visibility, uh we don't have any dense\n[00:49] fog advisories or anything like that.\n[00:51] Just was talking with our friends on the\n[00:53] desk and uh LaGuardia is still really\n[00:55] struggling. It's of our big three, it's\n[00:57] the one that's hurting the the the most.\n[00:59] Look at the Teterboro, it just went up\n[01:00] as far as visibility. They had a a\n[01:03] 59-minute ground delay due to visibility\n[01:05] earlier. Then there was a ground stop.\n[01:07] Now they're back to a delay. And it's\n[01:09] wind and visibility there. So, yeah, I\n[01:12] mean, you it's going to be challenging,\n[01:15] certainly not impossible. You're going\n[01:16] to get into a situation where it's just\n[01:18] a matter of making sure the planes are\n[01:20] there if we see major delays, but JFK\n[01:22] and Newark right now, okay. 62 in the\n[01:25] city and not that much warmer anywhere\n[01:27] else except out on the island. You have\n[01:29] that wind, and so we're actually that's\n[01:32] probably going to be the warm spot uh\n[01:34] for the afternoon. Now, the wind. Wind\n[01:36] is very important. You get a sense of\n[01:38] where the center of the low is. So, it's\n[01:40] come on board. It actually will make\n[01:42] another turn, and then it weakens, and\n[01:46] that's not until Monday. So, we are\n[01:48] stuck with this thing. Watch these winds\n[01:50] now stronger for parts of the Jersey\n[01:52] shore, and not as bad out on the East\n[01:54] End. Now, this is where we saw the 50s\n[01:56] and 60s. Those that wind is over. Here's\n[01:59] another snapshot. This is a composite\n[02:01] radar\n[02:03] and satellite. Look at this. The bands\n[02:05] stretch from almost, you know, the\n[02:07] border all the way down to the Delmarva.\n[02:11] So, it this is a big system, very well\n[02:14] organized, too. In the city, uh this\n[02:16] little batch is filling in right now.\n[02:18] We're on the west side and it's coming\n[02:19] down again. It's better in Queens right\n[02:21] now than it is in Manhattan or the\n[02:22] Bronx. And then over the sound into\n[02:24] parts of Westchester and Fairfield, it's\n[02:26] just showers. It's just dreary. I mean,\n[02:28] it's like the wipers can't make up their\n[02:30] mind. Are they on or are they on\n[02:32] intermittent? This heavier rain is to\n[02:34] our south, but this is what's going to\n[02:36] fill in. I was just talking to uh one of\n[02:39] my colleagues and friends on the desk\n[02:41] leaving the city at 4:00 to head out to\n[02:43] the island. Well, the concern is some of\n[02:45] this rain shield will fill in. So, let's\n[02:48] dive ahead. So, just uh setting this\n[02:50] ahead about an hour or so. Let let let\n[02:53] let let check on my times here. Um\n[02:55] this means it's probably a little wet at\n[02:57] Yankee Stadium, maybe okay uh for the\n[03:00] Giants, but there there'll be on again,\n[03:02] off again showers during the course of\n[03:04] both games. But, we're talking showers.\n[03:07] We don't see any convection with this,\n[03:08] so there's no lightning risk or anything\n[03:10] like that. I mean, I know the NFL's got\n[03:12] a different standard than the MLB, but\n[03:15] it's just going to be kind of gray and\n[03:16] wet at times. Now, games are over, but\n[03:19] you've got places to go. Here comes that\n[03:22] shield that that that leading edge of\n[03:24] that impulse that we showed you. So,\n[03:26] this is a pretty good estimate. Maybe\n[03:29] it'll verify a little farther to the\n[03:31] east, but this comes on board between\n[03:33] 5:00 and say 8:00 or 9:00 with pockets\n[03:36] of heavier rain. So, if you're wrapping\n[03:38] up a weekend trip or going out for the\n[03:40] evening, dinner tonight in Brooklyn,\n[03:42] yeah, anywhere, it's going to be a\n[03:43] little wet. And see how this takes over\n[03:45] tonight. So, Vanessa's coming in this\n[03:47] afternoon. She'll be watching this and\n[03:49] she'll take care of you on the the side\n[03:50] tonight. But, you're going to see more\n[03:52] rain falling a man right through Monday\n[03:54] morning. Early commuters,\n[03:56] yeah, have the rain gear ready and you\n[03:58] might need it coming and going. So, this\n[04:00] is for the kids. It's going to be wet\n[04:02] not for everybody in Jersey, but for the\n[04:03] city kids up in the Westchester, it's\n[04:06] still wet and then poof.\n[04:08] Lunchtime, dining outside. Well, if\n[04:10] you're a duck,\n[04:12] um look at that. There's more rain in\n[04:14] the parts of Connecticut. See there it\n[04:15] is. See what happened? It came here,\n[04:18] backed out, it wobbles a bit, then it\n[04:20] takes off. And so, it's not all the way\n[04:22] until\n[04:24] wow, Tuesday morning before we start to\n[04:26] see a break in this thing start to clear\n[04:28] out. So, just real quickly we're\n[04:31] updating this. Woah, now we're at 375\n[04:33] for Toms River. And if that makes its\n[04:35] way around, I mean, we could add another\n[04:37] inch to that. So, I do think we'll\n[04:39] pretty confident that there'll be a lot\n[04:41] of threes up to fives for totals for\n[04:44] this. The concern, even though the winds\n[04:46] are subsiding, is you still are dealing\n[04:48] with vulnerable trees. So, just be on\n[04:50] your game right now. I want to make sure\n[04:51] that everybody's ready for that. That's\n[04:53] why it's a first alert weather day\n[04:54] today. Just the heads-up too for Monday.\n[04:57] If you have a flight early tomorrow\n[04:58] morning or someplace you absolutely\n[05:00] positively have to be, have to factor in\n[05:02] extra time tomorrow morning. There's\n[05:04] going to be cleanup going on too. And\n[05:06] there's a better day for that. Tuesday,\n[05:08] Wednesday, Thursday of next week looks\n[05:10] so much better. It is getting better. I\n[05:13] mean, we're through with the worst of\n[05:15] this, but be mindful that there's still\n[05:17] rain. It's still going to slow you down\n[05:19] as you're heading out there and we're\n[05:20] not done with all of the heavy rain just\n[05:22] yet. So, that's why we're not done here.\n[05:24] We'll be back around 1:00 with another\n[05:26] update right here on CBS News New York.\n[05:31] >> This has been a first alert weather\n[05:33] special report.",
"summary": "This First Alert Weather Special reports a large, well-organized low-pressure system bringing dreary, wet conditions, gusty winds, and flight delays to the New York area. Rain will persist through Monday, with total accumulations potentially reaching 3 to 5 inches, before finally clearing by Tuesday.",
"channel_name": "CBS New York",
"channel_url": "https://www.youtube.com/channel/UCNZyLULUQBp5e9Q1cKtvk6Q",
"video_title": "Nor'easter impacting Tri-State Area | 12:15 p.m. Sunday update",
"video_url": "https://www.youtube.com/watch?v=7mn-Slhcwkg",
"views_count": 18170,
"duration_seconds": 338,
"release_date": "2026-09-27",
"thumbnail_url": "https://i.ytimg.com/vi/7mn-Slhcwkg/maxresdefault.jpg",
"hashtags": [
"CBSN New York",
"CBS News New York",
"John Elliott",
"First Alert Weather",
"Nor'easter",
"New York",
"Local TV"
],
"like_count": 126,
"comment_count": 17,
"description": "CBS News New York's John Elliott has a look at the First Alert Weather forecast. \n\nFor video licensing inquiries, contact: licensing@veritone.com",
"keywords": [],
"categories": [
"News & Politics"
],
"is_live": false,
"availability": "public",
"language": "en",
"chapters": [],
"series": null,
"season": null,
"episode": null,
"subtitles": [
{
"language_code": "en",
"language": "English (auto-generated)",
"is_generated": true,
"is_translatable": true
}
]
}
]

The Actor also saves the complete output array into the APIFY key-value store using the file name specified in file_output.

How the YouTube transcript extraction works

The Actor follows this strategy:

  1. If direct_url is provided, process each direct YouTube URL and extract transcript and metadata.
  2. If query is provided, search YouTube videos using the selected time filter.
  3. For each video, try to find a manually created transcript in the languages you requested.
  4. If no manual transcript is available, try an auto-generated transcript.
  5. If no requested language is available, fall back to the first available transcript.
  6. Save transcript text and structured metadata in the output dataset.

This makes the Actor practical for multilingual transcript scraping, targeted video extraction, and broad topic monitoring.

Who is this Actor for?

This Actor is a good fit for:

  • SEO specialists
  • marketers
  • data engineers
  • content analysts
  • AI data learning teams
  • researchers
  • agencies monitoring YouTube niches
  • users who need transcripts from specific YouTube links

Limitations

The actor has the following limitations:

  • Public videos only: Private or restricted videos are not supported.
  • Transcripts required: The Actor can only extract data from videos that have transcripts enabled.
  • Age-gated videos not supported: Videos marked as age-restricted (18+, 14+, etc.) cannot be accessed without YouTube account authentication. The Actor will return empty metadata and transcript for age-gated videos because YouTube blocks unauthenticated access to these videos for protection purposes.

To work with age-gated videos, you would need to implement YouTube account authentication, which is not currently supported by this Actor.

Summary

If you need a YouTube transcript scraper for APIFY that extracts subtitles, transcript text, hashtags, thumbnails, views, duration, and release date from either search results or direct video URLs, this Actor gives you a clean starting point with structured JSON output and proxy support.