# YouTube Transcript & Video Analytics Scraper (`bgfc97/youtube-transcript-analytics`) Actor

Pull the full transcript (captions) of any public YouTube video plus video and channel analytics - no login, no API key. Returns full-text transcript + timestamped segments, views, likes, channel and publish date. Pure HTTP, great for AI/content pipelines.

- **URL**: https://apify.com/bgfc97/youtube-transcript-analytics.md
- **Developed by:** [Bruno](https://apify.com/bgfc97) (community)
- **Categories:** Videos, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$4.00 / 1,000 video processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Transcript & Analytics Scraper

Pull the **full transcript** of any public YouTube video — plus rich video and channel
analytics — with no login and no YouTube Data API key. Built for AI/content pipelines:
feed real spoken content straight into summarizers, RAG stores, translation or repurposing
workflows, alongside honest view/like/upload metadata.

### What it does

- **Transcript (the premium feature)** — reads the video's actual caption track (manual or
  auto-generated), in your preferred language when available, and returns both a clean
  full-text `transcript` and timestamped `transcriptSegments` (`{start, dur, text}`). Two
  independent retrieval paths are tried automatically for every video: (1) YouTube's direct
  caption (`timedtext`) API in three response formats, and (2) as a fallback, the same
  internal `get_transcript` call YouTube's own website makes to render the transcript panel.
- **Video metadata** — title, channel, view count, like count, upload/publish date,
  category, duration, description, keywords, thumbnails, family-safe flag, available
  countries, and the list of all caption languages the video actually has.
- **Channel scraping** — give it a channel URL or `@handle` and it lists the channel's
  recent videos (title, view count text, published-time text, duration, thumbnail) and —
  if you want — fetches the full metadata + transcript for each of those videos too.

### Input

```json
{
  "videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
  "channelUrls": ["https://www.youtube.com/@MrBeast"],
  "languages": ["en"],
  "maxVideos": 30,
  "includeTranscript": true,
  "proxyConfiguration": { "useApifyProxy": true },
  "timeoutSecs": 30
}
```

- `videoUrls` — watch URLs, `youtu.be` links, `/shorts/` links, or bare 11-char video IDs.
- `channelUrls` — channel URLs (`/@handle`, `/channel/UC...`, `/c/Name`, `/user/Name`) or
  bare `@handle`.
- At least one of `videoUrls` / `channelUrls` is required.
- `languages` — ordered language-code preference for the transcript (e.g. `["en","pt"]`).
  A real (human-made) caption track in one of these languages is preferred; then an
  auto-generated one in these languages; and if neither exists the actor **honestly falls
  back** to whichever caption track the video actually has, reporting exactly which one via
  `transcriptLanguage` / `transcriptIsAutoGenerated` — it never invents transcript text.
- `maxVideos` — how many recent videos to pull per channel (first page of the Videos tab).
- `includeTranscript` — set `false` to skip transcripts and only get metadata/listings
  (faster, cheaper on proxy bandwidth).
- `proxyConfiguration` — Apify Proxy; **datacenter proxy works fine for YouTube**, no need
  for residential.

### Output

One dataset item per video, plus one summary item per channel.

**Video item** (from `videoUrls`, or expanded from a channel when `includeTranscript=true`) —
metadata fields are always real when the video loads; the `transcript*` fields reflect
whichever of the two retrieval paths above actually succeeded for that video **at request
time** (see "Known limitation" below):

```json
{
  "type": "video",
  "videoId": "dQw4w9WgXcQ",
  "title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
  "author": "Rick Astley",
  "viewCount": 1820040849,
  "likeCount": null,
  "publishDate": "2009-10-24T23:57:33-07:00",
  "category": "Music",
  "keywords": ["rick astley", "..."],
  "channelSubscriberCountText": "4.54M subscribers",
  "availableCaptionLanguages": [
    { "languageCode": "en", "name": "English", "isAutoGenerated": false },
    { "languageCode": "en", "name": "English (auto-generated)", "isAutoGenerated": true },
    { "languageCode": "pt-BR", "name": "Portuguese (Brazil)", "isAutoGenerated": false }
  ],
  "has_transcript": true,
  "transcript": "We're no strangers to love. You know the rules and so do I...",
  "transcriptSegments": [{ "start": 18.32, "dur": 3.02, "text": "We're no strangers to love" }],
  "transcriptLanguage": "en",
  "transcriptIsAutoGenerated": false,
  "transcriptSource": "timedtext",
  "transcriptError": null
}
```

When neither retrieval path succeeds for a given video (see limitation below), you instead
honestly get `"has_transcript": false, "transcript": null, "transcriptSource": null,
"transcriptError": "innertube: HTTP 400 fetching ..."` — with every metadata field above
still fully populated and real.

**Channel item:**

```json
{
  "type": "channel",
  "channelUrl": "https://www.youtube.com/@MrBeast",
  "title": "MrBeast",
  "subscriberCountText": "410M subscribers",
  "videoCountText": "700 videos"
}
```

If a video has no captions at all (or YouTube would not serve them to this request),
`has_transcript` is `false`, `transcript` is `null`, and `transcriptError` explains exactly
why — **the actor never fabricates transcript text.** `availableCaptionLanguages` is always
returned honestly from the video's own player data, independent of whether the transcript
body itself could be fetched, so you can always see which languages exist even on a run
where the text couldn't be pulled.
If a video/channel fails to load (blocked, region-locked, removed, not found), you get an
honest `{ input, error }` item instead of partial/guessed data.

### Known limitation (read before relying on 100% transcript coverage)

YouTube has been progressively tightening anti-bot checks on caption-serving endpoints.
In testing, both retrieval paths above can currently return an empty/`400 FAILED_PRECONDITION`
response for some videos even though the video genuinely has captions (confirmed via
`availableCaptionLanguages`) — this reproduces identically with or without a proxy, so it is
YouTube-side request validation, not a proxy quality issue, and not specific to this actor's
code path. When this happens the item is still fully honest: `has_transcript: false` and a
`transcriptError` describing which mechanism failed and why, never invented text. Metadata,
view/like counts, dates, and channel listings are unaffected and reliable regardless.
If YouTube's enforcement eases (it fluctuates) or for videos it doesn't trigger on, the
actor returns the real transcript automatically — no input change needed.

### Notes

- Source: YouTube's own public, server-rendered pages (`ytInitialPlayerResponse` /
  `ytInitialData`) and YouTube's own caption/transcript endpoints — no private API, no login.
- Every HTTP request uses a fresh proxy session and retries (up to 4 attempts) if YouTube
  responds with a block (403/429 or an "unusual traffic" page).
- Channel scraping reads only the **first page** of the Videos tab (no pagination token
  yet) — the returned `channelItem.notes` field says so honestly when the channel has more
  videos than were returned. Both of YouTube's current channel-grid JSON layouts (classic
  `videoRenderer` and the newer `lockupViewModel`) are supported.
- `likeCount` is best-effort (parsed from the page's own accessibility label) and is `null`
  when YouTube doesn't expose it in a reliably parseable way — never guessed.

### ⭐ Enjoying this Actor?

A quick **rating/review** helps others find it. Want multi-page channel pagination or
playlist support added? Open a ticket on the **Issues** tab.

# Actor input Schema

## `videoUrls` (type: `array`):

YouTube watch URLs (https://www.youtube.com/watch?v=...), short youtu.be links, /shorts/ links, or bare 11-character video IDs to scrape individually (metadata + transcript). Leave empty if you only want channel videos. At least one of videoUrls or channelUrls is required.

## `channelUrls` (type: `array`):

Channel URLs (https://www.youtube.com/@handle, /channel/UC..., /c/Name, /user/Name) or bare @handles whose recent videos should be listed and scraped. Leave empty if you only want individual videos. At least one of videoUrls or channelUrls is required.

## `languages` (type: `array`):

Language codes in order of preference for the transcript, e.g. \["en", "pt", "es"]. A real (manually created) caption track in one of these languages is picked first; if none exists, an auto-generated one in these languages is used; if still none, the actor honestly falls back to whatever caption track the video actually has (and reports which one via transcriptLanguage / transcriptIsAutoGenerated) rather than fabricating text.

## `maxVideos` (type: `integer`):

Maximum number of recent videos to fetch per channel (from the channel's Videos tab, most recent first, first page only).

## `includeTranscript` (type: `boolean`):

If true (default), fetch and parse the caption track (transcript) for every video — both videos passed in videoUrls and every video discovered via channelUrls. If false, only metadata/listing fields are returned (faster, fewer requests, cheaper on your proxy bandwidth).

## `proxyConfiguration` (type: `object`):

Apify Proxy configuration used for all requests to YouTube. Using Apify Proxy (datacenter proxy is sufficient for YouTube) is strongly recommended to avoid rate limiting and IP blocks.

## `timeoutSecs` (type: `integer`):

Timeout in seconds for each individual HTTP request made to YouTube (page loads and caption-track downloads), 5-90.

## Actor input object example

```json
{
  "languages": [
    "en"
  ],
  "maxVideos": 30,
  "includeTranscript": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "timeoutSecs": 30
}
```

# Actor output Schema

## `dataset` (type: `string`):

One row per video: full transcript + timestamped segments, plus title, author, channel, views, likes, length and publish date.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("bgfc97/youtube-transcript-analytics").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("bgfc97/youtube-transcript-analytics").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call bgfc97/youtube-transcript-analytics --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bgfc97/youtube-transcript-analytics"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BtJS21b7Kzv3x5B1R/builds/yCYfI2rtVd03jECUN/openapi.json
