# YouTube AI Video Summarizer — Gemini Brief, Q\&A & Transcript (`yugenox/youtube-ai-video-summarizer`) Actor

AI summaries of YouTube videos with Google Gemini: a timestamped brief (summary, key points, chapters, topics), answers to your own questions, JSON-schema extraction and an optional AI transcript. Video links, channels or a search. A brief is $0.02 per 10 minutes of video. No login.

- **URL**: https://apify.com/yugenox/youtube-ai-video-summarizer.md
- **Developed by:** [Yugenox Corp](https://apify.com/yugenox) (community)
- **Categories:** AI, Videos, Social media
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $20.00 / 1,000 ai video read (per 10 min)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube AI Video Summarizer — Gemini Brief, Q\&A & Transcript

**Turn YouTube videos into briefs, answers and transcripts with AI.** Paste video links, a channel or a search. Google's Gemini AI watches and listens to each video and gives you a summary, the key points and chapters with clickable timestamps, answers to your own questions, your own JSON fields, and if you want, a full timestamped transcript.

> 🤖 Gemini AI reads the video itself (picture + sound) · ⏱️ Every moment links to that second of the video · 🌍 Output in 25 languages · 🗣️ Transcripts also for videos without captions · 💵 A brief is $0.02 per 10 minutes of video, transcripts $0.004 per minute · 🚫 No login, no API key

### 🎯 What you get for each video

| Output | What it is |
|---|---|
| **AI brief** (on by default) | `tldr` (one sentence), `summary` (at most 120 words), 3–8 `keyPoints` with timestamps, `chapters` with timestamps (YouTube's own when the creator made them, otherwise written by the AI), `topics` |
| **Your questions** | up to 5 questions answered from each video, each answer with the timestamp where it is said or shown. If the video doesn't answer a question, it says so and you are not charged for it |
| **Your JSON schema** | your own fields (for example products shown, prices, a verdict), filled from each video into `extracted` |
| **AI transcript** | a timestamped transcript of everything said, in the language spoken, written by the AI from the audio. Works for videos without captions |

Every timestamp comes as `startSeconds`, a readable `timestamp` (`3:16`) and a `url` such as `https://youtu.be/MBRqu0YOH14?t=196` that opens the video at that second. The AI's timestamps are checked against the video's length, and any that fall outside the video are removed.

The brief, your questions and your schema all come from **one AI read** of each video. YouTube's own data is in the same row: title, channel, publish date, views, likes, comment count, length, description and thumbnail.

### 🚀 How to use it

**Summarize videos.** Paste links: watch links, `youtu.be` links, Shorts or video ids.

```json
{ "videoUrls": ["https://www.youtube.com/watch?v=MBRqu0YOH14", "https://youtu.be/dQw4w9WgXcQ"] }
```

**Ask your own questions**, and get the brief in Spanish:

```json
{ "videoUrls": ["https://www.youtube.com/watch?v=MBRqu0YOH14"], "questions": ["What does optimistic nihilism mean?", "Which books are recommended?"], "outputLanguage": "es" }
```

**Pull structured data** from product reviews, using your own JSON schema (or just `{"field": "what to put in it"}`):

```json
{ "searchQueries": ["best robot vacuum 2026"], "maxVideos": 5, "includeBrief": false,
  "extractionSchema": { "products": "Products reviewed, with brand", "winner": "The product the reviewer recommends", "pricesMentioned": "Prices said in the video" } }
```

**Get transcripts** of a channel's latest videos up to 30 minutes long:

```json
{ "channels": ["https://www.youtube.com/@kurzgesagt"], "maxVideos": 10, "maxVideoMinutes": 30, "includeBrief": false, "includeTranscript": true }
```

Other options: `channelSortBy` (`newest` or `popular`), `searchSortBy` (`relevance`, `date` or `views`), `country` (the country searches run from), and `outputLanguage` (`en`, `es`, `pt`, `fr`, `de`, `hi`, `ja`, … or `original`, which uses the language spoken in the video).

### 📦 Output

A real row from 2026-10-04, with the lists shortened:

```json
{
  "id": "MBRqu0YOH14",
  "url": "https://www.youtube.com/watch?v=MBRqu0YOH14",
  "title": "Optimistic Nihilism",
  "channelName": "Kurzgesagt – In a Nutshell",
  "date": "2017-07-26",
  "viewCount": 18765885,
  "likes": 980826,
  "durationSeconds": 446,
  "duration": "7:26",
  "aiStatus": "done",
  "spokenLanguage": "en",
  "tldr": "Optimistic nihilism suggests that because the universe has no inherent meaning or purpose, we are free to create our own meaning and enjoy our brief existence.",
  "summary": "Humanity's existence is brief and insignificant in the vast, ancient universe. While this realization can cause existential dread, it also offers a liberating perspective. …",
  "keyPoints": [
    { "text": "The realization that the universe has no inherent meaning can be liberating rather than depressing.", "startSeconds": 196, "timestamp": "3:16", "url": "https://youtu.be/MBRqu0YOH14?t=196" }
  ],
  "chapters": [
    { "title": "The vastness of the universe", "startSeconds": 0, "timestamp": "0:00", "url": "https://youtu.be/MBRqu0YOH14?t=0" },
    { "title": "Optimistic nihilism", "startSeconds": 182, "timestamp": "3:02", "url": "https://youtu.be/MBRqu0YOH14?t=182" }
  ],
  "chaptersSource": "ai",
  "topics": ["Philosophy", "Nihilism", "Existentialism", "Universe", "Humanity"],
  "answers": [
    { "question": "What does optimistic nihilism mean?", "found": true, "answer": "Optimistic nihilism is the idea that because the universe has no inherent meaning or purpose, we are free to create our own meaning and enjoy our lives.", "startSeconds": 196, "timestamp": "3:16", "url": "https://youtu.be/MBRqu0YOH14?t=196" }
  ],
  "transcript": [
    { "startSeconds": 1, "timestamp": "0:01", "text": "Human existence is scary and confusing." }
  ],
  "transcriptText": "Human existence is scary and confusing. A few hundred thousand years ago, we became conscious …",
  "transcriptLanguage": "en",
  "aiModel": "gemini-3.1-flash-lite",
  "aiCharged": { "ai-video": 1, "ai-answer": 1, "ai-transcript-minute": 8 },
  "input": "https://www.youtube.com/watch?v=MBRqu0YOH14",
  "inputType": "url"
}
```

- `aiStatus` tells you what happened. `done` means you got everything you asked for. `partial` means you got part of it. `skipped` means the video was not sent to the AI (too long, live or upcoming, or its AI read would not finish before your run's timeout). `refused` means the AI could not open the video because it isn't public. `failed` means the AI request failed. `off` means the AI was unavailable on this run. `aiNote` explains anything missing in plain words.
- `aiCharged` shows what each row was charged.
- `chaptersSource` is `youtube` when the creator made chapters (those are used exactly as they are), otherwise `ai`.
- `answers[].found` is `false` when the video doesn't answer the question. Those answers are free.
- The AI writes the brief and answers. They are usually right, but check anything important by following the timestamp link.

The dataset has five table views: Briefs, Key moments, Answers, Transcript and YouTube details. The run's key-value store holds `RUN_REPORT` (how the run ended, the videos skipped and why, and what was charged), plus `FAILED_INPUTS` when needed (channels that don't exist, unavailable videos, and inputs YouTube blocked, none of them charged).

### 💵 Pricing

You pay per event, per video. Minutes are counted on the whole video.

| Event | When | Price |
|---|---|---|
| Actor start | once per run (per GB of memory; the default 256 MB = 1) | $0.002 |
| AI video read (per 10 min) | per started 10 minutes of each video the AI read for your brief, questions or JSON schema (one read covers all three) | $0.02 |
| AI answer | each question the video answers; your JSON schema filled = 1 per started 1,000 characters of filled JSON (a typical 5–10 field schema = 1) | $0.002 |
| AI transcript minute | per started minute of each video transcribed (only when AI transcript is on) | $0.004 |

Examples:

- A brief of a 7-minute video costs $0.002 + $0.02 = **$0.022**.
- A 25-minute video with a brief and 2 answered questions costs $0.002 + 3 × $0.02 + 2 × $0.002 = **$0.066**.
- A transcript of a 7½-minute video costs $0.002 + 8 × $0.004 = **$0.034**.
- Briefs of a channel's 5 latest videos (each under 10 minutes) cost $0.002 + 5 × $0.02 = **$0.102**.

**What is never charged:**

- videos longer than your maximum length, and videos skipped because your run's timeout left too little time for the AI;
- live streams and upcoming premieres;
- private, deleted or unavailable videos, and videos the AI cannot open;
- videos where the AI failed or gave an unusable answer;
- questions the video doesn't answer, and a JSON schema the video gives nothing for or the AI rejects;
- everything on a run where the AI is unavailable.

One exception: a transcript is charged for the minutes the AI listened to, even when nobody speaks and the transcript comes back empty.

If you set a maximum cost per run, the actor never sends a video to the AI unless your limit covers everything that video could be charged.

### ❓ FAQ

**Do I need a YouTube account, cookies, or a Gemini or OpenAI API key?**
No. The actor reads YouTube's public pages like any logged-out visitor, and the AI runs on our Gemini account. You don't need any key.

**How does the AI read the video?**
It uses Gemini's official YouTube video input: Google fetches the public video and the AI watches it (about one frame every 5 seconds, at low resolution) and listens to the full soundtrack. Nothing is downloaded. Because the AI hears the audio itself, transcripts and summaries also work for videos without captions.

**Which videos work?**
Public videos and Shorts. Private and deleted videos are listed in `FAILED_INPUTS`. Videos the AI is not allowed to open come back with `aiStatus` `refused`: unlisted, members-only, age-restricted and region-locked videos, and some live-stream replays. Neither kind is charged. The maximum length is 180 minutes; the default is 60, and longer videos are skipped and not charged.

**What do I send to Google, and is my data private?**
To make each video's output, the actor sends Google's Gemini API the video's public link, its title, channel name, length, description (first 1,500 characters) and chapter list, and your questions and JSON schema. Nothing else about you or your account is sent. This goes through Google's paid Gemini API under Google's terms. YouTube data is read without logging in.

**Is it legal to use this on YouTube videos?**
The actor only reads what anyone can see on YouTube without signing in. It does not download or redistribute videos. It returns summaries, short quotes and transcripts for analysis. Collecting public data is generally legal, but you're responsible for how you use the results. Transcripts and summaries can be covered by the creator's copyright, so use them for research, analysis, accessibility or your own notes, not to republish someone else's content. Results can include names and channel handles, which can count as personal data, so follow privacy laws such as GDPR, PIPEDA and CCPA. If you're unsure, ask a lawyer. More on this: [Is web scraping legal?](https://blog.apify.com/is-web-scraping-legal/)

**How accurate is it?**
The AI is good at summarizing, but it can make mistakes, especially with names, numbers and timestamps. Every point and answer has a timestamp link so you can check it in seconds. Timestamps outside the video are removed automatically. The transcript is written by the AI from the audio, not copied from YouTube's captions.

**Can I get the brief in another language?**
Yes. Set `outputLanguage` to one of 25 languages, or to "Same as the video". The transcript always stays in the language spoken in the video.

**What if the AI is unavailable?**
You still get rows with YouTube's details for the first 5 videos, with `aiStatus` set to `off`. The run stops there and nothing is charged beyond the start. Run it again later.

**What happens if YouTube blocks a request?**
Each request is retried on fresh IPs, starting with Apify datacenter IPs and then the proxy in your input (residential by default). A video that stays blocked is listed in `FAILED_INPUTS` and not charged. If YouTube blocks the run from the start, the run stops at once and says so, instead of spending your money.

**What if my run's timeout is short?**
The actor stops starting new videos shortly before your timeout and keeps every row already saved. A video whose AI read would not finish in the time left (estimated from its length) comes back with YouTube's details, `aiStatus` `skipped` and a note, is not charged, and the actor goes on to the next video. The AI usually takes 5–15 seconds per video but sometimes 2 minutes, so give runs with many videos a few minutes.

**Is this affiliated with YouTube or Google?**
No. It is not affiliated with, endorsed by or sponsored by YouTube or Google. YouTube and Gemini are trademarks of Google LLC.

# Actor input Schema

## `videoUrls` (type: `array`):

YouTube videos for the AI to read: watch links, youtu.be links, Shorts links or 11-character video ids. Only public videos can be read. A channel or search link pasted here is read as a channel or a search. One per line.

## `channels` (type: `array`):

Channels whose latest (or most popular) videos to read: a channel link (https://www.youtube.com/@NASA), an @handle, a channel id (UC…) or a channel name. One per line.

## `searchQueries` (type: `array`):

Keywords to search YouTube videos for; the top results are read. One per line.

## `maxVideos` (type: `integer`):

How many videos to read from each channel and each search (video links are always read in full). Videos longer than "Max video length" are passed over and do not count.

## `maxVideoMinutes` (type: `integer`):

Videos longer than this are not sent to the AI and are not charged (a video you gave by link still gets a row with YouTube's details). Up to 180 minutes.

## `includeBrief` (type: `boolean`):

A one-sentence TL;DR, a summary (at most 120 words), 3–8 key points with timestamps, chapters with timestamps (YouTube's own when the video has them) and topic tags. Every time comes with a youtu.be link that opens the video at that moment.

## `questions` (type: `array`):

Up to 5 questions the AI answers from each video (at most 300 characters each), with the timestamp where the answer is. A question the video does not answer says so and is not charged. One per line.

## `extractionSchema` (type: `object`):

Optional: a JSON schema such as {"type":"object","properties":{"products":{"type":"array","items":{"type":"string"}},"verdict":{"type":"string","enum":\["buy","skip"]}}}, or simply {"field": "what to put in it"}. The AI fills it from each video into "extracted". Supported: type, properties, items, required, enum, description (up to 40 fields, 4 levels deep). Text only, never code.

## `includeTranscript` (type: `boolean`):

A timestamped transcript of everything said, written by the AI from the video's audio in the language spoken — also for videos without captions. Charged per started minute of video.

## `outputLanguage` (type: `string`):

The language the brief, the answers and the extracted text are written in, whatever language the video is in. "Same as the video" keeps the language spoken in the video. The transcript always stays in the language spoken.

## `channelSortBy` (type: `string`):

Which videos of a channel to read: the latest uploads or the most popular ones.

## `searchSortBy` (type: `string`):

Order of search results: most relevant first (YouTube's default), newest first, or most viewed first.

## `country` (type: `string`):

Two-letter country code (US, GB, IN, BR, DE …). Searches run as YouTube does for people in that country.

## `proxyConfiguration` (type: `object`):

YouTube's pages are read through Apify datacenter IPs first (cheapest); this proxy (residential by default) is used only for the kinds of request YouTube blocks on datacenter IPs. An account without residential proxies falls back to Apify datacenter IPs automatically. The AI reads the videos itself and never uses the proxy.

## Actor input object example

```json
{
  "videoUrls": [
    "https://www.youtube.com/watch?v=MBRqu0YOH14"
  ],
  "maxVideos": 5,
  "maxVideoMinutes": 60,
  "includeBrief": true,
  "includeTranscript": false,
  "outputLanguage": "en",
  "channelSortBy": "newest",
  "searchSortBy": "relevance",
  "country": "US",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All videos read in this run, with the AI output.

## `run` (type: `string`):

Status and statistics for this run.

## `report` (type: `string`):

RUN_REPORT: how the run ended, videos per input, skipped videos and why, AI calls and what was charged.

## `failedInputs` (type: `string`):

FAILED_INPUTS: channels not found, unavailable videos and inputs YouTube blocked (none of them charged). Absent when there were none.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://www.youtube.com/watch?v=MBRqu0YOH14"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("yugenox/youtube-ai-video-summarizer").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videoUrls": ["https://www.youtube.com/watch?v=MBRqu0YOH14"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("yugenox/youtube-ai-video-summarizer").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://www.youtube.com/watch?v=MBRqu0YOH14"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call yugenox/youtube-ai-video-summarizer --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,yugenox/youtube-ai-video-summarizer"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UtDkMGukf4XXOgPy1/builds/jeE3zhirfvLJlaKfk/openapi.json
