# YouTube AI Transcriber | No Captions Needed (`andok/youtube-ai-transcriber`) Actor

AI watches the video and transcribes every word — works on videos WITHOUT captions. Speakers, timestamps, SRT/VTT/JSON. $0.15/min (or $0.03/min with your Gemini key).

- **URL**: https://apify.com/andok/youtube-ai-transcriber.md
- **Developed by:** [Andok](https://apify.com/andok) (community)
- **Categories:** AI, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $14.00 / 1,000 video minute transcribed (your gemini key)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## YouTube AI Transcriber — no captions needed

Gemini AI **watches the actual video** — audio and frames — and writes a verbatim transcript with **timestamps and speaker labels**. It works on videos that have **no captions at all**, in any spoken language, and exports ready-to-use **SRT and WebVTT subtitle files** plus JSON and plain text.

Caption scrapers can only copy captions that already exist. This actor transcribes the video itself.

### What it does

- ✅ Transcribes videos **without captions** — AI listens to and watches the video
- ✅ **Speaker labels** — real names when identifiable on screen, otherwise Speaker 1/2/…
- ✅ **Timestamps** on every segment (`MM:SS` / `HH:MM:SS`)
- ✅ **4 output formats**: SRT subtitles, WebVTT subtitles, JSON segments, plain text — each as a downloadable file
- ✅ Any spoken language, up to **150 minutes** per video
- ✅ No API key needed — or bring your own Gemini key and pay 5× less

### Pricing

Charged **per minute of video** (minimum 5 minutes per video), plus a $0.02 actor start fee.

| Video length | No API key (bundled) | With your Gemini key |
|---|---|---|
| Short (≤5 min) | $0.75 | $0.15 |
| 10 min | $1.50 | $0.30 |
| 30 min | $4.50 | $0.90 |
| 60 min | $9.00 | $1.80 |

**Bring your own key and save 5×**: create a free key at [aistudio.google.com/apikey](https://aistudio.google.com/apikey) and paste it into the `geminiApiKey` field (stored encrypted, used only for your run). Google token costs (a few cents per hour of video) are then billed to your Google account.

You are only charged when a transcript is actually delivered — failed or timed-out runs charge nothing beyond the start fee.

### Input

| Field | Type | Description |
|---|---|---|
| `videoUrl` | string (required) | Public YouTube video URL (watch, youtu.be, or Shorts link) |
| `includeSpeakers` | boolean | Label each segment with the speaker (default `true`) |
| `outputFormats` | array | Any of `json`, `srt`, `vtt`, `text` (default: all) |
| `language` | string | Optional language hint, e.g. `"Estonian"` |
| `model` | string | `gemini-3.7-flash` (default) or `gemini-3.5-flash-lite` |
| `geminiApiKey` | secret string | Optional — your Google AI Studio key for 5× cheaper runs |

### Output

One dataset item per run:

```json
{
  "videoUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
  "videoTitle": "Me at the zoo",
  "durationSeconds": 19,
  "minutesCharged": 5,
  "language": "English",
  "segmentCount": 4,
  "transcript": "[00:01] Jawed Karim: Alright, so here we are in front of the, uh, elephants...",
  "segments": [
    { "start": "00:01", "end": "00:04", "speaker": "Jawed Karim", "text": "Alright, so here we are in front of the, uh, elephants." }
  ],
  "srtUrl": "https://api.apify.com/v2/key-value-stores/…/records/transcript-jNQXAC9IVRw.srt",
  "vttUrl": "…", "jsonUrl": "…", "textUrl": "…",
  "status": "completed"
}
```

The SRT/VTT/JSON/text files are saved to the run's key-value store with public download URLs — drop the SRT straight into a video editor or player.

### Use it via API

```bash
curl -X POST "https://api.apify.com/v2/acts/andok~youtube-ai-transcriber/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"videoUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw"}'
```

### FAQ

**How is this different from YouTube transcript scrapers?** Scrapers download captions that already exist. This actor has Gemini watch the video and transcribe it from the audio and visuals — so it works when captions are missing, auto-captions are garbage, or you need speaker labels.

**What videos are supported?** Public, finished YouTube videos up to 120 minutes (150 minutes with your own key). Private, unlisted, and live videos are not supported.

**What languages?** Any language Gemini understands — the transcript is written in the video's original spoken language.

**How long does a run take?** Typically well under the video's own duration — a 10-minute video usually transcribes in 1–3 minutes.

**What if the run fails?** You are never charged per-minute fees for a failed or undelivered transcript.

# Actor input Schema

## `videoUrl` (type: `string`):

Public YouTube video to transcribe (youtube.com/watch, youtu.be, or Shorts link). Private and unlisted videos are not supported.

## `includeSpeakers` (type: `boolean`):

Label each segment with the speaker (real names when identifiable, otherwise Speaker 1, Speaker 2, …).

## `outputFormats` (type: `array`):

Transcript files saved to the run's key-value store, each with a public download URL in the results.

## `language` (type: `string`):

Expected spoken language, e.g. "Estonian". The transcript is always in the video's original language; this only helps accuracy.

## `model` (type: `string`):

Bundled-key runs allow gemini-3.7-flash and gemini-3.5-flash-lite. The other models require your own geminiApiKey.

## `geminiApiKey` (type: `string`):

Bring your own Google AI Studio key (aistudio.google.com/apikey) and pay only $0.03 per video minute instead of $0.15; Google token costs are then billed to your Google account. Leave empty to use the bundled key.

## Actor input object example

```json
{
  "videoUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
  "includeSpeakers": true,
  "outputFormats": [
    "json",
    "srt",
    "vtt",
    "text"
  ],
  "model": "gemini-3.7-flash"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `resultsCsv` (type: `string`):

No description

## `files` (type: `string`):

No description

## `run` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw"
};

// Run the Actor and wait for it to finish
const run = await client.actor("andok/youtube-ai-transcriber").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw" }

# Run the Actor and wait for it to finish
run = client.actor("andok/youtube-ai-transcriber").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw"
}' |
apify call andok/youtube-ai-transcriber --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,andok/youtube-ai-transcriber"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xQHG77yMpzgaqBXzj/builds/zg9eFPmELSddotwNl/openapi.json
