# Twitch & Kick VOD Transcript Scraper - SRT & Text (`deapi/twitch-kick-transcript-scraper`) Actor

Transcribe Twitch and Kick VODs, clips and full broadcasts to text with AI speech recognition. Word-level timestamps, speaker labels for co-streams, SRT and VTT subtitles, 90+ languages, up to 10 hours per recording. Mix Twitch and Kick in one bulk run. No login, no API key, no downloads.

- **URL**: https://apify.com/deapi/twitch-kick-transcript-scraper.md
- **Developed by:** [deAPI](https://apify.com/deapi) (community)
- **Categories:** Videos, AI, Social media
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $9.00 / 1,000 transcript, per minute of videos

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

Get an accurate, timestamped **Twitch transcript** or **Kick transcript** from any finished broadcast. Paste one VOD link or a hundred, and the **Twitch & Kick VOD Transcript Scraper** returns the full text, sentence or word-level timings, speaker labels, and ready-to-use SRT and VTT subtitles. No Twitch or Kick login, no API key, no cookies, nothing to download.

[Twitch](https://www.twitch.tv) and [Kick](https://kick.com) publish no captions at all. There is no auto-caption track to scrape, no transcript tab, no subtitle file — which is why a six-hour broadcast is, today, completely opaque to search, to research and to any tool that reads text. This Actor listens to the audio and transcribes it, so a stream becomes something you can read, search, quote and clip from.

### Why this one

A handful of tools can now put a Twitch or Kick VOD through speech recognition. What they hand back is a block of text with rough timings. This Actor times the transcript **to the word** and tells you **who was speaking** — and both of those are what turn a wall of stream text into something you can actually cut clips from or attribute quotes to. It also takes Twitch and Kick in the same run, so a research batch spanning both platforms is one job and one dataset, not two.

| | Caption scrapers | Other stream transcribers | **Twitch & Kick VOD Transcript Scraper** |
|---|---|---|---|
| Works on Twitch VODs | ❌ no caption track exists | ✅ | ✅ |
| Works on Kick VODs | ❌ | ✅ | ✅ |
| **Twitch *and* Kick in one run** | ❌ | ❌ one platform each | ✅ mixed freely |
| **Word-level timings** | ❌ | ❌ | ✅ every word, with confidence |
| **Speaker labels for co-streams** | ❌ | ❌ | ✅ included |
| Twitch clips | ❌ | Sometimes | ✅ same run, same rate |
| Handles a 10-hour broadcast | ❌ | Usually capped well below | ✅ up to 10 hours, no default cap |
| Subtitle files, properly split | ❌ | Raw blocks | ✅ SRT + VTT, two-line cues |
| Dead or expired VOD detected before you are charged | — | ❌ | ✅ always |

### What you get for every recording

- 📝 **Full transcript** — clean, punctuated text of everything spoken in the broadcast.
- ⏱️ **Timestamps at the level you choose** — one per sentence, one for every single word, or none at all if you just want the text. Sentence timing is in the base price.
- 🗣️ **Speaker labels** — who said what, for co-streams, podcasts, watch-alongs, interviews and any broadcast with guests in the call.
- 🎬 **SRT and VTT subtitles** — properly split, two lines, readable. Drop them straight into a video editor, a player or your CMS.
- 🌍 **90+ languages** — auto-detected, or forced when the audio sits under loud game sound.
- 🎮 **Broadcast metadata** — title, channel, stream date, view count, thumbnail, and on Twitch the game or category the stream was listed under.
- 🔢 **Confidence scores** — per segment and per word, so you can flag the passages worth a human check.

### How to transcribe a Twitch or Kick VOD

1. Open the Actor and paste one or more links into **Twitch or Kick video URLs**.
2. Pick a **timestamp level** — segment is the default and suits subtitles and reading.
3. Switch on **speaker labels** if the broadcast has more than one voice.
4. Click **Start**. Results appear row by row; a long broadcast reports its progress as it goes.
5. Download the dataset as JSON, CSV or Excel, or pull it from the [Apify API](https://docs.apify.com/api/v2).

#### Which links work

| Link | Works |
|---|---|
| `twitch.tv/videos/1234567890` | ✅ Twitch VOD |
| `m.twitch.tv/videos/1234567890` | ✅ rewritten for you |
| `twitch.tv/channelname/clip/ClipSlug` | ✅ Twitch clip |
| `clips.twitch.tv/ClipSlug` | ✅ resolved for you |
| `kick.com/channelname/videos/<video-id>` | ✅ Kick VOD |
| `kick.com/video/<video-id>` | ✅ resolved for you |
| `twitch.tv/channelname` | ❌ that is a channel page |
| A stream that is live right now | ❌ see below |

Paste the same recording twice in any two of those forms and it is transcribed — and billed — **once**. Share links carrying `?t=` collapse onto the same recording too.

### How fast is it?

Every job pays a fixed setup cost of roughly a minute and a half, then transcribes very fast — about 38× real time. Measured end to end:

| Recording | Finished in |
|---|---|
| 4-minute Twitch VOD | ~1 minute 45 |
| 30-minute Twitch VOD | ~2 minutes 25 |
| 1-hour Kick VOD | ~5 minutes |
| 10-hour broadcast (the ceiling) | ~17 minutes |

Several recordings run in parallel, so a batch of twenty short clips finishes in about the time one of them takes.

### Live streams are not supported, and cannot be

A broadcast in progress has no final length, and length is what a transcription job is priced and scheduled from. There is nothing to transcribe until the stream ends and the VOD exists. Paste a channel URL and the Actor tells you so immediately, at no charge, rather than failing halfway through a run.

**Twitch and Kick both delete VODs.** Retention runs from 7 days to 60 depending on the channel's plan and platform, so a broadcast you want the text of is worth transcribing sooner rather than later. Expired, private, subscriber-only and deleted recordings are all detected **before anything is charged**.

### Pricing

Charged **per minute of the recording**, rounded up, with a three-minute minimum. Nothing else is metered: platform usage is on us, and subtitles, confidence scores and sentence timestamps are all in the base price.

| What you asked for | Price per minute |
|---|---|
| Plain text, or sentence timestamps | **$0.010** |
| Word-level timings, or speaker labels, or both | **$0.017** |
| Broadcast metadata (optional add-on) | $0.005 per recording |

Word timings and speaker labels sit on the same higher rate, and **picking both costs no more than picking one**.

| Recording | Sentence timing | Word timing or speakers |
|---|---|---|
| A 30-second clip (billed as 3 min) | $0.030 | $0.051 |
| A 20-minute highlight | $0.20 | $0.34 |
| A 1-hour stream | $0.60 | $1.02 |
| A 3-hour stream | $1.80 | $3.06 |
| A 6-hour stream | $3.60 | $6.12 |

**A note on Kick:** Kick broadcasts are long. Most VODs on the platform run three to six hours, and very few are under an hour — so budget for a Kick transcript the way you would for a full stream, not for a clip. Twitch VODs and clips span the whole range.

Recordings that fail — expired, private, deleted, region-locked, too long — are **never charged**. Neither are recordings skipped because a run cost limit or a free-plan allowance was reached. Every one of them still gets a row explaining what happened, so nothing disappears silently from your dataset.

Set **maximum cost per run** to whatever you are comfortable with. The Actor checks the price of each recording before it starts work and skips anything the remaining budget cannot cover, so a run cannot overshoot the cap you set.

**Free plan:** one recording per run, 20 minutes maximum, 20 transcribed minutes per calendar month, resetting at 00:00 UTC on the 1st. Any paid Apify plan removes all three limits.

### Output example

One row per recording. Fields shown abbreviated:

```json
{
  "url": "https://www.twitch.tv/videos/69409460",
  "platform": "twitch",
  "sourceType": "vod",
  "sourceId": "69409460",
  "status": "success",
  "hasSpeech": true,
  "durationSeconds": 241,
  "language": "en",
  "languageConfidence": 0.978,
  "wordCount": 430,
  "timestampLevel": "word",
  "speakers": ["SPEAKER_00", "SPEAKER_01", "SPEAKER_02", "SPEAKER_03"],
  "text": "That fucking hitbox is bullshit. Come on, we just got 30 seconds, just hold it...",
  "segments": [
    {
      "start": 0.95,
      "end": 35.31,
      "text": "That fucking hitbox is bullshit. Come on, we just got 30 seconds...",
      "avg_logprob": -0.247,
      "speaker": "SPEAKER_02",
      "words": [
        { "word": "That", "start": 2.051, "end": 2.171, "score": -0.2866, "speaker": "SPEAKER_02" }
      ]
    }
  ],
  "srt": "1\n00:00:00,950 --> 00:00:03,525\n[SPEAKER_02] That fucking hitbox\nis bullshit.\n\n...",
  "vtt": "WEBVTT\n\n00:00:00.950 --> 00:00:03.525\n[SPEAKER_02] That fucking hitbox\nis bullshit.\n\n...",
  "billedMinutes": 5,
  "pricingOption": "precision"
}
```

With **broadcast metadata** switched on, each row also carries:

```json
{
  "metadata": {
    "platform": "twitch",
    "title": "whats my fkin name",
    "channel": "loltyler1",
    "categories": ["League of Legends"],
    "streamedOn": "2016-05-29",
    "viewCount": 135827,
    "thumbnail": "https://static-cdn.jtvnw.net/..."
  }
}
```

A very long broadcast produces more text than a single dataset row is allowed to hold. When that happens the heavy fields are written to the run's key-value store and linked as `textUrl`, `srtUrl`, `vttUrl` and `segmentsUrl`, and the row is flagged `oversized: true`. Nothing is ever truncated.

### What people use this for

| Use case | Why it needs a transcript |
|---|---|
| **Clip mining** | Find the moment someone said the thing, by searching text instead of scrubbing six hours of timeline. Word-level timings give you the exact cut point. |
| **Stream SEO** | Publish a searchable text version of a VOD so it can be indexed, quoted and cited. Streams are invisible to search engines today. |
| **Repurposing** | Turn one broadcast into show notes, a newsletter, chapter markers, shorts scripts and social posts. |
| **Community moderation and safety** | Review what was actually said on a broadcast, with timestamps, instead of relying on reports. |
| **Esports and creator research** | Track how a scene talks about a patch, a roster move or a sponsor across many channels. |
| **Sponsorship verification** | Confirm a read happened, when, and in what words. |
| **Accessibility** | Ship real subtitles for VOD re-uploads on platforms that accept SRT. |
| **RAG and AI agents** | Feed timestamped, speaker-labelled stream text into a vector store or an LLM pipeline. |

### Frequently asked questions

#### Can I get a transcript of a Twitch VOD?

Yes. Paste the VOD link — `twitch.tv/videos/1234567890` — and the Twitch & Kick VOD Transcript Scraper returns the full text with timestamps, plus SRT and VTT subtitles. Twitch publishes no caption track, so the audio is transcribed with a speech recognition model rather than copied from anywhere.

#### Does it work on Kick?

Yes, the same way. Paste `kick.com/channelname/videos/<video-id>`. Twitch and Kick links can be mixed freely in the same run — most transcript tools cover one platform each, so a cross-platform batch normally means two separate jobs. Bear in mind that Kick VODs are usually whole multi-hour broadcasts.

#### Can I transcribe a live stream?

No. A broadcast in progress has no final length, and there is no recording to work from until it ends. Wait for the VOD to appear on the channel's Videos tab, then paste that link.

#### How long a broadcast can it handle?

Up to 10 hours per recording — long enough for almost any single stream — and there is no lower default cap to raise first. Anything longer is rejected up front, at no charge, with a message saying so. Split the broadcast or transcribe the clips instead.

#### Can I transcribe Twitch clips?

Yes. Both `twitch.tv/channelname/clip/ClipSlug` and the short `clips.twitch.tv/ClipSlug` form work, at the same per-minute rate with the three-minute minimum.

#### Can I transcribe many VODs in bulk?

Yes. Paste as many links as you like into one run; several are transcribed in parallel and each produces its own row. Duplicates are collapsed before anything is charged, so pasting the same VOD in two different link formats bills once.

#### Can I get SRT or VTT subtitles?

Yes, both, free of charge, on every recording with timing. They are properly split into readable two-line cues rather than dumped as one long block, and speaker labels are included in the cue text when diarization is on.

#### Can I tell who is speaking on a co-stream?

Yes — switch on **Label speakers**. Every segment, and every word when word timings are on, carries a `SPEAKER_00`, `SPEAKER_01` … tag, and the row lists all speakers detected. On a four-way Twitch co-stream this separates the voices cleanly. No caption-based tool can do this, because there is no caption track to read speakers from.

#### What happens if a VOD has no speech?

You get a row with `hasSpeech: false` and an empty transcript, not an error. Silence, music-only and pure gameplay audio are valid outcomes. The recording is still billed, because the audio was still processed.

#### Why did a recording fail?

Almost always because it is gone or unreachable: Twitch and Kick delete VODs after a retention window, and the recording may also be private, subscriber-only, deleted or region-locked. Failures are detected before the work starts and are never charged. The reason is in the row's `error` field.

#### Which languages are supported?

More than 90, auto-detected by default. Set the language explicitly when a broadcast mixes languages or the speech sits under loud game audio — it noticeably improves accuracy.

#### Do I need a Twitch or Kick account?

No. No login, no cookies, no API key, no browser extension. Public recordings only.

#### Can I run this on a schedule or from my own code?

Yes. Every Apify Actor has a [REST API](https://docs.apify.com/api/v2), a [scheduler](https://docs.apify.com/platform/schedules), [webhooks](https://docs.apify.com/platform/integrations/webhooks) and official [JavaScript and Python clients](https://docs.apify.com/api/client/js). A common pattern is a daily run over a channel's newest VODs, writing straight into a search index or a vector store.

#### Do you have this for other platforms?

Yes — [TikTok Transcript Scraper](https://apify.com/deapi/tiktok-transcript-scraper) does the same job for TikTok videos, with the same word timings, speaker labels and subtitle output.

#### Is this legal?

The Actor reads publicly available recordings, the same ones any viewer can open. You are responsible for what you do with the output — quoting, publishing and re-using someone's broadcast is subject to copyright and to the platform's terms, exactly as it is if you transcribe it by hand.

# Actor input Schema

## `videoUrls` (type: `array`):

Links to **finished** recordings. Twitch VODs (`twitch.tv/videos/1234567890`), Twitch clips (`twitch.tv/channel/clip/ClipSlug`) and Kick VODs (`kick.com/channel/videos/<video-id>`) all work, mixed freely in one run. Short forms — `m.twitch.tv`, `clips.twitch.tv/Slug`, `kick.com/video/<id>` — are rewritten for you. A **channel page is rejected**: a live broadcast has no final length and cannot be transcribed. Maximum 10 hours per recording. Dead, private, subscriber-only and expired VODs are detected before anything is charged. **Free plan: 1 recording per run, 20 minutes maximum, 20 transcribed minutes per calendar month (resets 00:00 UTC on the 1st). Paid plans remove all three limits.**

## `timestampLevel` (type: `string`):

How finely the transcript is timed. **Segment** gives one timestamp per sentence — right for subtitles, reading and jumping into a stream at the moment something was said, and included in the base price. **Word** times every single word — right for clip cutting, highlight detection and search. **None** returns plain text only. Word moves the recording onto the precision rate ($0.017 a minute against $0.010), and choosing speaker labels as well costs nothing extra on top.

## `diarize` (type: `boolean`):

Detect who is speaking and tag every line with a speaker (SPEAKER\_00, SPEAKER\_01, …). Use it for co-streams, podcasts, interviews, watch-alongs and any broadcast with guests in the call. Speaker labels need timing to attach to, so switching this on together with **None** still returns sentence-level timing. Moves the recording onto the precision rate ($0.017 a minute against $0.010); if you already picked Word timestamps, adding speakers costs nothing extra.

## `language` (type: `string`):

Leave on auto-detect unless the broadcast mixes languages or the audio sits under loud game sound — naming the language then makes the transcript noticeably more accurate. Auto-detect covers 90+ languages; the dropdown lists the 29 most requested, and anything outside it is still transcribed on auto-detect.

## `includeSubtitles` (type: `boolean`):

Return ready-to-use subtitle files as `srt` and `vtt` fields — paste them straight into a video editor, a player or a CMS. Free of charge. Subtitles need timing to exist, so they are produced at every timestamp level except None, and also at None when speaker labels are switched on. On very long broadcasts the files are saved to the run's key-value store and linked as `srtUrl` and `vttUrl` instead of being inlined, so nothing is ever truncated.

## `includeMetadata` (type: `boolean`):

Attach the recording's own data to each row: title, channel, description, stream date, thumbnail, view count and — on Twitch — the game or category the broadcast was listed under. Which fields are present varies by platform; Twitch and Kick do not expose the same set. Turns one run into transcript + broadcast context for content research. Charged as an add-on of $0.005 per recording.

## `maxConcurrency` (type: `integer`):

How many recordings are transcribed at the same time. Higher finishes a large batch sooner. It has no effect on price — you are charged per minute of source either way, and platform usage is on us.

## Actor input object example

```json
{
  "videoUrls": [
    "https://www.twitch.tv/videos/69409460"
  ],
  "timestampLevel": "segment",
  "diarize": false,
  "language": "auto",
  "includeSubtitles": true,
  "includeMetadata": false,
  "maxConcurrency": 3
}
```

# Actor output Schema

## `transcripts` (type: `string`):

One row per recording: the full transcript text, timed segments (with word-level timings and speaker labels when requested), SRT and VTT subtitles, detected language, broadcast length, word count and the minutes billed. Every URL you submit produces a row. Recordings that could not be transcribed carry status 'failed' with an 'error' field, and are not charged. Recordings left out because a free-plan allowance or a run cost limit was reached carry status 'skipped' with the reason in 'error', and are not charged either. On very long broadcasts the heavy fields are saved as files and linked from 'textUrl', 'srtUrl', 'vttUrl' and 'segmentsUrl'.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://www.twitch.tv/videos/69409460"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("deapi/twitch-kick-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoUrls": ["https://www.twitch.tv/videos/69409460"] }

# Run the Actor and wait for it to finish
run = client.actor("deapi/twitch-kick-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://www.twitch.tv/videos/69409460"
  ]
}' |
apify call deapi/twitch-kick-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,deapi/twitch-kick-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/luAEA1axMvU86NuJe/builds/qbkWn6T5Sqbpnb5mH/openapi.json
