# Video to Text: TikTok, Instagram Reels & MP4 Transcriber (`fguiraud/video-to-text-transcriber`) Actor

Convert video to text: transcribe TikTok videos, Instagram Reels, X and Facebook videos and MP4, MOV, WEBM files (or Google Drive / Dropbox links) to a full transcript, timestamps and SRT/VTT subtitles in 99 languages, or translate to English. Open-source Whisper, no API key, pay per minute.

- **URL**: https://apify.com/fguiraud/video-to-text-transcriber.md
- **Developed by:** [Fernando Guiraud](https://apify.com/fguiraud) (community)
- **Categories:** Videos, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.00 / 1,000 audio minutes

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Video to Text Transcriber do?

**Video to Text Transcriber** converts **video to text**: it transcribes **TikTok videos, Instagram Reels, X (Twitter) and Facebook videos** and **MP4, MOV, WEBM, MKV, AVI** files into a **full transcript**, **timestamped segments**, **Markdown with timestamps** and **readable SRT/VTT subtitles** (42 characters per line, 2 lines, ready to upload), in **99 languages**, and can **translate the speech to English**. Paste a **TikTok, Instagram, X or Facebook link**, a direct link to a file, or a **Google Drive, Dropbox or OneDrive share link**; audio files work too.

It runs **open-source Whisper** on Apify, so there is no OpenAI key and no subscription. You **pay per minute of video**, and failed files or videos without speech are never billed.

It runs on the Apify platform, so you also get an API, scheduling, integrations (Make, Zapier, n8n, Google Drive, Notion) and access for **AI agents through the [Apify MCP server](https://mcp.apify.com)**.

![Sample output: real rows from a run of this Actor](https://fernandoguiraud16-coder.github.io/data-tools/assets/outputs/output-video-to-text-transcriber.png)

> **Independent tool**, not affiliated with or endorsed by the platforms it reads. TikTok, Instagram, X and Facebook are trademarks of their respective owners; they are named only to describe the public data this Actor collects. It only accesses publicly available content.

### Why use it?

- 📱 **TikTok and Instagram Reels**: the spoken words of any public video, with its title, author, date, views and likes. Study hooks and scripts of viral videos, or repurpose your own.
- 🎓 **Courses, webinars and lectures**: searchable text and notes from every recording.
- 🎥 **Content repurposing**: turn videos into blog posts, newsletters, show notes and social posts.
- 🧑‍💼 **Meetings and interviews**: transcripts of Zoom, Teams or Meet recordings, with an optional **AI summary, chapters and action items** (Claude, with your own Anthropic key).
- 🔎 **Research and compliance**: make video archives searchable.
- 🤖 **RAG and AI assistants**: timestamped text chunks ready for a vector database.

### How to convert video to text

1. Click **Try for free**.
2. Paste links to TikTok videos, Instagram Reels or video files (or Google Drive / Dropbox share links).
3. Choose the **spoken language** (or `auto`) and what to return: text, SRT subtitles, Markdown with timestamps…
4. Click **Start**. With the default model, transcription runs several times faster than real time.
5. Download the transcript as JSON, CSV or Excel, or the `.txt` / `.srt` files from the links in the results.

### Input

| Field | Description | Default |
|---|---|---|
| `sources` | TikTok / Instagram links, video (or audio) URLs and share links | required (or `profiles` / `base64Files`) |
| `profiles` | TikTok accounts (`@username` or profile link): transcribe their newest videos | - |
| `maxVideosPerProfile` | Newest videos per TikTok account | 10 |
| `onlyNewVideos` | Skip TikTok videos transcribed by earlier runs | on |
| `language` | Spoken language code, or `auto` | `auto` |
| `task` | `transcribe` or `translate` (to English) | `transcribe` |
| `model` | `base` (recommended), `small` (most accurate) or `tiny` | `base` |
| `outputs` | `text`, `markdown`, `segments`, `srt`, `vtt`, `chunks` | text, srt |
| `subtitleMaxChars` | Characters per subtitle line (0 = raw Whisper segments) | 42 |
| `vocabulary` | Names and jargon to spell correctly | - |
| `aiInsights` + `anthropicApiKey` | AI summary, chapters and action items | off |
| `maxDurationMinutes` | Transcribe (and bill) at most N minutes per file | 240 |

```json
{
  "sources": [{ "url": "https://example.com/webinar.mp4" }],
  "outputs": ["text", "markdown", "srt"]
}
```

### Output

One record per video. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

```json
{
  "source": "https://example.com/webinar.mp4",
  "status": "ok",
  "language": "en",
  "durationSeconds": 310.6,
  "transcribedSeconds": 60,
  "billedMinutes": 1,
  "text": "My fellow Americans, Prime Minister Nakasoni of Japan, will be visiting me here at the White House next week. …",
  "srt": "1\n00:00:07,500 --> 00:00:11,840\nMy fellow Americans, Prime Minister\nNakasoni of Japan, will be visiting me\n…",
  "stats": { "words": 140, "speedFactor": 15.2 }
}
```

### TikTok, Instagram, X and Facebook video to text

Paste the link of a public **TikTok video**, **Instagram Reel / video post**, **X (Twitter) post with a video** or **Facebook video / Reel**. The Actor downloads only the audio track, transcribes it with Whisper and adds the video details:

```json
{
  "source": "https://www.instagram.com/reel/C0hQSaMpD97/",
  "status": "ok",
  "language": "en",
  "durationSeconds": 60.02,
  "billedMinutes": 1,
  "text": "Hey Space Nerds, I got a special delivery for you. So picture this. You're on a military range in the middle of the Utah desert. …",
  "video": {
    "platform": "instagram",
    "author": "NASA Goddard",
    "uploadDate": "2023-12-06",
    "likes": 8391,
    "comments": 17
  }
}
```

#### Transcribe a whole TikTok account

Put accounts in **TikTok accounts** (`profiles`) instead of pasting links one by one. The Actor lists the newest videos of each account and transcribes them, one row per video with its date, views and likes:

```json
{
  "profiles": ["@nasa"],
  "maxVideosPerProfile": 10,
  "onlyNewVideos": true
}
```

Schedule it weekly: with `onlyNewVideos` (on by default) videos transcribed in earlier runs are skipped, so you only pay for new ones. Videos with only music are marked as having no speech and are not billed. Instagram accounts cannot be listed without a login, so for Instagram paste the Reel links.

A typical 15 to 60-second video costs **$0.006** (one minute). Private videos, videos that need a login and videos that cannot be downloaded are not billed; the result explains why. YouTube is not supported here: use [YouTube Transcript Scraper](https://apify.com/fguiraud/youtube-transcript-scraper).

### How much does video transcription cost?

| Event | Price |
|---|---|
| Run start (per GB of memory, default 4 GB) | $0.0005 |
| Minute of video, base or tiny model | **$0.006** ($0.36 per hour) |
| Minute of video, small model | **$0.012** ($0.72 per hour) |
| AI insights per video (optional; Claude usage billed to your own key) | $0.01 |

A 1-hour webinar costs about **$0.36** with the base model. Failed files and videos without speech are free.

### How accurate is it? Whisper tiny vs base vs small

We measured the three models ourselves, with the same code this Actor runs:

| Model | Word error rate (lower is better) | Speed on Apify (default 4 GB) | Price per minute |
|---|---|---|---|
| `tiny` | 11.3% | about 20–30× faster than real time | $0.006 |
| `base` (default) | 8.2% | about 11× faster than real time | $0.006 |
| `small` | **5.4%** | about 4× faster than real time | $0.012 |

Word error rate was measured on 73 clean English recordings from the LibriSpeech test set (8.6 minutes of audio). Speed was measured on a 2.6-minute speech recording in Apify runs (tiny: scheduled test runs). Accents, background noise and other languages give higher error rates, so choose `small` for those. A 60-minute video takes about 5–6 minutes with `base` and about 15 minutes with `small`.

### Use it with AI agents (MCP)

Add `https://mcp.apify.com?tools=fguiraud/video-to-text-transcriber` to Claude, Cursor or any MCP client and ask: *"Transcribe this webinar and give me the key takeaways: https://…/webinar.mp4"*.

### Related tools

- Need **subtitles** with proper line lengths for publishing? [SRT Subtitle Generator](https://apify.com/fguiraud/srt-subtitle-generator).
- **YouTube** videos? [YouTube Transcript Scraper](https://apify.com/fguiraud/youtube-transcript-scraper) returns existing captions in seconds.
- Podcasts by name or RSS feed: [Podcast Transcript Scraper](https://apify.com/fguiraud/podcast-transcript-scraper).

### FAQ and limitations

- **Which links work?** TikTok, Instagram, X (Twitter), Facebook, direct file links and cloud-drive share links. Other video platforms (YouTube, Vimeo) are not downloaded; use the YouTube tool above for YouTube.
- **Who is speaking?** Speaker labels are not included yet.
- **Long files**: up to 240 minutes per file by default; set `maxDurationMinutes` to cap cost.
- Found a problem or need a feature? Open an issue on the **Issues** tab. Replies within 48 hours. If the Actor saved you time, a quick review helps others find it.

# Changelog

This Actor's version history is a separate document: https://apify.com/fguiraud/video-to-text-transcriber/changelog.md

# Actor input Schema

## `sources` (type: `array`):

Links to your videos: TikTok, Instagram (Reels, posts), X (Twitter) and Facebook (videos, Reels, Watch) links, or direct links to media files (MP4, MOV, MKV, WEBM, MP3, M4A, WAV and more). Google Drive, Dropbox, OneDrive and GitHub share links work too (shared as 'Anyone with the link'). YouTube is not supported: use YouTube Transcript Scraper for YouTube videos.

## `profiles` (type: `array`):

TikTok accounts to transcribe, as '@username' or a profile link (https://www.tiktok.com/@username). The newest videos of each account are transcribed (see 'Videos per account'). Instagram profiles are not supported: paste the Reel links in 'Video URLs' instead.

## `maxVideosPerProfile` (type: `integer`):

How many of the newest videos to transcribe from each TikTok account.

## `onlyNewVideos` (type: `boolean`):

Skip TikTok videos already transcribed by previous runs of this Actor in your account (remembered in a key-value store named 'transcriber-state-video'). Ideal for a weekly schedule: you only pay for new videos.

## `base64Files` (type: `array`):

Short media files without a URL, e.g. from an AI agent: \[{"fileName": "memo.m4a", "content": "<base64>"}]. Keep the total input under ~9 MB; use URLs for longer recordings.

## `model` (type: `string`):

'base': good accuracy, fast (recommended). 'small': best accuracy, especially for accents, noisy audio and non-English speech; slower and billed at a higher per-minute price. 'tiny': fastest draft quality.

## `language` (type: `string`):

ISO code of the spoken language (en, es, de, fr, pt, it, ja, zh, ...) or 'auto' to detect it. Setting it avoids misdetection on short clips.

## `task` (type: `string`):

'transcribe': text in the spoken language. 'translate': translate the speech to English text.

## `vocabulary` (type: `array`):

Words the speech recognition should favour: people and company names, product names, technical terms (e.g. 'Kubernetes', 'Dr. Nguyen', 'Apify'). Improves spelling of rare words.

## `outputs` (type: `array`):

'text': full transcript split into paragraphs at pauses. 'segments': timestamped segments. 'srt' / 'vtt': ready-to-use subtitle files. 'chunks': ~chunkSize-character passages with start/end times and a token estimate, ready for vector databases. 'markdown': paragraphs prefixed with their start time, e.g. '**\[00:01:23]** ...'.

## `subtitleMaxChars` (type: `integer`):

Characters per subtitle line in SRT/VTT (42 is the broadcast and YouTube standard; 32-37 for vertical video). 0 keeps Whisper's long raw segments.

## `subtitleMaxLines` (type: `integer`):

Maximum lines shown at once in each subtitle (when line length is set).

## `subtitleMaxDuration` (type: `number`):

Longest time a single subtitle stays on screen (when line length is set).

## `aiInsights` (type: `boolean`):

Analyse each transcript with Claude: title, summary, key points, chapters with start times, action items and topics (in 'insights'). Requires your Anthropic API key; Claude usage is billed to your Anthropic account, plus one small 'AI insights' event per file.

## `anthropicApiKey` (type: `string`):

Your key from console.anthropic.com. Stored as a secret input; used only to call Claude for this run.

## `insightsModel` (type: `string`):

'claude-opus-5': best quality (default). 'claude-sonnet-5': cheaper, great for meetings and podcasts. 'claude-haiku-4-5': cheapest (transcripts up to ~2 hours).

## `insightsInstructions` (type: `string`):

Optional, e.g. 'Summarise in Spanish', 'Focus on decisions and owners', 'Chapters every ~10 minutes'.

## `saveFiles` (type: `boolean`):

Save the transcript (.txt) and subtitles (.srt / .vtt, if selected in outputs) as files in the run's key-value store; the result includes their download links.

## `wordTimestamps` (type: `boolean`):

Add start/end times for every word inside each segment (for karaoke-style captions or precise search). Slightly slower.

## `chunkSize` (type: `integer`):

Target size of 'RAG chunks'.

## `skipSilence` (type: `boolean`):

Detect speech first and skip silent parts. Faster and reduces hallucinated text in long pauses.

## `maxDurationMinutes` (type: `integer`):

Only the first N minutes of each file are transcribed (and billed).

## `maxFileSizeMb` (type: `integer`):

Larger files are skipped (not billed).

## `failOnError` (type: `boolean`):

Mark the run as FAILED when a file cannot be transcribed. Useful for pipelines and monitoring.

## Actor input object example

```json
{
  "sources": [
    {
      "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422"
    },
    {
      "url": "https://www.instagram.com/reel/C0hQSaMpD97/"
    }
  ],
  "maxVideosPerProfile": 10,
  "onlyNewVideos": true,
  "model": "base",
  "language": "auto",
  "task": "transcribe",
  "outputs": [
    "text",
    "srt"
  ],
  "subtitleMaxChars": 42,
  "subtitleMaxLines": 2,
  "subtitleMaxDuration": 6,
  "aiInsights": false,
  "insightsModel": "claude-opus-5",
  "saveFiles": true,
  "wordTimestamps": false,
  "chunkSize": 1000,
  "skipSilence": true,
  "maxDurationMinutes": 240,
  "maxFileSizeMb": 1000,
  "failOnError": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `files` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sources": [
        {
            "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422"
        },
        {
            "url": "https://www.instagram.com/reel/C0hQSaMpD97/"
        }
    ],
    "outputs": [
        "text",
        "srt"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fguiraud/video-to-text-transcriber").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sources": [
        { "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422" },
        { "url": "https://www.instagram.com/reel/C0hQSaMpD97/" },
    ],
    "outputs": [
        "text",
        "srt",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("fguiraud/video-to-text-transcriber").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sources": [
    {
      "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422"
    },
    {
      "url": "https://www.instagram.com/reel/C0hQSaMpD97/"
    }
  ],
  "outputs": [
    "text",
    "srt"
  ]
}' |
apify call fguiraud/video-to-text-transcriber --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fguiraud/video-to-text-transcriber"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/oW3CZs4SiiHUmTLZv/builds/1hasK8z2wMjN5G9gJ/openapi.json
