Video Transcript & AI Summary — YouTube & TikTok
Pricing
from $5.00 / 1,000 transcript (captions)s
Video Transcript & AI Summary — YouTube & TikTok
Transcripts with timestamps for YouTube and TikTok videos, from captions or speech recognition. Arabic dialects and English, optional translation and AI summary with chapters and keywords.
Pricing
from $5.00 / 1,000 transcript (captions)s
Rating
0.0
(0)
Developer
Al Moutasem Nabil
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
0
Monthly active users
6 days ago
Last modified
Categories
Share
Transcripts with timestamps for YouTube and TikTok videos, from captions or speech recognition. Arabic dialects and English, optional translation and AI summary with chapters and keywords.
تفريغ الفيديو وترجمة وملخص — يوتيوب وتيك توك. يدعم اللهجات الخليجية والشامية والمصرية، والملخص يُكتب بالعربية الفصحى مع الحفاظ على اللهجة كما نُطقت في النص الأصلي.
Who it's for
- Content teams repurposing long videos into clips, posts and newsletters.
- Researchers and analysts who need searchable text from many videos at once.
- Marketers studying hooks and calls to action across competitors' short-form video.
- AI agents and automations that need a transcript as a synchronous API call (see Standby below).
What makes it different
- YouTube and TikTok in one schema. Both return the same fields, and channel or playlist URLs expand to their latest videos.
- Captions first, speech recognition second. When a platform serves captions the Actor uses them: exact, already timed, and far cheaper than transcribing audio. When it does not, the audio is transcribed, so every public video with speech gets a transcript.
- Arabic as a first-class language. Dialect is preserved verbatim in the transcript; summaries are written in Modern Standard Arabic. Keyword extraction normalizes alef, yaa, taa marbuta and tashkeel so "أحمد" and "احمد" count as one word.
- Clean sentences, not caption fragments. Auto-captions arrive as an overlapping rolling window ("so today", "so today we're", "so today we're building"); the Actor merges them back into sentences while keeping the start time of the first fragment, so timestamps stay accurate.
- Nothing is stored or rehosted. Audio is extracted at 16 kHz mono only for transcription and deleted immediately, on failures too. The output is text.
Input
This is the Console prefill: one short YouTube video and one TikTok, with an AI summary. It finishes in under a minute and costs about $0.05:
{"videoUrls": ["https://www.youtube.com/watch?v=jNQXAC9IVRw","https://www.tiktok.com/@scout2015/video/6718335390845095173"],"summarize": true,"proxyConfiguration": {"useApifyProxy": true}}
Add "translateTo": "ar" (or any ISO code) for a translation, "summaryLanguage": "ar" for an Arabic
summary, and "language": "ar" to help speech recognition on dialect. YouTube channel and playlist
URLs are expanded to their most recent videos (maxVideosPerSource).
Output
One dataset item per video. Every field is documented in .actor/dataset_schema.json, and the
Console shows three views: Transcripts, Summaries and Errors.
{"platform": "youtube","videoId": "aircAruvnKk","url": "https://www.youtube.com/watch?v=aircAruvnKk","title": "But what is a neural network?","author": "3Blue1Brown","publishedAt": "2017-10-05T00:00:00.000Z","durationSec": 1134,"language": "en","transcriptSource": "captions","asrProvider": null,"segments": [{ "startSec": 12.4, "endSec": 15.8, "text": "This is a 3, and it's sloppily written." }],"text": "This is a 3, and it's sloppily written.\nBut it's still recognisable.","wordCount": 1842,"translation": null,"summary": null,"processingMs": 4210,"scrapedAt": "2026-09-07T01:20:00.000Z"}
transcriptSource tells you what you paid for: captions and auto-captions are the cheap path,
asr means the audio was transcribed per minute.
Every video that produces no transcript, and every link that is not a YouTube or TikTok video, is
stored as an error item with the failing stage and a readable reason (for example
Video is 64 minutes, longer than maxDurationMinutes=60) — and is never charged. A requested
translation or summary that fails is reported the same way.
Sample output
Real items from the example input above (one YouTube video transcribed from captions, one TikTok by speech recognition, both with an AI summary), verified on the Apify platform in September 2026. Long text and lists are shortened for this page; field names and values are otherwise exactly as the Actor returns them.
[{"platform": "youtube","videoId": "jNQXAC9IVRw","url": "https://www.youtube.com/watch?v=jNQXAC9IVRw","title": "Me at the zoo","author": "jawed","publishedAt": null,"durationSec": 19,"language": "en","transcriptSource": "captions","asrProvider": null,"segments": [{"startSec": 1.2,"endSec": 3.3600000000000003,"text": "All right, so here we are, in front of the elephants"},{"startSec": 5.318,"endSec": 7.974,"text": "the cool thing about these guys is that they have really..."},{"startSec": 7.974,"endSec": 15.732999999999999,"text": "really really long trunks and that's cool (baaaaaaaaaaahhh!!)"}],"text": "All right, so here we are, in front of the elephants\nthe cool thing about these guys is that they have really...\nreally really long trunks and that's cool (baaaaaaaaaaahhh!!)\nand that's pretty much all there is to say","wordCount": 39,"translation": null,"summary": {"oneLine": "A quick look at zoo elephants highlighting their long trunks.","bullets": ["The video opens with the presenter standing in front of the elephants.","It notes that the most notable feature is their very long trunks, describing them as cool."],"chapters": [{"startSec": 1,"title": "Introducing the elephants"},{"startSec": 5,"title": "Talking about their long trunks"}],"keywords": ["elephants", "zoo", "long trunks"],"hook": "All right, so here we are, in front of the elephants","cta": null,"sentiment": "positive"},"processingMs": 6563,"scrapedAt": "2026-09-13T02:07:14.393Z"},{"platform": "tiktok","videoId": "6718335390845095173","url": "https://www.tiktok.com/@scout2015/video/6718335390845095173","title": "Scramble up ur name & I’ll try to guess it😍❤️ #foryoupage #petsoftiktok #aesthetic","author": "Scout, Suki & Stella","publishedAt": null,"durationSec": null,"language": "en","transcriptSource": "asr","asrProvider": "groq","segments": [{"startSec": 0,"endSec": 10,"text": "I'm out."}],"text": "I'm out.","wordCount": 2,"translation": null,"summary": {"oneLine": "I'm out.","bullets": ["I'm out."],"chapters": [{"startSec": 0,"title": "Intro"}],"keywords": ["scramble name", "guess name", "TikTok challenge"],"hook": "I'm out.","cta": null,"sentiment": "neutral"},"processingMs": 30569,"scrapedAt": "2026-09-13T02:07:38.399Z"}]
Pricing
Pay-per-event. You pay for finished work, never for a video the Actor could not read.
| Event | Price | When |
|---|---|---|
| Video processed | $0.002 | Per video whose metadata was fetched. |
| Transcript from captions | $0.005 | Per video transcribed from the platform's own captions. |
| Transcript minute (speech recognition) | $0.005 | Per audio minute, rounded up, only after it succeeds. |
| Translation | $0.01 | Per video translated. |
| AI summary | $0.02 | Per video summarized. |
Worked examples:
- 100 videos served with captions: 100 × $0.002 + 100 × $0.005 = $0.70.
- One 10-minute TikTok without captions, plus a summary: $0.002 + 10 × $0.005 + $0.02 = $0.072.
- The example input above (one YouTube video from captions, one TikTok by speech recognition, two summaries): 2 × $0.002 + 2 × $0.005 + 2 × $0.02 = $0.054.
Apify's own apify-actor-start fee and platform compute are billed separately by your plan. Default
memory is 1024 MB, which is enough for every path including audio extraction.
Scheduling and integrations
- Fill in the input and Save as task.
- Add a Schedule to the task for a recurring pull.
- Under Integrations, send results to Slack, Google Sheets, a webhook, Make or Zapier on Run succeeded.
- From n8n or Make, call the Apify node with this Actor and read the dataset, or use Standby below for a synchronous single-video call.
Use as an API (Standby)
The Actor runs in Standby mode, so an AI agent or a workflow can get one transcript per HTTP request instead of starting a run and polling:
GET https://almoutasem-nabil--video-transcript-summary.apify.actor/?url=https://youtu.be/jNQXAC9IVRw&summarize=1Authorization: Bearer <your Apify token>
Query parameters mirror the input fields (url, language, translateTo, summarize,
summaryLanguage, includeTimestamps, forceAsr, maxDurationMinutes). The response is JSON with
items and errors, charged with the same events as a normal run. GET /health returns {"ok":true}.
Limitations, honestly
- Keep the default datacenter proxy. Without a proxy YouTube refuses requests from the Apify
platform. With it, YouTube captions are read directly; a video YouTube will not serve captions for
is transcribed from audio instead and billed per minute (
transcriptSource: "asr"). maxDurationMinutesis the cost control. Longer videos are skipped with a free error item instead of running up a per-minute bill.- AI summaries of very long videos are based on roughly the first 25 minutes of speech. Translations cover the whole transcript.
- Heavy dialect reduces accuracy. Naming the language (
"ar") instead ofautomeasurably helps; auto-captions on dialect-heavy content are often worse thanforceAsr. - Live streams are rejected — there is no finished transcript to return.
- Private, age-restricted or deleted videos produce an error item.
Legal note
The Actor reads publicly available videos and their published captions. It never logs in and sends no cookies. Audio is downloaded only as a temporary intermediate for transcription and deleted immediately; no media is stored or rehosted, and the output is text. Only public channel names and handles are recorded — no viewer names, no comments, no personal contact details. Use the output in line with each platform's terms and the copyright that applies to the source video. YouTube and TikTok are trademarks of their owners; this Actor is not affiliated with them.
Roadmap
Not available yet:
- Instagram Reels and X (Twitter) videos. Neither platform publishes captions and both block requests from Apify's IPs, so they cannot be offered reliably today. Links to them get a free "not available yet" error item.