Instagram Transcript Scraper — Reels, Posts & IGTV to Text
Pricing
from $15.00 / 1,000 transcript delivereds
Instagram Transcript Scraper — Reels, Posts & IGTV to Text
Instagram reel transcripts from a handle, reel URL, or chained dataset. On-device Whisper returns the full spoken text, hook3s, SRT/VTT subtitles and creator metadata. No login. Silent and music-only reels are not charged. Flat $0.015 per delivered transcript, any length.
Pricing
from $15.00 / 1,000 transcript delivereds
Rating
0.0
(0)
Developer
Muhamed Didovic
Maintained by CommunityActor stats
1
Bookmarked
162
Total users
81
Monthly active users
3 hours ago
Last modified
Categories
Share
Turn any public Instagram Reel into a swipe-file row: the spoken script, the hook (hook3s — first 3 seconds, always included), and optionally the on-screen text Whisper cannot hear (caption cards, prices, URLs burned into the frames). Paste a handle, reel URLs, or chain another Instagram scraper's dataset. No login. Flat $0.015 per delivered transcript — silent / music-only reels are not charged. Tick Read on-screen text, hook and structure to add onScreenText, hookOnScreen, ctaText and beats, or Summary, chapters and keywords to get the reel written up in its own language.
Why use this Instagram Transcript Scraper?
- No Instagram login or cookies required — works on any public reel/post/IGTV URL.
- Hook on every row —
hook3sis the first 3 spoken seconds, the line that stops the scroll. Always included, no extra charge. - On-screen text (optional) — caption cards, prices, step numbers and URLs burned into the frames. Whisper cannot hear these; tick the vision pass to get
onScreenText,hookOnScreenandctaText. - Summary, chapters and keywords (optional) — Haiku 4.5 writes them from the transcript, in the language the reel is spoken in, with chapter starts taken from the transcript's own timestamps.
- Real speech-to-text, not just captions — Instagram exposes no caption track to the public, so this actor downloads each video's audio and transcribes it with Whisper AI.
- Timestamped segments — every transcript comes with a
segmentsarray (start,end,text) you can turn into subtitles or jump-to-moment links. - 30+ languages — set the language or let Whisper auto-detect it per video.
- Rich metadata included — username, display name, caption, like/comment counts, thumbnail and upload date alongside the transcript.
- Pay per transcript — $0.015 per delivered transcript (a 2★ volume listing charges $0.026). Image-only, private, silent and music-only items are skipped or written as
no_speechand not billed. - Fresh Whisper, not cached captions — every run transcribes the audio. Failed items go to the
ERRORSrecord with a reason instead of a charged empty row. - Handle + reel URLs —
nasa,@nasa, or a reel/post/IGTV URL. No fake channel-caption scrape.
What does it do?
Give it a creator handle, reel URLs, or a dataset from another Instagram scraper. For each video the actor:
- Resolves the public video and its metadata (cookieless).
- Downloads the audio.
- Runs on-device Whisper speech-to-text.
- Pushes one row: metadata + full transcript + timestamped segments.
Supported input
| Input | Example |
|---|---|
| Creator handle / profile | nasa, @nasa, https://www.instagram.com/nasa/ |
| Reel URL | https://www.instagram.com/reel/Dabg9j1xNHM/ |
| Post URL (video) | https://www.instagram.com/p/ABC123xyz/ |
| IGTV URL | https://www.instagram.com/tv/ABC123xyz/ |
| Dataset chain | datasetId of a finished IG scraper run, or pasted datasetItems |
Direct .mp4 | Instagram CDN videoUrl from another scraper |
Image-only posts are detected and skipped (there is no audio to transcribe). Empty input transcribes one sample reel so "Start" still produces a row.
Use cases
- Content research & repurposing — turn a creator's reels into searchable text, blog drafts or scripts.
- Social listening & brand monitoring — analyze what's actually said in videos, not just captions.
- Accessibility — ready-made
.srtand.vttsubtitle files with every transcript. - Competitive & influencer analysis — mine talking points, hooks and CTAs across many videos.
- Dataset building — collect speech transcripts for NLP, trend detection or LLM fine-tuning.
How it works
The actor uses yt-dlp to resolve each public Instagram video and pull its audio without any login, then transcribes the audio locally with faster-whisper (a fast CTranslate2 build of OpenAI Whisper). Because transcription runs on-device, your data never goes to a third-party API. Instagram traffic is routed through Apify Residential proxies by default.
Input configuration
| Field | Type | Default | Description |
|---|---|---|---|
startUrls | array | — | Reel / post / IGTV URLs, or a profile URL. |
handles | array | — | Creator usernames (nasa, @nasa). Lists recent reels, newest first. |
maxReelsPerHandle | integer | 12 | How many of each creator's latest reels to transcribe. |
datasetId | string | — | Default dataset of a finished IG scraper run; reel URLs are picked out of each row. |
datasetItems | array | — | Paste scraper rows instead of chaining by ID. |
watchlistId | string | — | Named store that remembers transcribed shortcodes across scheduled runs. |
newReelsOnly | boolean | false | Skip reels already on that watchlist. Needs watchlistId. |
videoUrls | array | — | Direct .mp4 / .webm links (CDN videoUrl from another scraper). |
whisperModel | select | tiny | tiny (fast, reliable), base or small (more accurate, slower). |
language | string | auto | ISO code (en, es, ar, …) or auto to detect per video. |
includeSegments | boolean | true | Include the segments array of timestamped chunks. |
maxItems | integer | 25 | Maximum videos to transcribe this run. |
maxDurationSec | integer | 300 | Skip videos longer than this (compute guard). 0 disables. |
maxConcurrency | integer | 2 | Videos transcribed in parallel (1–4). |
summarize | boolean | false | Haiku 4.5 writes a summary, chapters with start times and keywords, in the reel's own language. Charged per summary delivered. |
proxy | object | Residential | Proxy configuration. |
Example input
{"handles": ["nasa"],"maxReelsPerHandle": 12,"watchlistId": "nasa-daily","newReelsOnly": true,"startUrls": [{ "url": "https://www.instagram.com/reel/Dabg9j1xNHM/" }],"whisperModel": "tiny","language": "auto","includeSegments": true,"maxItems": 25}
What was not delivered, and why
Every reel the input asked for ends as a charged dataset row or as an entry in the run's ERRORS record (key-value store), never neither. A reel that was not delivered also gets an uncharged row in the dataset with its status set to the reason and an error saying what to change, so an integration that reads only the dataset (the API, Make, n8n) sees it too. The record is written on every run, with an empty items list when everything was delivered, so an integration can read it unconditionally. Each entry carries a reason, a plain-language detail saying what happened and what to change, and a retryable flag:
reason | What happened | Re-run as-is? |
|---|---|---|
private_or_gone | Instagram would not serve the reel to a logged-out client: private, deleted, region-locked | No |
no_audio | An image post; nothing to transcribe | No |
no_speech | Music or silence, including a transcript made only of the filler speech models invent over music ("Thank you.", "♪"). The row is in the dataset with status: no_speech, not charged | No |
too_long | Over your maxDurationSec (a value from 1 to 9 would skip every reel, so the run uses the default 300 instead and says so) | Raise the limit, or set 0 to turn it off |
deadline | The run would have hit its timeout before this reel finished | Raise the timeout or lower the batch |
transport | The fetch failed on the network after retries on fresh proxy exits, or Instagram throttled every exit tried | Yes, in a few minutes |
transcription_failed | Whisper or the download failed | Yes |
The run's status line carries the same arithmetic, for example "Delivered 9 of 10 reels. 1 not delivered and not charged: 1 private, removed or not found — see the ERRORS record." A run that delivers nothing because every fetch failed on the network is marked FAILED so it stands out in your run list; a run whose misses describe the reels themselves finishes SUCCEEDED. Reels left out by the watchlist are counted separately on the status line and never charged.
Output
One row per video. Example (abridged):
{"id": "Dabg9j1xNHM","shortcode": "Dabg9j1xNHM","url": "https://www.instagram.com/reel/Dabg9j1xNHM/","username": "nasaadmin","userId": "79712351002","fullName": "NASA Administrator Jared Isaacman","caption": "For 250 years, America has inspired generations to dream bigger…","durationSec": 58,"thumbnailUrl": "https://instagram.f…/735552979_….jpg","timestamp": 1783294802,"uploadDate": "20260705","likeCount": 56053,"commentCount": 1923,"viewCount": null,"isVideo": true,"transcript": "For 250 years, the United States of America has been the hope, the promise, the light and the glory of all the nations of the world…","hook3s": "For 250 years, the United States of America has been the hope","hookStartSeconds": 0.46,"status": "ok","language": "en","segments": [{ "start": 0.46, "end": 5.46, "text": "For 250 years, the United States of America has been the hope" }],"srt": "1\n00:00:00,460 --> 00:00:05,460\nFor 250 years, the United States of America has been the hope\n","vtt": "WEBVTT\n\n00:00:00.460 --> 00:00:05.460\nFor 250 years, the United States of America has been the hope\n","srtFileUrl": "https://api.apify.com/v2/key-value-stores/…/records/Dabg9j1xNHM.srt","vttFileUrl": "https://api.apify.com/v2/key-value-stores/…/records/Dabg9j1xNHM.vtt","transcriptSource": "whisper","whisperModel": "tiny","transcribedAt": "2026-07-15T09:49:00.990Z"}
Key output fields
| Field | Description |
|---|---|
transcript | Full spoken text of the video. |
hook3s | Spoken text that starts in the first 3 seconds (always present). |
hookStartSeconds | Offset of that hook, or null if there is no speech. |
status | ok for a transcript. Anything else is an uncharged row: no_speech, or the reason a reel was not delivered (too_long, private_or_gone, no_audio, deadline, transport, transcription_failed), with error saying what to change. Only ok rows are charged. |
segments | Array of {start, end, text} timestamped chunks (subtitle-ready). |
srt / vtt | The transcript as ready-made SubRip and WebVTT subtitle text. |
srtFileUrl / vttFileUrl | Direct download links to the .srt / .vtt files. |
language | Detected (or specified) spoken language. |
username / fullName | Creator handle and display name. |
caption | The post's written caption. |
likeCount / commentCount | Engagement metrics when available. |
durationSec | Video length in seconds. |
transcribedAt | UTC timestamp of transcription. |
summary | Two to four sentences on what the reel covers, in the reel's language. Only with summarize. |
summaryChapters | {start, end, startTime, title} per chapter, timed from the transcript itself. |
summaryKeywords | The names, topics and terms that matter most, most important first. |
summaryError | Why a summary is missing on a row that asked for one. The transcript is unaffected. |
FAQ
Coming from another Instagram transcript actor? If a listing charged you and returned cached or empty text, or failed on a handle / channel scrape, this actor transcribes the audio with on-device Whisper on every run. Price is $0.015 per delivered transcript (vs $0.026 on the 2★ volume listing). Silent and music-only reels are written with status: no_speech and not charged. A handle (nasa, @nasa) lists recent reels; there is no channel-scrape mode that invents captions. Failed items land in the ERRORS key-value record with a reason — they are not billed.
What does it cost? A flat $0.015 per delivered transcript plus a $0.005 run start — a 10-minute IGTV costs the same as a 15-second Reel. Silent / music-only reels (status: no_speech) and skipped items are never charged.
Do I need an Instagram account or cookies? No. The actor works on public videos without any login.
Can I get subtitle files? Yes. Every row carries the full srt and vtt text plus srtFileUrl / vttFileUrl download links — drop them straight into CapCut, Premiere, DaVinci Resolve or Instagram's own caption upload.
Why not just read Instagram's captions? Instagram does not expose a caption/subtitle track to logged-out clients. This actor transcribes the actual audio, so you get text even for videos the creator never captioned.
What about image posts or carousels with no video? They're detected and skipped (reported as skipped, not charged) — there is no audio to transcribe.
Which languages are supported? 30+ via Whisper. Set language for best speed/accuracy, or use auto.
Why is tiny the default model? It is ~5× faster than larger models and reliable on Apify's CPU. Use base/small for higher accuracy on clear speech; they are slower and can time out on long videos.
Support
- Found a bug or need a new field? Open a ticket on the Issues tab of this actor — it's the fastest way to reach me and I actively maintain this scraper.
- Email: muhamed.didovic@gmail.com
- Website: muhamed-didovic.github.io
Additional Services
Need something beyond the standard output? I build and maintain custom scrapers and data pipelines. Happy to help with:
- Transcripts from other platforms (TikTok, YouTube, Douyin) in the same output shape
- Scheduled, incremental transcript feeds wired into your CMS, data warehouse, or vector store
- Custom fields, summaries, translations, or export formats tailored to your workflow
- Private or dedicated actors for high-volume or compliance-sensitive use
Reach out via the Issues tab or email and describe what you need.
Explore More Scrapers
More of my actors in the same space:
- Social Video Transcript Scraper — transcribe TikTok and Instagram videos in one run, same Whisper pipeline
- Douyin Scraper — Chinese TikTok profiles, videos and search
- YouTube Shorts Scraper — Shorts metadata at scale
- Instagram Profile Scraper — profiles, followers, bio and contact data
Browse the full catalog — job boards, business directories, review sites, social platforms — on my profile: memo23 on Apify.
🤖 For AI Agents & LLM Apps
Compact reference for AI agents calling this actor via the Apify MCP server or the Apify API (actor: memo23/instagram-transcript-scraper).
Purpose: Transcribe public Instagram Reels, video posts and IGTV to text with on-device Whisper AI, returning one row per video with metadata, hook3s, a full transcript and timestamped segments.
Minimal input:
{"handles": ["nasa"],"maxReelsPerHandle": 5,"maxItems": 5}
Output: one row per video — id, shortcode, url, username, userId, fullName, caption, durationSec, thumbnailUrl, timestamp, uploadDate, likeCount, commentCount, viewCount, isVideo, transcript, hook3s, hookStartSeconds, status, language, segments [{ start, end, text }], srt, vtt, srtFileUrl, vttFileUrl, transcriptSource, whisperModel, transcribedAt. With summarize: summary, summaryChapters [{ start, end, startTime, title }], summaryKeywords, summaryError.
Behaviors an agent should know:
- Accepts
handles/ profile URLs, reel/p//tv/URLs,datasetId/datasetItems, and direct.mp4videoUrls. Empty input transcribes one sample reel (charged if speech is delivered) unlesswatchlistId/newReelsOnlyemptied the list. watchlistId+newReelsOnlypersist seen shortcodes in a named KV store so a scheduled run only transcribes new reels.newReelsOnlywithoutwatchlistIdis ignored.- Image-only posts are skipped (not charged).
status: no_speechrows are written and not charged. - Always set
maxItems; transcription is the slow, billed step, so start small. whisperModelistiny(default, fastest) /base/small;base/smallare more accurate but slower and can time out on long videos.languageaccepts an ISO code (en,es,ar) orauto; setting it is faster and more accurate than auto-detect.maxDurationSec(default 300) skips videos longer than the limit;includeSegmentstoggles the timestampedsegmentsarray.summarize: trueaddssummary,summaryChaptersandsummaryKeywords, written by Haiku 4.5 from the transcript in the reel's own language. Chapter starts come from the transcript's timestamps, so they always land inside the recording. A reel with no speech is not summarized; a summary that fails leaves the transcript intact and puts the reason insummaryError.- Flat $0.015 per delivered transcript plus a $0.005 run start; skipped and
no_speechitems are not charged.summarizebills its own event per summary delivered, and only when one is delivered.
⚠️ Disclaimer
This actor collects only publicly available data from Instagram and is intended for lawful uses such as research, accessibility and content analysis. You are responsible for how you use the output, including compliance with Instagram's Terms of Service, applicable copyright, and data-protection laws (e.g. GDPR/CCPA) where relevant. Do not use transcripts to infringe copyright or process personal data unlawfully.
SEO Keywords
Instagram transcript scraper, Instagram reel transcript, Instagram video to text, transcribe Instagram reels, Instagram speech to text, Instagram captions scraper, IGTV transcript, Instagram subtitles generator, Whisper Instagram, Instagram audio transcription, Instagram reels text extractor, social media transcription, Instagram AI transcript extractor, reel to text converter, Instagram video transcription API, extract text from Instagram videos, Instagram content analysis, Instagram VTT subtitles, bulk Instagram transcription, Instagram NLP dataset, influencer content research, no-code Instagram scraper.