YouTube Shorts Scraper - Growth & Audience Insights
Pricing
from $5.00 / 1,000 shorts
YouTube Shorts Scraper - Growth & Audience Insights
Scrape YouTube Shorts from channels and keyword searches. Get video metrics, available transcripts, comments, replies and optional author avatars. Track growth across runs and discover audience questions, product requests and content ideas backed by real comments.
Pricing
from $5.00 / 1,000 shorts
Rating
0.0
(0)
Developer
Kelopr_bk
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
▶️ YouTube Shorts Scraper
Find Shorts. Understand their growth. Hear what viewers are asking for.
📊 Video metrics · 📝 Transcripts · 💬 Comments · 🎯 Audience requests
Collect public Shorts from channels or keyword searches. Export video metrics, creator details, available transcripts and optional comment samples. Track the same Shorts across runs to measure growth, and turn explicit viewer questions into an evidence-backed research list.
New here? Keep the starter example, click Start, then open Overview. It collects a small sample of real Shorts. Increase the result limit when you are ready.
⚡ Choose what you need
| Your task | Settings | What you receive |
|---|---|---|
| Collect video statistics quickly | Fast | Titles, URLs, views, likes, total comment counts, dates, duration and channel details |
| Read what is said in each Short | Full | Everything in Fast, plus available public caption text |
| Read audience reactions | Turn on Collect comment texts | A separate Comments table with text, authors, likes and source links |
| Follow conversations | Also turn on Include replies | Replies linked to parent comments, within your limits |
| Keep profile pictures | Turn on Save commenter avatars | Saved image files, once per author per run |
| Find questions and requests | Keep Find audience questions and requests on | Matching comments in Questions, plus grouped evidence in Audience requests |
| Track growth over time | Reuse a History workspace | Measured changes between runs and comparisons at similar video ages |
Fast / Full controls transcripts. Comments, replies and avatars are separate optional switches. Adding them increases the work even if Fast is selected.
🚀 Start in three steps
- Choose your sources. Enter channel handles/URLs, search phrases, or both. One phrase per line.
- Choose the depth and limits. Select Fast or Full, set Shorts per source, and enable comments or avatars if needed.
- Run and export. Open the result view you need, then download JSON, CSV or Excel.
maxResultsShorts applies per channel or search phrase. maxTotalResults is an additional whole-run cap. Duplicates and records excluded by the publication-date filter do not consume the saved-result allowance.
🧩 Copy-ready examples
1. Statistics for a channel
{"channels": ["MrBeast"],"mode": "fast","maxResultsShorts": 50}
2. Search a niche and collect transcripts
{"searchQueries": ["space facts"],"mode": "full","maxResultsShorts": 30,"transcriptLanguage": "auto"}
3. Research audience questions
{"searchQueries": ["camera review"],"mode": "fast","maxResultsShorts": 10,"includeComments": true,"maxCommentsPerShort": 50,"maxCommentsTotal": 500,"commentsSort": "newest","analyzeAudience": true,"downloadAvatars": false}
4. Transcripts, conversations and saved avatars
{"channels": ["MrBeast"],"mode": "full","maxResultsShorts": 10,"includeComments": true,"maxCommentsPerShort": 20,"maxCommentsTotal": 200,"includeReplies": true,"maxRepliesPerComment": 3,"downloadAvatars": true,"maxAvatarDownloads": 200}
5. Measure growth on later runs
{"channels": ["MrBeast"],"mode": "fast","maxResultsShorts": 50,"trackHistory": true,"historyKey": "my-creator-watchlist","minObservationIntervalSeconds": 900,"comparisonAgeHours": 24}
Run the same input again later. The Actor does not schedule itself or wait for another measurement. Use Apify scheduling separately if you want recurring runs.
📂 Where to find the results
| View | Contents |
|---|---|
| Overview | One row per saved Short, with its main statistics |
| Analytics | Engagement rates, duration and lifetime average views/day |
| Growth | Observed view changes, growth rate and stage |
| Comparisons | Age-matched channel baselines and outlier ratios |
| Transcripts | Available caption text, language, status and opening text |
| Comments | All saved comment texts and optional author avatars |
| Questions | Only matching viewer comments; creator-authored questions are excluded |
| Audience requests | Grouped request types, sample sizes and evidence |
| Shared topics | Shared phrases/hashtags across different creator channels |
| Compact | Alternative core field names for compact exports |
Comments are stored in a separate dataset and linked to Shorts by videoId. Questions are a filtered research copy of matching comments. The OUTPUT record explains limits, coverage and failures. AUDIENCE and TOPICS contain aggregate results.
💬 Comment counts, texts and replies
Total comments and collected comments are different. commentsCount is the numeric counter shown by YouTube. commentsCollected is how many comment records this run actually saved for that Short. A video can have 12,481 comments while your sample contains 20.
- The total counter is collected even when comment texts are switched off.
maxCommentsPerShortincludes replies when replies are enabled.maxCommentsTotallimits the combined comment output across the run.commentsSortsupports Popular (top) and Newest (newest).- Each comment ID is saved once per run. Repeated authors remain separate comments.
- A checked reply sample can show an observed creator response. It cannot prove that no response exists outside that sample.
Comment rows include text, author name/handle/channel link, likes and their precision, reply count, pin/creator-heart flags when available, and a source link. YouTube often supplies a relative date such as “2 days ago”. We preserve it in publishedTimeText; an exact publishedAt is not invented.
🖼️ Optional avatars
Avatar downloads are off by default. Off means no commenter-image files are fetched or exposed for image previews. On saves one image per author per run and reuses it across that author's comments. authorAvatarUrl links to the saved file; avatarStatus explains failures, missing images or a reached limit.
The code supports a separate avatar_saved event only after successful file storage, and comment_saved only after successful comment storage. Events are used only when already configured on the platform. Monetization is not configured by the Actor itself.
🎯 What audience analysis tells you
The Actor finds explicit phrases about purchase links, prices, product models/features, tutorials, follow-ups, alternatives and reported problems. Each signal keeps the original comment and matched phrase. Aggregates count distinct author channel IDs and expose the number of sampled comments.
Phrase rules cover English, Russian, Spanish and Portuguese. This is transparent text matching, not a prediction that someone will buy, and not a general AI sentiment model. Unmatched comments remain in Comments. Generic praise does not become a buying signal.
📝 How transcripts work
Full mode attempts to retrieve public YouTube captions. auto selects an available track, preferring manual captions and then English. A specified language selects that language when available; the Actor does not silently substitute another language or generate a translation.
| Field | Meaning |
|---|---|
transcript | Caption text returned by YouTube |
transcriptLanguage | Language of the selected track |
transcriptIsAutoGenerated | Whether YouTube marks the track as automatic captions |
openingText | Caption segments starting within the first five seconds |
transcriptStatus | ok, no_captions, language_unavailable, empty, unavailable, timeout, error or not_requested |
Captions can contain recognition errors or sound labels. Word counts use whitespace-separated tokens and are not linguistic segmentation for every language. Missing captions never become text generated from the title or description.
📈 Growth and fairer comparisons
The first observation has no invented history. Two sufficiently spaced observations allow viewsDelta and observedViewsPerHour. Three allow comparison of recent growth rates. Four can support the defined second-wave heuristic. The default minimum interval is 15 minutes. Decreases in YouTube counters are labeled as corrections.
For age-matched comparisons, the default target is 24 hours after publication, with an explicit ±25% age window. The baseline uses at least five other Shorts from the same channel, one actual observation per Short. No observation near the requested age means a null ratio and a clear status. Starting to track an old Short cannot reconstruct its first-day views.
Shared-topic groups use the search phrase, hashtags or adjacent words from titles. They describe the collected sample, not all of YouTube, and do not infer the visual content of a video.
⚙️ Important settings and limits
| Setting | Default / meaning |
|---|---|
scrapeType | auto: use supplied channels and searches; optionally restrict to channels or search |
outputFormat | standard or compact; independent of Fast / Full |
sortChannelShortsBy | NEWEST, POPULAR or OLDEST for channel inputs |
oldestPostDate | Inclusive publication cutoff: ISO date, Unix seconds/milliseconds or a window such as 30 days |
trackHistory, historyKey | Remember observations and keep separate monitoring lists |
maxRunSeconds | 60 seconds by default; maximum 120 |
| Source inputs | Up to 20 channels/search phrases combined |
| Comments | Up to 500 per Short and 10,000 across a run |
| Avatars | Up to 1,000 distinct authors per run |
Separate resource protections bound source requests, response sizes, transfer and estimated cost, and stop prolonged collection without matching results. These protections cannot be increased through input fields. A result cap is a maximum, not a promise to fetch that many regardless of time, source availability or resource limits. OUTPUT.resourceBudget explains any resource stop.
❓ Common questions
Why are some fields null? YouTube may hide or omit them. Disabled comments, hidden likes and unavailable captions have explicit statuses. Zero and false values are preserved.
What is dataStatus: complete? Core video metadata is present, and likes/comments are available or explicitly hidden/disabled. Transcript, comment-text and avatar coverage are reported separately.
Why did the date filter return nothing? The discovered Shorts may all be older than the cutoff. The run reports the cutoff, counts and a correction suggestion. It never relaxes your filter silently.
Why did a run stop early? A source, saved-record, page, time, spending or owner resource limit may have been reached. Saved records remain available and the stop reason is reported.
Are all schema fields implemented? The preserved fields for translations, subtitles, AI summaries, collaborators and monetization can remain null. New transcripts are in transcript. Channel age restrictions are not inferred from a family-safety flag.
How is speed measured? Record the exact preset. Fifty metadata records and thirty Shorts with hundreds of comments and images are different workloads. Source response times and caption availability vary; see the validation report for measured runs.
Does it change prices or publish itself? No. This is currently a private development Actor with no configured monetization.