YouTube Comments Scraper — Top, Newest & Replies
Pricing
from $0.17 / 1,000 additional comment or replies
YouTube Comments Scraper — Top, Newest & Replies
Export public YouTube comments and replies from 1–10 videos. Compare videos with an equal per-video limit and a per-video coverage CSV; skip already-seen comment IDs on repeat runs. $0.0005 first delivered comment, then $0.000165 each; zero delivered rows = no Actor event charge.
Pricing
from $0.17 / 1,000 additional comment or replies
Rating
0.0
(0)
Developer
Datamule
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
2
Monthly active users
4 days ago
Last modified
Categories
Share
YouTube Comments Scraper — Comments & Replies
Export public YouTube comments and reply threads to JSON, CSV or Excel. No API key or login. Two common jobs: pull comments from one video, or compare audience feedback across several videos with an equal, bounded sample from each.
Compare audience feedback across videos
What you get: one Comments dataset (every row tagged with videoId) plus a Per-video coverage (CSV) in the Output tab that tells you how many comments each video actually contributed, so you know which videos are comparable.
- Paste 2–10 video URLs into
videoUrls. - Set
maxCommentsPerVideo(the equal share per video) andmaxCommentsto at least videos × share. - Run, then open Output → "Per-video coverage (CSV)", and filter the dataset by
videoId.
Ready-made starter: open the public Compare YouTube Audience Comments example and click Try for free. It saves the input below as a task in your Apify account (sign-up is free if you don't have one); replace the URLs with your own videos. Or paste the JSON below into the Input tab.
Tested input (an Apify cloud run of this Actor returned 60 rows, 20 per video, instead of 60 from the first video):
{"videoUrls":["https://www.youtube.com/watch?v=dQw4w9WgXcQ","https://www.youtube.com/watch?v=9bZkp7q19f0","https://www.youtube.com/watch?v=kJQP7kiw5Fk"],"maxComments":60,"maxCommentsPerVideo":20,"sort":"top","includeReplies":false}
Coverage CSV for that run, produced by the current coverage code from its stored rows (URL, cap, note and skippedAlreadySeen columns omitted):
| requestOrder | videoId | status | stopReason | deliveredTotal | topLevelDelivered | repliesDelivered |
|---|---|---|---|---|---|---|
| 1 | dQw4w9WgXcQ | available | per_video_cap | 20 | 20 | 0 |
| 2 | 9bZkp7q19f0 | available | per_video_cap | 20 | 20 | 0 |
| 3 | kJQP7kiw5Fk | available | result_cap | 20 | 20 | 0 |
How to read it: per_video_cap or source_exhausted means that video got its share (or everything the source returned). result_cap on the last video means the global maxComments was reached as it hit 20, so it also got its full share here. not_attempted, charge_limit or an unavailable status means that video is under-covered: leave it out or re-run it. Counts are rows stored in this run, not the video's total comments, and top-level and reply counts are shown separately.
First top-level row per video from the same run (real retained output, selected columns):
| videoId | text | likeCountText | replyCountText |
|---|---|---|---|
| dQw4w9WgXcQ | can confirm: he never gave us up | 319K | 962 |
| 9bZkp7q19f0 | This guy could've said 'Vote me for president' and everyone would've voted | 32K | 195 |
| kJQP7kiw5Fk | who's here after 9 billion?? | 85K | 973 |
This is a bounded sample in YouTube's own top order, not a representative or complete sample, and the Actor adds no sentiment scores. Do the analysis in your own tools.
Single video
{"videoUrls":["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"maxComments":100,"sort":"top","includeReplies":false}
Use sort: "newest" for recent feedback, or includeReplies: true for bounded thread context.
Price and key limits
Charge per run = $0.0005 for the first delivered comment + $0.000165 for each additional comment or reply ($0.165 per 1,000). The 60-row comparison above costs $0.0005 + 59 × $0.000165 = $0.010235. No start fee; zero delivered rows = no Actor event charge; partial results already delivered stay billable.
| Delivered comments and replies | Total Actor charge (USD) |
|---|---|
| 1 | $0.000500 |
| 100 | $0.016835 |
| 1,000 | $0.165335 |
- Up to 10 videos per run;
maxCommentsis a global cap (1–10,000, default 100) across all videos and replies. - Set a positive Apify maximum charge to cap spend. Do not enter zero: Apify treats zero as unset. A positive limit below $0.0005 delivers no rows.
- Uses YouTube's anonymous web comments surface: no completeness guarantee, no proxy rotation, no livestream chat, no private/age-gated/members-only videos.
Input and coverage
videoUrls: 1–10 HTTPS YouTube video URLs or 11-character IDs. You can paste a list copied from a spreadsheet or chat into one row: links separated by commas, semicolons, spaces or new lines are split into separate videos, and wrapping quotes/brackets, a trailing period and invisible characters are removed. Still at most 10 distinct videos per run. Duplicate videos are deduplicated; an unrecognized piece is skipped and named in the log andSUMMARY, while the valid ones still run. Shorts/live URLs are accepted, but livestream chat is not supported.maxComments: integer 1–10,000, default 100; a global hard cap across all videos and replies, not per video. Inputs are processed sequentially; without a per-video cap the first video may consume the global cap. No unlimited mode.maxCommentsPerVideo: optional integer 1–10,000, empty by default (no per-video limit). Caps comments and replies together for each video; when reached, the Actor stops fetching that video and moves to the next.maxCommentsand your Apify maximum charge still take precedence.sort:top(default) ornewest. The source may put a pinned comment first even in newest mode. There is no chronological sorting of relative dates by this Actor.includeReplies: false by default. When true, each page of top-level comments is emitted before following its reply threads in source order; a popular first thread may consume the remaining cap (or the video's per-video cap). Comments and replies count equally toward the cap and price.excludeCommentIds: optional list of up to 5,000 exactcommentIdvalues (top-level or reply IDs) from an earlier run. Checked before any YouTube request. See "Repeat monitoring" below.
Repeat monitoring: skip comments you already have
First run (keep the dataset):
{"videoUrls":["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"maxComments":100,"sort":"newest"}
Open the run's Output tab and download "Already-seen comment IDs" (key-value store record SEEN_COMMENT_IDS, a JSON array). Paste it as-is into excludeCommentIds on the next run. Via API: GET https://api.apify.com/v2/key-value-stores/<defaultKeyValueStoreId>/records/SEEN_COMMENT_IDS. Repeat each time: the new list already contains the IDs you passed in, so it keeps growing run to run (up to the 5,000-ID limit):
{"videoUrls":["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"maxComments":100,"sort":"newest","excludeCommentIds":["Ugzge340dBgB75hWBm54AaABAg","UgwxEqSh_78DU_MOKpt4AaABAg","UgxPYbA7XtJzYXPPPHV4AaABAg"]}
Comments whose commentId matches exactly are dropped before storage: they are not in the dataset, not charged, and do not count toward maxComments or maxCommentsPerVideo. Only new stored rows are billed, at the normal price. Matching is exact source ID only; edited or reposted text with a new ID is treated as new. The list lives only in your run input; nothing is shared between runs or users.
How the list is built: IDs of rows actually stored in this run's dataset come first (dataset order; top-level and reply IDs alike, including rows restored after an Apify restart), then the excludeCommentIds you passed in (input order), exact duplicates removed. Rows cut off by your charge limit were never stored, so they are not listed. It is capped at 5,000 IDs, the input limit: past that, the oldest carried IDs are dropped, SEEN_COMMENT_IDS_INFO.droppedOverLimit says how many and complete becomes false. Rows without a valid ID are counted in unkeyedRowsNotExported, never listed. The record holds IDs only (no comment text or author data), costs nothing and makes no extra YouTube requests. It is not the video's full comment history, and an edited comment keeps its ID, so it stays skipped.
Limits: the Actor does not crawl forever looking for new rows. The usual request/time limits still apply, and a video stops after 5 consecutive source pages with no new rows (stopReason: "already_seen_window"). That means "stopped looking", not "no new comments exist", and the Actor makes no claim to have seen the full or newest history. SUMMARY.skippedAlreadySeen and the skippedAlreadySeen coverage column count distinct excluded IDs actually observed in this run (blank when no exclusions were given); after an Apify restart/migration the count covers only the resumed attempt, while rows already stored are still never re-delivered or re-charged. If every fetched comment was already seen, the run succeeds with zero new rows and zero charge, which differs from an unavailable video (reported as skipped/failed as usual). sort: "newest" is the best fit for monitoring; with top the order shifts and new comments may sit deeper than the window.
Run it from your own code (Python, standard library only)
This script does the copy-and-paste step for you. It starts a bounded run, waits up to 5 minutes, saves the new comments to comments_<runId>.json, and keeps the returned SEEN_COMMENT_IDS in a local seen_comments.json that it sends as excludeCommentIds next time. Set APIFY_TOKEN in your environment first; the script sends it as a Bearer header and never prints it. Run it again whenever you want the next batch; no scheduler is included.
import hashlib, json, os, time, urllib.requestTOKEN = os.environ["APIFY_TOKEN"]API = "https://api.apify.com/v2"STATE = "seen_comments.json" # your local file; keep it between runssettings = {"videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],"sort": "newest", "maxComments": 20, "includeReplies": False}def call(method, path, body=None):req = urllib.request.Request(API + path, method=method,data=None if body is None else json.dumps(body).encode(),headers={"Authorization": f"Bearer {TOKEN}", "Content-Type": "application/json"})with urllib.request.urlopen(req, timeout=90) as r:return json.load(r)def main():key = hashlib.sha256(json.dumps(settings, sort_keys=True).encode()).hexdigest()[:16]state = json.load(open(STATE)) if os.path.exists(STATE) else {}seen = state.get(key, {}).get("ids", [])run_input = dict(settings, excludeCommentIds=seen) if seen else settings# maxComments caps rows; maxTotalChargeUsd separately caps what this run can cost you.run = call("POST", "/acts/datamule~youtube-comments-scraper/runs?maxTotalChargeUsd=0.01", run_input)["data"]deadline = time.time() + 300 # stop waiting after 5 minuteswhile run["status"] in ("READY", "RUNNING") and time.time() < deadline:run = call("GET", f"/actor-runs/{run['id']}?waitForFinish=60")["data"]if run["status"] != "SUCCEEDED":raise SystemExit(f"run {run['id']} ended {run['status']}; {STATE} left unchanged")rows = call("GET", f"/datasets/{run['defaultDatasetId']}/items?clean=true")with open(f"comments_{run['id']}.json", "w") as f:json.dump(rows, f, ensure_ascii=False, indent=1)kvs = f"/key-value-stores/{run['defaultKeyValueStoreId']}/records/"try:ids, info = call("GET", kvs + "SEEN_COMMENT_IDS"), call("GET", kvs + "SEEN_COMMENT_IDS_INFO")summary = call("GET", kvs + "SUMMARY")except OSError as e:raise SystemExit(f"run {run['id']}: could not read its records ({e}); {STATE} left unchanged")ok = (isinstance(ids, list) and all(isinstance(i, str) and i for i in ids)and isinstance(info, dict) and info.get("ids") == len(ids) <= 5000and {r["commentId"] for r in rows if r.get("commentId")} <= set(ids))if not ok:raise SystemExit(f"run {run['id']}: seen-ID records missing or invalid; {STATE} left unchanged")state[key] = {"settings": settings, "ids": ids, "lastRun": run["id"]}with open(STATE + ".tmp", "w") as f:json.dump(state, f)os.replace(STATE + ".tmp", STATE) # atomic: the old file survives a crashstops = sorted({v.get("stopReason") for v in summary.get("videos", []) if v.get("stopReason")})print(f"run {run['id']}: {len(rows)} new rows -> comments_{run['id']}.json; "f"{len(ids)} IDs saved; stop reasons {stops}")if info.get("droppedOverLimit"):print(f"5,000-ID limit: {info['droppedOverLimit']} oldest IDs dropped; they may come back")main()
How it behaves:
- The first run for a settings block sends no exclusions and saves the IDs it got. Later runs send only the IDs saved for the same
settings; changing a video,sort,maxCommentsorincludeRepliesstarts a separate list in the same file. seen_comments.jsonis replaced (atomically) only after the run SUCCEEDED and bothSEEN_COMMENT_IDSandSEEN_COMMENT_IDS_INFOlook valid. A failed, aborted or timed-out run, a network error or a missing record exits and leaves the file as it was (the comments file may already be written in that case). A run still RUNNING at the deadline keeps going in the cloud, bound by the charge cap; if it later succeeds, run the script again and those comments are fetched afresh, because its IDs were never saved.- Deleting or losing
seen_comments.jsonmeans the next run starts over and re-delivers (and re-bills) comments you already have. 0 new rowswith a SUCCEEDED run is a normal result: no unseen comment turned up within the window this run looked at. Withalready_seen_windowin the stop reasons, the run stopped after 5 pages of known comments; that is not proof that no new comments exist.- Past 5,000 IDs the oldest ones are dropped (the script prints how many), so very old comments may be delivered again.
maxComments: 20andmaxTotalChargeUsd=0.01keep each run small; raise them if a video gets more than 20 comments between your runs, or you will miss some. The output is the comments as YouTube's page returns them. No sentiment, topic or sampling claim is made.
Output and provenance
Every row retains the source commentId, requested videoId, explicit parentCommentId for replies, and commentUrl. publishedTimeText is the exact relative source string; publishedAt is null because this source does not provide an absolute posting date. scrapedAt is the actual UTC observation time, not posting time. likeCountText and replyCountText retain source display values; numeric fields are null for rounded values such as 312K.
Limits and diagnostics
Uses the anonymous YouTube WEB comments surface. It is undocumented and may change, block cloud IPs, omit comments, or personalize ranking. No completeness or continuous-availability guarantee is made. No retries, proxy rotation or access-control workarounds. Age-gated, private, unavailable, restricted and members-only sources are out of scope.
SUMMARY in the default key-value store describes videos and stop reasons. With maxCommentsPerVideo set, each video also reports deliveredCount, and stopReason: "per_video_cap" when its cap was reached. Explicit comments-disabled responses produce no dataset row and no event. If every requested video is unavailable/restricted/unrecognized the run fails rather than looking like an empty success. In a mixed batch, valid videos are still exported and the run succeeds; skipped video IDs are listed in SUMMARY.skippedVideos and the run log. Partial rows may remain after a failure; inspect run status and SUMMARY.
Finite safety bounds: 150 HTTP requests, 8 MB decoded per response, 80 MB decoded total and 240 seconds of extraction per process lifetime; platform timeout 300 seconds. A safety-limit failure is not source exhaustion. Failed storage/charging calls propagate; they are never swallowed. Diagnostic records are not billed comments.
Per-video coverage CSV
Every run also writes COVERAGE, a CSV in the default key-value store, linked in the run's Output tab as "Per-video coverage (CSV)". It has one row per distinct requested video, in input order: requestOrder, videoId, videoUrl, status, stopReason, deliveredTotal, topLevelDelivered, repliesDelivered, maxCommentsPerVideo, maxComments, note, skippedAlreadySeen. Counts come from the rows actually stored in this run's dataset (including rows restored on resume), so they add up to the dataset. They are not the video's total comment count, and they say nothing about sentiment or how representative the sample is. Blank means not set or unknown; it never stands in for zero. Status values: available, comments_disabled, unavailable_or_restricted, empty_or_unrecognized, restored_complete (already at its cap on resume) and not_attempted (the global cap or charge limit was reached first, or the run failed before that video). Stop reasons include per_video_cap, result_cap, charge_limit, source_exhausted and already_seen_window. Comments and replies draw on one pool, limited by the per-video cap, the global maxComments and your charge limit. The CSV is not billed and adds no dataset rows. Apart from videoId and videoUrl, which stay unchanged so they still match the dataset, cells that start with =, +, - or @ get a leading apostrophe so spreadsheets don't run them as formulas. The dataset text itself is never changed.
On restart, physical dataset rows restore the global cap, per-video counts and dedupe keys; a video already at its per-video cap is not fetched again. Paid runs also reconcile stored event counts; ambiguous billing/storage history fails for operator review instead of replaying charges. Storage and charging are not a distributed exactly-once transaction. Do not blindly resurrect a failed run; inspect its records first.
Independent implementation; not affiliated with YouTube or Google. Use public content lawfully and respect privacy and platform terms.