# YouTube Comments Scraper (`mlg14/youtube-comments-scraper`) Actor

Scrape public YouTube video comments and replies from video URLs. Export text, commenter details, likes, reply counts, creator signals, and video context.

- **URL**: https://apify.com/mlg14/youtube-comments-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** Videos, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Comments Scraper

Scrape public YouTube comments and replies from one or more video URLs, then export the results as JSON, CSV, or a spreadsheet. This YouTube comments scraper collects the conversation text, public commenter identity, displayed engagement, reply relationships, and video context without requiring a video data key.

Use it when you need a repeatable way to capture a discussion around a specific video. Each result is a flat record with a stable comment identifier, so records can be joined, deduplicated, compared across runs, or loaded into a table. The actor follows public comment pagination and optional reply threads until it reaches the limit you set or the public continuation ends.

### What data can you extract from YouTube?

Each comment and reply has the same set of fields. A reply has `type` set to `reply` and `replyToCid` set to its parent comment ID. A top-level comment has `type` set to `comment` and a null parent ID. The dataset may also contain an error item for an input that could not provide comments; error items are described below.

| Field | Description | Example |
| --- | --- | --- |
| `comment` | Public comment text. | `But the link said free robux 😢` |
| `cid` | Stable comment or reply ID supplied by the site. | `UgwmOpWybGLAIIhhA854AaABAg` |
| `commentUrl` | Direct watch-page link to the comment. | `https://www.youtube.com/watch?v=dQw4w9WgXcQ&lc=UgwmOpWybGLAIIhhA854AaABAg` |
| `author` | Public display name. | A public handle |
| `authorChannelId` | Public commenter channel ID. | `UCdOyiHRpksv1dAMVQWmss3A` |
| `authorChannelUrl` | Public commenter channel link. | A public channel URL |
| `authorAvatarUrl` | Public profile image link. | An image URL |
| `authorIsVerified` | Whether a verified badge is shown. | `false` |
| `videoId` | Video ID shared by all results from that video. | `dQw4w9WgXcQ` |
| `pageUrl` | Canonical video watch link. | `https://www.youtube.com/watch?v=dQw4w9WgXcQ` |
| `commentsCount` | Video-wide count, if the public page exposes one. | `null` |
| `replyCount` | Displayed number of replies to this comment, parsed as a number. | `1` |
| `voteCount` | Displayed likes parsed as an approximate number. | `320000` |
| `voteCountText` | Original displayed like count. | `320K` |
| `publishedTimeText` | Relative age as shown on the page. | `22 minutes ago` |
| `authorIsChannelOwner` | Whether the commenter owns the video channel. | `false` |
| `hasCreatorHeart` | Whether the creator heart is shown. | `true` |
| `isPinned` | Whether a top-level comment is marked as pinned. | `true` |
| `type` | `comment` or `reply`. | `comment` |
| `replyToCid` | Parent ID for a reply; null for a top-level comment. | `null` |
| `title` | Video title. | The title shown on the watch page |

The site currently supplies relative comment age in the public comment response. `publishedTimeText` preserves that wording; it is not an exact publication timestamp. The site also abbreviates large like and reply counts. `voteCountText` preserves the displayed value, while `voteCount` expands abbreviations such as `K` into an approximate integer. Do not treat `320K` as an exact measurement of 320,000.

`commentsCount` is included for a stable shape alongside other comment exports, but the tested public page did not provide the video-wide count. It is null when unavailable. A null count does not mean the video has zero comments; the dataset can still contain many individual records.

### How to scrape YouTube comments

1. Paste one or more public video URLs into **Video URLs**. A normal watch link is the simplest input. Short links and links to Shorts, live pages, or embedded videos are also accepted.
2. Set **Maximum comments per video**. The default is 100. This limit counts both top-level comments and replies when replies are enabled.
3. Choose **Newest first** to monitor recent discussion, or **Top comments** to collect the featured ranking. Leave **Include replies** on when conversations beneath comments matter.
4. Run the actor. It reads the watch page, follows comment continuation pages, and saves each result as a dataset item.
5. Download the dataset in your preferred format. Use `cid` to remove duplicates if you combine multiple runs for the same video.

For a first check, use a video with a visible public comment section and a limit of 40 or 100. A video with disabled comments, a region restriction, an age wall, or an unavailable watch page may produce an error item. Those conditions are tied to the video and the public visitor view, not to the requested limit.

### Input

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `startUrls` | URL list | Required | Public video URLs. Each item can contain a `url` property. The actor also accepts a bare 11-character video ID in a list item. |
| `maxComments` | Integer | `100` | Maximum number of comment and reply records per video. Minimum is 1. |
| `sortCommentsBy` | Choice | `NEWEST_FIRST` | `NEWEST_FIRST` or `TOP_COMMENTS` for top-level comments. |
| `includeReplies` | Boolean | `true` | Follow public reply threads when they are available. |
| `oldestCommentDate` | Text | Empty | Optional UTC date, such as `2026-09-01`, or a relative age, such as `7 days`. This filter uses displayed relative age and is approximate. |
| `maxItems` | Integer | `0` | Overall dataset cap across all input videos. Zero means there is no overall cap. |
| `proxyConfiguration` | Object | Automatic | Optional network configuration for watch pages and comment requests. |

A realistic input for recent public discussion is:

```json
{
  "startUrls": [
    { "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ" }
  ],
  "maxComments": 40,
  "sortCommentsBy": "NEWEST_FIRST",
  "includeReplies": true
}
```

The per-video and overall limits work together. With three URLs, `maxComments` set to 100, and `maxItems` set to 180, the actor can return at most 100 items for any one video and at most 180 items for the whole run. If the first two videos fill the overall cap, the third video is not fetched. If one video exposes only 12 public comments, the actor moves on after those 12; a requested maximum is a ceiling, not a promise of that many records.

The date filter changes the top-level comment order to newest first. A date such as `2026-09-01` is interpreted in UTC. A relative value such as `7 days` is calculated from the start time of the run. Because the source says things like `5 days ago` instead of returning an exact timestamp, the cutoff is approximate near its boundary. A reply to an older top-level comment may also be absent when its parent was excluded by the cutoff.

### Output example

The following values come from a successful 40-item run against a public video. This shortened record omits the profile image URL and several other fields shown in the table above.

```json
{
  "comment": "But the link said free robux 😢",
  "cid": "UgwmOpWybGLAIIhhA854AaABAg",
  "commentUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ&lc=UgwmOpWybGLAIIhhA854AaABAg",
  "videoId": "dQw4w9WgXcQ",
  "pageUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "commentsCount": null,
  "replyCount": 1,
  "voteCount": 1,
  "voteCountText": "1",
  "publishedTimeText": "22 minutes ago",
  "authorIsChannelOwner": false,
  "hasCreatorHeart": false,
  "isPinned": false,
  "type": "comment",
  "replyToCid": null
}
```

Each record can be read independently. `pageUrl` identifies the video, while `commentUrl` points to that specific comment. For replies, `replyToCid` connects the reply to its parent. This makes it possible to build a thread view after export without relying on dataset order. The video title and public commenter fields are repeated on each item so a CSV export remains useful without a separate lookup table.

### Use cases

- **Community monitoring:** Collect recent comments on your own videos to see recurring questions, requests, and points of confusion. Use a modest limit on a regular schedule, then compare comment IDs against earlier runs.
- **Campaign response tracking:** Capture the public discussion beneath a launch video and separate top-level reactions from replies. The reply relationship helps distinguish a new opinion from a response to an existing conversation.
- **Editorial research:** Review comments on a specific topic video for language viewers use, examples they mention, and questions left unanswered. Keep the original comment URL when passing a finding to an editor.
- **Moderation review:** Export public text and engagement to prioritize manual review. The actor does not classify harmful content or make moderation decisions; it supplies the source records for your own workflow.
- **Audience analysis:** Compare the questions and reactions under several public videos. Use `videoId` and `cid` as stable keys, while remembering that featured ordering and visible counts can change between runs.
- **Creator interaction review:** Identify comments marked as pinned, posted by the channel owner, or hearted by the creator. These flags describe the public page state at collection time.

For longitudinal work, keep the run date with your exported file. Relative ages such as `2 days ago` will read differently in a later run even for the same comment. A stable comment ID lets you identify the same item over time, but the visible like count and creator signals may change.

### How much does it cost to scrape YouTube comments?

The configured result charge is **$1 per 1,000 dataset items**, or **$0.001 per delivered item**. A normal comment, a reply, and an error item each occupy one dataset item. The actual bill can depend on the account and run settings; check the run estimate before collecting a large dataset.

| Delivered items | Result charge |
| ---: | ---: |
| 100 | $0.10 |
| 1,000 | $1.00 |
| 5,000 | $5.00 |

For example, ten videos returning 100 items each produce up to 1,000 records and a $1.00 result charge. Five videos returning only 20 public items each produce 100 records and a $0.10 result charge, even if the requested per-video maximum was higher. An overall `maxItems` cap is useful when you need a hard ceiling on record volume. It stops data collection as soon as the dataset reaches that cap.

Reply-heavy videos can require more requests for the same number of output records. In a verified 40-item run, the actor completed successfully with all required fields filled. Run time and infrastructure usage vary with video availability, public page behavior, reply depth, and network conditions. Do not multiply the short verification run time into a guarantee for thousands of comments.

### Tips for best results

Choose the sort order for your question. **Newest first** is suitable when you want current reactions; **Top comments** follows the site's featured ranking and can surface older high-engagement entries. Sort order applies to top-level comments. Replies are collected from the reply threads attached to those comments, so a reply is not globally sorted against every top-level entry.

Leave replies enabled when the discussion itself matters. The actor gathers a spread of top-level comments before traversing reply threads, so a single popular pinned comment does not consume the entire small output limit. If you only need independent top-level opinions, turn replies off. That generally reduces requests and gives your limit entirely to top-level comments.

Start with a small maximum to verify the video and fields, then increase it. Popular videos may expose long chains of continuations. The actor stops at the requested per-video limit, the overall cap, or the end of public continuations. It also tracks comment IDs while processing a video so repeated entries from pagination are not added twice.

Use more than one URL when comparing videos, but remember that the overall cap is processed in input order. Put the highest-priority videos first if `maxItems` may stop the run early. If a particular video has restricted or disabled comments, other input videos can still return normal records. Inspect error items before treating a low result count as a sign of low audience activity.

Keep both `voteCount` and `voteCountText` when analyzing engagement. The text field records what a visitor saw; the numeric field is convenient for sorting and charting. Large values are rounded by the source display, so use them as approximations. The same caution applies to `replyCount` where the display uses a compact value.

### Limits

Only public comments visible to an anonymous visitor can be collected. Private, deleted, hidden, disabled, age-gated, or region-restricted discussions may not appear. Moderation and ranking can also change what a visitor sees. The actor does not sign in, reveal private details, or bypass permissions attached to an account.

The source currently provides relative publication age rather than an exact UTC timestamp for individual comments. `oldestCommentDate` is a convenience filter based on that visible age, not an exact historical query. For strict time-series boundaries, retain the run time and apply your own review to comments near the cutoff.

The tested watch page did not disclose a video-wide total comment count, so `commentsCount` is null. A requested maximum may exceed the number of public items reachable through the site's continuation chain. Reply counts and like counts can be abbreviated in the interface. A value such as `1K` is expanded approximately in the numeric output.

An unavailable source produces an item with `input`, `url`, `error`, and `note`. These fields let downstream processing separate failures from normal comments. Check `error` before interpreting an item as comment data. A failed video does not erase records already collected from other URLs in the same run.

The comment endpoint and response shape are controlled by the site and can change. A layout or response change can temporarily affect extraction, pagination, or a field's availability. The actor uses the current watch-page configuration rather than a fixed client version or hard-coded continuation token, which helps it follow routine changes but cannot guarantee permanent compatibility.

### Structured requests

The input names are intended to make scheduled and programmatic runs straightforward. A request such as “collect the newest 100 comments and replies from these two video URLs” maps to `startUrls`, `maxComments`, `sortCommentsBy`, and `includeReplies`. A request such as “capture up to 500 top-level comments across five videos” maps to `startUrls`, `includeReplies: false`, and `maxItems: 500`.

When passing the output into another system, key each record by `cid` and group it by `videoId`. Keep `type` and `replyToCid` so replies do not become disconnected from their parent. Check for `error` items before counting comments, and treat null fields as missing public values rather than zero. These conventions work for both one-off exports and repeated monitoring.

### FAQ

#### Can I scrape comments from more than one video?

Yes. Add multiple entries to `startUrls`. The actor processes them in order and applies `maxComments` separately to each video. `maxItems` limits the entire run. Duplicate video IDs in the input are skipped during a run.

#### Are replies included?

Yes, by default. Set `includeReplies` to `false` when you want only top-level comments. Every included reply has `type: "reply"` and a parent ID in `replyToCid`. The per-video comment limit includes both kinds of item.

#### Can I get the exact date a comment was posted?

The tested public response exposes relative wording such as `22 minutes ago` or `1 year ago`. The actor preserves that wording in `publishedTimeText`. It does not invent an exact timestamp from an imprecise relative label.

#### Why is the video-wide comment count null?

The current public watch page and comment response did not expose a reliable total for the tested video. The `commentsCount` key stays in the output for a stable schema and is null when no total is supplied. The number of dataset items is the number actually collected, not the total number on the video.

#### What happens if comments are disabled or a video is unavailable?

The actor emits an error item for that input when it cannot retrieve public comments. The item includes the input URL and a short explanation. Other valid URLs in the same input can still produce comments.

#### Do I need a proxy?

The default network configuration is set up for public watch-page and continuation requests. The successful verification runs returned comments without a challenge. Availability can differ by network, country, and time, so a blocked request may require a different route or may remain inaccessible to an anonymous visitor.

#### Is it legal to collect comments?

Use public data for a legitimate purpose. Respect the site's terms, applicable privacy rules, and any obligations around storing or sharing people's comments and profile details. Avoid collecting or using personal data in ways the commenters would not reasonably expect.

### Integrations

The dataset can be downloaded directly or read programmatically. Schedule repeated runs to monitor a video, send a completion notification, or load the flat records into a spreadsheet or database. Use `cid` for deduplication and `videoId` for grouping when combining exports across runs. A webhook can signal that fresh comments or an error item are ready for processing.

### Support

Open an issue on the actor's Issues tab with the video URL, the input used, and the observed error or missing field. Include a run link when available so the failing public response can be checked. Requests for additional public fields are welcome.

# Actor input Schema

## `startUrls` (type: `array`):

Public YouTube video URLs. Watch, Shorts, live, embedded, and short links are accepted.

## `maxComments` (type: `integer`):

Stop after this many comments and replies per video.

## `sortCommentsBy` (type: `string`):

Order top-level comments by featured ranking or newest first.

## `includeReplies` (type: `boolean`):

Follow public reply threads as well as top-level comments.

## `oldestCommentDate` (type: `string`):

Optional UTC date such as 2026-09-01 or relative age such as 7 days. YouTube supplies relative times, so this filter is approximate and selects newest-first order.

## `maxItems` (type: `integer`):

Hard cap across all videos. Zero means no total cap.

## `proxyConfiguration` (type: `object`):

Proxy settings for public watch pages and comment requests.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    }
  ],
  "maxComments": 100,
  "sortCommentsBy": "NEWEST_FIRST",
  "includeReplies": true,
  "maxItems": 0,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `comments` (type: `string`):

Comments, replies, and any error items in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/youtube-comments-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ" }] }

# Run the Actor and wait for it to finish
run = client.actor("mlg14/youtube-comments-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    }
  ]
}' |
apify call mlg14/youtube-comments-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/youtube-comments-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TXa6zP4BaEggxG5Mz/builds/a9G6cDEJq5ehgERP0/openapi.json
