# YouTube Scraper (`mlg14/youtube-scraper`) Actor

Scrape public YouTube videos from search results, channels, playlists, and direct URLs, with titles, views, channels, dates, thumbnails, and optional detailed video metadata.

- **URL**: https://apify.com/mlg14/youtube-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** Videos, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Scraper

Scrape YouTube video listings from search results, public channels, and playlists, or collect metadata for a direct video URL. Export YouTube data to JSON, CSV, or Excel for research, monitoring, and reporting; the actor offers a practical YouTube API alternative for public page data without requiring an API key.

The actor follows listing pages and their continuation requests. A search term can yield more than the first screen of videos, a channel can go beyond its initial grid, and a playlist can continue past the first 100 entries when YouTube provides more items. It stores one flat dataset item per unique video. Optional detail fetching visits each video page for its full description and other public counters.

### What data can you extract from YouTube?

Every result is a video record. Listing pages provide the fields that are visible there; a direct video URL or the **Fetch full video details** option can add richer fields. Values that the public page does not expose are `null`. The actor does not invent numbers or replace missing data with estimates.

| Field | Description | Example |
| --- | --- | --- |
| `id` | Stable video ID | `QXeEoD0pB3E` |
| `url` | Watch URL; playlist items retain their list and position | `https://www.youtube.com/watch?v=QXeEoD0pB3E&list=...&index=1` |
| `fromYTUrl` | Source URL that produced the item | `https://www.youtube.com/playlist?list=...` |
| `title` | Public video title | `#0 Python for Beginners \| Programming Tutorial` |
| `thumbnailUrl` | Public thumbnail image URL | `https://i.ytimg.com/vi/QXeEoD0pB3E/hqdefault.jpg` |
| `channelName` | Name shown for the publishing channel | `Example Channel` |
| `channelUrl` | Link to the channel page | `https://www.youtube.com/@example` |
| `channelId` | Stable channel identifier when exposed | `UC...` |
| `date` | Date or relative age shown on the source page | `8y ago` |
| `publishedAt` | Exact publication date from a video page | `2018-01-01` |
| `viewCount` | Public views, rounded if a listing abbreviates the number | `9900000` |
| `duration` | Listing duration; detailed pages may give seconds as text | `1:06` |
| `description` | Full public video description from a video page | `Lesson overview...` |
| `descriptionSnippet` | Short excerpt shown in search results | `Learn the basics...` |
| `likes` | Public like count when supplied by the video page | `12000` |
| `commentsCount` | Public comment count when supplied by the video page | `1200` |
| `numberOfSubscribers` | Public subscriber count when shown | `250000` |
| `videoType` | Content class | `video`, `short`, or `live` |
| `playlistId` | Source playlist ID, if applicable | `PLsyeobzWxl7poL9JTVyndKe62ieoN-MZ3` |
| `playlistIndex` | One-based position among collected playlist items | `1` |

The dataset is flat, so each row can be filtered or grouped without first expanding nested objects. You can group by `channelId` to compare channels, use `id` to deduplicate across runs, and retain `fromYTUrl` to trace where each video was found. `url` is the direct watch link, while `fromYTUrl` records the search or collection page. A video that appears in multiple supplied sources is emitted once per run.

### How to scrape YouTube

1. Enter one or more search terms, one or more public YouTube URLs, or both. Supported URLs include watch pages, short links, channel video pages, playlists, and search result pages.
2. Set **Maximum videos per source** to control how far the actor paginates each source. Set **Maximum videos total** to cap the entire run.
3. Enable **Fetch full video details** if you need the complete description, exact publication date, and the public counters that appear on video pages.
4. Run the actor and inspect the dataset. Download JSON, CSV, or Excel from the dataset view, or retrieve items through the platform API.

A search term behaves like entering that phrase in YouTube's search box. A direct video URL produces one video item if the video is publicly available. A channel URL without a tab is directed to its Videos tab. To target public Shorts or live listings for a channel, supply the `/shorts` or `/streams` tab URL. A playlist URL produces its video entries in playlist order.

### Input

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `searchQueries` | String array | Empty | Terms to search on YouTube. Each term is a separate source. |
| `startUrls` | URL list | Empty | Public watch, channel, playlist, short, or search result URLs. |
| `maxResults` | Integer | `50` | Maximum videos from each source. `0` collects available results up to the actor's page safety limit. |
| `maxItems` | Integer | `100` | Maximum output rows across all sources. `0` removes this total cap. |
| `includeVideoDetails` | Boolean | `false` | Visit each video page to collect richer metadata. |
| `proxyConfiguration` | Proxy settings | Platform proxy | Connection settings for public page requests. |

Provide at least one search term or URL. The actor processes supplied URLs first and then search terms. If you provide both, `maxItems` still applies across the complete run. For a direct watch URL, `maxResults` does not change the one-video result. The total cap can stop processing before later sources are reached, so set it high enough when you want every source represented.

A realistic input for a focused search plus a direct video is:

```json
{
  "startUrls": [
    {"url": "https://www.youtube.com/watch?v=jfKfPfyJRdk"}
  ],
  "searchQueries": ["python tutorial"],
  "maxResults": 34,
  "maxItems": 35,
  "includeVideoDetails": false
}
```

For a playlist, pass its normal `playlist?list=` URL in `startUrls`. For a channel, pass a `/@handle/videos` URL. A `youtu.be` short link is normalized to a watch URL. The actor accepts direct links with the main, mobile, or short-link host; unrelated domains are rejected. Public videos that are removed, private, age restricted, or unavailable in the run's region may not produce an item.

### Output example

This shortened item came from a verified playlist run. Optional fields and the long thumbnail query string are omitted here; the dataset contains the actual URL and all documented keys.

```json
{
  "id": "QXeEoD0pB3E",
  "url": "https://www.youtube.com/watch?v=QXeEoD0pB3E&list=PLsyeobzWxl7poL9JTVyndKe62ieoN-MZ3&index=1",
  "fromYTUrl": "https://www.youtube.com/playlist?list=PLsyeobzWxl7poL9JTVyndKe62ieoN-MZ3",
  "title": "#0 Python for Beginners | Programming Tutorial",
  "date": "8y ago",
  "viewCount": 9900000,
  "duration": "1:06",
  "videoType": "video",
  "playlistId": "PLsyeobzWxl7poL9JTVyndKe62ieoN-MZ3",
  "playlistIndex": 1
}
```

The listing displayed an abbreviated view count. The numeric value above reflects that abbreviation rather than an exact current count. Turn on detail fetching to request the video page's more precise public count. Published dates, descriptions, likes, comments, and subscriber totals can still be absent when the page does not expose them.

### Use cases

- **Topic discovery:** Search phrases related to a market, product category, or research question, then review titles and publication ages in one table.
- **Channel monitoring:** Collect recent uploads from public channel video pages and compare the latest dataset with an earlier run using `id`.
- **Playlist inventory:** Export playlist titles, URLs, positions, and view counts for catalog work or content audits.
- **Video reporting:** Capture public views and metadata for a known set of video URLs at regular intervals.
- **Editorial research:** Filter search results by channel or title terms after export, and follow `url` back to the public video.
- **Collection quality checks:** Use `fromYTUrl` and `playlistIndex` to verify that a playlist export has the expected ordering and source.

For recurring monitoring, store the results of each run rather than assuming view counts or search rankings remain fixed. A view count is a snapshot at collection time. Relative date labels are as displayed by YouTube at that time; they are not calculated publication dates. If you need an exact date for each result, enable detail fetching and inspect `publishedAt`.

### How much does it cost to scrape YouTube?

The configured result event is **$0.001 per item**, equivalent to **$1 per 1,000 delivered results**. A 100-result run has a result charge of $0.10; 1,000 results are $1; 5,000 results are $5. Platform usage is included in the configured result price. Your maximum charge setting can stop a run when its allowed event budget is reached.

These examples count delivered dataset items, not page requests. A playlist that returns 105 videos is 105 result events. A source that yields no public videos contributes no result events. Duplicate videos within one run are emitted once. Enabling full detail fetching adds requests and may make the run slower, but the result event count still follows emitted videos. Check the run's charge summary for the actual amount billed.

The default `maxItems` value of 100 is useful for a small trial. To estimate a larger export, multiply the intended item count by $0.001. If you have several sources, the total cap is shared; for example, four search terms with `maxResults` set to 50 can produce at most 200 items, and a lower `maxItems` value will stop sooner.

### Tips for best results

Use a specific search phrase if broad searches include unrelated videos. Start with 30 to 50 results, inspect the dataset, then increase `maxResults` only when the first batch matches your intent. Search order follows the public results returned by YouTube for the run's location and time. The actor does not apply a separate ranking or add unsupported filters.

For channel exports, prefer the explicit Videos, Shorts, or Live tab URL that matches the collection you want. A channel's root URL is converted to its Videos tab. Each tab can show a different set of items. For playlists, use the playlist page rather than a watch URL carrying a `list` parameter; the watch URL is handled as a single video.

Leave detail fetching off for broad discovery. Listing pages already provide the title, source, channel, relative date, thumbnail, views when shown, and often duration. Enable it when a project needs descriptions or exact dates and can accept one extra watch-page request per video. A direct video URL always uses its watch page, even with the option off.

Use `maxItems` as a run-wide safeguard. If you are collecting many channels or terms, raise it only after deciding how many rows you want to pay for. Data from earlier sources may consume the cap before the actor reaches later sources. To ensure balanced coverage, run separate jobs for separate topics or channels.

### Limits

YouTube changes page layouts, public counters, and continuation responses. The actor reads public JSON embedded in pages and follows the continuation requests those pages expose. Its safety limit is 200 continuation pages per source. It stops earlier when the source has no more continuation token, the source limit is reached, or the total result limit is reached.

A successful test collected 35 channel videos across the first grid and a continuation, 34 search results after a direct video, and 105 playlist videos across the initial 100 and a continuation. These are observed examples, not a guarantee that every channel, query, or playlist will have that many accessible results. Search rankings and counts can vary by region and time.

Listing counts can be abbreviated, so `viewCount` may be rounded. Some listings omit views for newly published or live content. Video pages may omit likes, comments, subscribers, or captions, and this actor does not download subtitles or comments. It also does not collect private videos, account-only information, viewer identities, or contact information. Channel tabs and playlists may contain unavailable entries; those are skipped when no public video metadata is returned.

The actor deduplicates by video ID within a run. If the same video appears in a playlist and a search result, only the first occurrence is retained, including its original `fromYTUrl`. For separate provenance records per source, run the sources separately. The playlist position is the position among collected entries and can differ from a displayed index if unavailable entries are hidden.

### Agent workflows (MCP)

The input and output schemas make the actor usable in an agent workflow. A connected client can provide search terms or URLs, start a run, then read dataset items. Keep prompts narrow enough that the intended source and row count are explicit.

- “Search YouTube for `python tutorial`, collect 30 public videos, and summarize the channels and view counts in a table.”
- “Collect the first 100 public videos from this playlist URL and return each title, watch URL, and playlist position.”

A client should treat optional fields as nullable. If it needs a full description or exact publication date, it should set `includeVideoDetails` to `true`. If it needs a time series, it should schedule separate runs and join their datasets by `id`.

### FAQ

#### Is scraping YouTube legal?

This actor reads public page data. Follow the site's terms and applicable privacy and copyright rules for your use case. Do not use exported data to identify, profile, or contact individuals without an appropriate basis. The actor does not access private or login-only content.

#### Do I need to configure a proxy?

The default input uses the platform proxy. You can change the proxy settings if your run environment requires it. Public pages and continuation responses can still vary by location or temporarily fail, so review run logs if a particular source produces fewer items than expected.

#### How fast is a run?

Speed depends mainly on source count, requested results, and whether full details are enabled. Listing pages return many videos per request. Detail mode adds a video-page request for each emitted item. A small listing-only test is the fastest way to check a new source.

#### Can I schedule or monitor the actor?

Yes. Schedule repeated runs through the platform and inspect each run's status and dataset. For monitoring, keep the previous dataset and compare video IDs, view counts, or titles. A scheduled run should use explicit caps so a growing channel or broad query does not unexpectedly expand the export.

#### Can I export to a spreadsheet?

Yes. Download the dataset as CSV or Excel, or connect the dataset API to a spreadsheet workflow. Each result is a flat object, making columns such as `title`, `channelName`, `viewCount`, and `url` straightforward to work with.

#### Why is a field empty?

The source may not show that field on a listing page, or a video's public page may not expose it. For descriptions and exact dates, enable full details. For likes, comments, or subscribers, even a detail request can return `null`. A missing value is kept missing rather than inferred.

#### Can this collect every video from a channel or playlist?

It follows public continuation tokens until it reaches the requested limit or the source stops returning more items, with a 200-page safety limit. Public availability, regional differences, hidden videos, and site changes can reduce coverage. Test a representative source before planning a large export.

### Integrations

Use the platform API to start runs and read datasets from another application. Webhooks can react to completed runs, while scheduling supports recurring collection. Exported JSON or CSV can feed spreadsheet, database, reporting, or automation workflows. Store historical datasets if you want to measure changes over time; the actor emits a current snapshot rather than maintaining history internally.

### Support

Open an issue on the Issues tab; we reply within 24h and add fields on request. Include the input source, expected result count, and a run link when reporting a problem. A public URL that demonstrates a missing field or a changed listing layout helps reproduce it quickly.

# Actor input Schema

## `searchQueries` (type: `array`):

Search YouTube for each term and collect video results.

## `startUrls` (type: `array`):

Public video, channel, playlist, Shorts, or search result URLs. Both URLs and search terms are processed.

## `maxResults` (type: `integer`):

Maximum videos per search term, channel, playlist, or search URL. Zero collects available results up to the page safety limit.

## `maxItems` (type: `integer`):

Hard cap across all sources. Zero has no total cap.

## `includeVideoDetails` (type: `boolean`):

Visit each video page for its full description, exact publication date, current views, likes when available, comments count, and subscriber count. Slower than listing-only mode.

## `proxyConfiguration` (type: `object`):

Proxy connection for public YouTube pages. The default uses Apify Proxy.

## Actor input object example

```json
{
  "searchQueries": [
    "python tutorial"
  ],
  "startUrls": [
    {
      "url": "https://www.youtube.com/@GoogleDevelopers/videos"
    }
  ],
  "maxResults": 50,
  "maxItems": 100,
  "includeVideoDetails": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped items in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "python tutorial"
    ],
    "startUrls": [
        {
            "url": "https://www.youtube.com/@GoogleDevelopers/videos"
        }
    ],
    "maxResults": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/youtube-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["python tutorial"],
    "startUrls": [{ "url": "https://www.youtube.com/@GoogleDevelopers/videos" }],
    "maxResults": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("mlg14/youtube-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "python tutorial"
  ],
  "startUrls": [
    {
      "url": "https://www.youtube.com/@GoogleDevelopers/videos"
    }
  ],
  "maxResults": 50
}' |
apify call mlg14/youtube-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/youtube-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Nem8oAaPw2k0gtmJl/builds/acYrzr7KtD864CWBf/openapi.json
