# YouTube Scraper Pro: Videos & Channels (`smart_albatross/youtube-scraper-pro`) Actor

Fast YouTube scraper for videos, channels, search, playlists, comments, and transcripts. Exact limits, restart recovery, no API key.

- **URL**: https://apify.com/smart\_albatross/youtube-scraper-pro.md
- **Developed by:** [Dev](https://apify.com/smart_albatross) (community)
- **Categories:** Videos, Social media
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 primary youtube results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Scraper Pro

Extract public YouTube videos, Shorts, live streams, channels, playlists, search results, comments, and timestamped transcripts—without a YouTube Data API key or a browser.

Use it for creator research, content monitoring, lead discovery, trend analysis, dataset building, and YouTube workflow automation. Start with URLs, search queries, or both in the same run.

### Why choose this YouTube scraper?

| What you need | What this Actor does |
| --- | --- |
| Predictable pricing | Charges only for successful primary records. No start fee. Comments, transcripts, notices, and errors have no additional event charge. Platform usage is included. |
| Broad coverage | One Actor handles videos, Shorts, streams, channels, playlists, hashtags, search, comments, and transcripts. |
| Reliable bulk runs | Exact concurrent limits, global deduplication, isolated source failures, retry handling, automatic checkpoints, and restart-safe dataset recovery. |
| Fast, efficient collection | HTTP-first extraction runs without a browser and uses 256 MB by default. |
| Integration-ready data | Stable `recordType` values, normalized fields, native dataset views, machine-readable errors, and a compact run summary. |
| Safe cost controls | Per-source and run-wide record caps work together with Apify's maximum charge per run. |

In a paid cloud benchmark, the Actor produced exactly 1,000 unique results and 1,000 matching billed events in 136.5 seconds at 256 MB, with zero retries, 429s, 5xx responses, or network errors. Batched paid writes made this identical workload about 30% faster than the previous paid build. A separate forced-reboot test resumed a run, recovered the dataset tail, and finished with no duplicate records. Actual speed varies with source type, enrichment, network conditions, and YouTube responses.

### Pricing

You pay only when the Actor stores a successful primary result: a `video`, `short`, `stream`, `channel`, or `playlist` record.

| Apify plan | Price per 1,000 primary results |
| --- | ---: |
| Free | $0.60 |
| Bronze | $0.50 |
| Silver | $0.45 |
| Gold | $0.40 |
| Platinum | $0.40 |
| Diamond | $0.40 |

Included at no additional event charge:

- Actor starts and empty searches
- Comments and comment pagination
- Timestamped transcript enrichment
- `notice` and `error` records
- Apify platform usage during the run

For example, 10,000 primary results cost $6.00 on Free, $5.00 on Bronze, or $4.00 on Gold and higher plans. Set **Maximum total charge** when starting a run to enforce an additional billing ceiling.

### Quick start

Scrape a channel and a direct video in one run:

```json
{
  "startUrls": [
    { "url": "https://www.youtube.com/@Apify" },
    { "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ" }
  ],
  "channelSections": ["videos", "shorts", "live"],
  "maxResultsPerSource": 50,
  "maxTotalResults": 100,
  "includeComments": true,
  "maxCommentsPerVideo": 20,
  "includeTranscript": true
}
```

Search YouTube without opening every result page:

```json
{
  "searchQueries": [
    "web scraping tutorials",
    "Apify YouTube automation"
  ],
  "searchResultTypes": ["video", "short", "channel", "playlist"],
  "uploadDate": "month",
  "searchSort": "relevance",
  "maxResultsPerSource": 100,
  "maxTotalResults": 200
}
```

`startUrls` and `searchQueries` can be combined. Each URL or query is processed independently, so one unavailable target does not discard successful results from other sources.

### Supported YouTube URLs

- Videos: `youtube.com/watch?v=...` and `youtu.be/...`
- Shorts, live streams, and embeds: `/shorts/...`, `/live/...`, `/embed/...`
- Channels: `/@handle`, `/channel/...`, `/c/...`, and `/user/...`
- Playlists: `/playlist?list=...`
- Hashtags: `/hashtag/...`
- Search pages: `/results?search_query=...`

### What you can extract

#### Videos, Shorts, and streams

- ID, canonical URL, title, description, publication date, and duration
- View, like, and comment counts when YouTube exposes them
- Channel identity, handle, verification, URL, and thumbnails
- Video thumbnails, tags, category, live/upcoming flags, and availability information
- Optional expanded details for search, channel, hashtag, and playlist results

#### Channels and playlists

- Channel identity, description, subscriber and video counts, and thumbnails
- Videos, Shorts, and Live channel tabs with continuation-page pagination
- Playlist identity, metadata, item counts, and paginated playlist items

#### Comments and transcripts

- Top or newest comments with author, votes, replies, pinned/hearted flags, owner flag, and deep link
- Timestamped transcript segments with text, start/end milliseconds, and available languages
- Transcript fallback through signed public caption tracks when YouTube's transcript endpoint is unavailable

### Input options

| Input | Purpose | Default |
| --- | --- | --- |
| `startUrls` | Video, channel, playlist, hashtag, or search URLs | `[]` |
| `searchQueries` | Plain-text YouTube searches | `[]` |
| `searchResultTypes` | `video`, `short`, `channel`, `playlist` | video + short |
| `maxResultsPerSource` | Exact cap for every URL or query | `50` |
| `maxTotalResults` | Exact cap for the entire run | `1000` |
| `channelSections` | Channel tabs to collect | all three |
| `includeChannelRecord` | Store channel metadata before its content | `true` |
| `includeComments` | Fetch comments for detailed videos | `false` |
| `maxCommentsPerVideo` | Comment cap per enriched video | `20` |
| `commentsSortBy` | Top comments or newest first | top |
| `includeTranscript` | Add timestamped transcript segments | `false` |
| `enrichExpandedVideos` | Open every expanded list result for full details | `false` |
| `searchSort` | Relevance or popularity | relevance |
| `uploadDate` | Today, week, month, year, or all | all |
| `duration` | Under 3, 3–20, over 20 minutes, or all | all |
| `language` / `country` | Localize YouTube responses | `en` / `US` |
| `maxConcurrency` | Parallel top-level sources | `4` |
| `maxRequestRetries` | Retries for transient failures | `3` |
| `maxSources` | Maximum URLs and queries combined | `200` |
| `maxPagesPerSource` | Maximum paginated YouTube pages per source | `100` |
| `maxRequestsPerSource` | Maximum YouTube HTTP attempts per source | `500` |
| `requestTimeoutMs` | Timeout for each HTTP attempt | `20000` |
| `proxyConfiguration` | Optional Apify or custom proxy | disabled |

`maxResultsPerSource` and `maxTotalResults` include every stored record type, including free comments, notices, and errors. This makes the dataset size deterministic as well as the bill.

The source, page, and request limits also bound work when YouTube returns empty or repeated pages. They can be raised for larger jobs. A `PAGE_LIMIT_REACHED` or `REQUEST_LIMIT_REACHED` notice means available results may be partial; repeated URLs or queries that were already covered by another source do not produce a misleading `NO_RESULTS` notice.

### Fast list mode or full enrichment

The default list path returns metadata already present on search, channel, hashtag, and playlist pages. This is the fastest and least expensive way to collect large result sets.

Enable `enrichExpandedVideos` when every expanded video needs full video-page fields, comments, or transcripts. Enrichment adds at least one request per video. Direct video URLs always use the full-detail path.

### Output format

Every dataset row contains a stable type and source context:

```json
{
  "recordType": "video",
  "id": "dQw4w9WgXcQ",
  "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "title": "Example video",
  "viewCount": 123456,
  "channel": {
    "id": "UC...",
    "name": "Example channel",
    "url": "https://www.youtube.com/channel/UC..."
  },
  "source": {
    "kind": "search",
    "value": "web scraping tutorials",
    "index": 0
  },
  "scrapedAt": "2026-09-27T08:30:00.000Z"
}
```

`recordType` is one of:

- `video`, `short`, or `stream`
- `channel` or `playlist`
- `comment`
- `notice` for a valid source with no public results
- `error` for an isolated source failure

The default dataset contains the records. The `OUTPUT` key-value-store record contains source outcomes, counts, duplicate and recovery metrics, first-result latency, duration, HTTP request/retry status, and whether a result or billing limit was reached.

### Reliability behavior

- Retries HTTP 408, 425, 429, and 5xx responses with exponential backoff and jitter
- Deduplicates globally by record type and stable ID or canonical URL
- Reserves output capacity before concurrent writes, preventing limit overshoot
- Saves state at startup, in source batches, periodically, and before migration
- Recovers rows written after the last checkpoint following a restart
- Preserves partial results when one source, comment page, or enrichment fails
- Keeps free comments eligible for output after the final affordable video, as long as the dataset and per-source record caps still have room
- Reports a comment fetch failure as a separate free error row tied to the video, without rewriting or charging the video again
- Returns stable error codes for rate limits, network failures, unavailable/private content, transcript failures, invalid input, and extraction changes

### API example

Run synchronously and receive dataset rows:

```bash
curl "https://api.apify.com/v2/acts/smart_albatross~youtube-scraper-pro/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{
    "searchQueries": ["web scraping tutorials"],
    "maxTotalResults": 25
  }'
```

For larger jobs, start the Actor asynchronously and read the default dataset after the run finishes.

### Responsible use and limitations

This Actor requests public YouTube data only. You are responsible for complying with applicable law, privacy and copyright rules, contractual restrictions, and YouTube's terms.

- Fields vary by locale, region, account state, content type, and YouTube experiments.
- Some list pages expose abbreviated counts or omit fields; full enrichment can provide more detail.
- Private, deleted, members-only, age-restricted, and region-restricted content may be unavailable or partial.
- Transcripts require public captions or transcripts for the selected video.
- The Actor does not download video or audio media.

# Actor input Schema

## `startUrls` (type: `array`):

Video, Shorts, live, channel, @handle, playlist, hashtag, or search-result URLs. URLs and search queries can be combined in one run.

## `searchQueries` (type: `array`):

Terms to search on YouTube. Each query has its own result limit and error isolation.

## `searchResultTypes` (type: `array`):

Kinds of items to return for search queries.

## `maxResultsPerSource` (type: `integer`):

Hard limit across every record emitted by one input URL or query, including metadata, comments, notices, and errors.

## `maxTotalResults` (type: `integer`):

Run-wide hard cap across every source and record type. Prevents accidental unbounded jobs.

## `channelSections` (type: `array`):

Content tabs to scrape for channel URLs.

## `includeChannelRecord` (type: `boolean`):

Emit a separate channel record before videos from a channel URL.

## `includeComments` (type: `boolean`):

Fetch comments for direct video URLs. Search/channel expansion does not fetch comments unless enrichExpandedVideos is enabled.

## `maxCommentsPerVideo` (type: `integer`):

Hard comment limit for each enriched video.

## `commentsSortBy` (type: `string`):

Sort comments by YouTube's top ranking or newest first.

## `includeTranscript` (type: `boolean`):

Fetch the available YouTube transcript for direct videos and place timestamped segments on the video record.

## `enrichExpandedVideos` (type: `boolean`):

Open every expanded video to collect full details and optional comments/transcripts. Slower but richer.

## `searchSort` (type: `string`):

Ordering applied to search queries.

## `uploadDate` (type: `string`):

YouTube upload-date filter for search queries.

## `duration` (type: `string`):

YouTube duration filter for search queries.

## `language` (type: `string`):

YouTube interface language used for localized text, e.g. en, es, de, hi.

## `country` (type: `string`):

Two-letter region code used for localized results, e.g. US, IN, GB.

## `maxConcurrency` (type: `integer`):

Parallel top-level sources. Lower this when using fragile proxies or large enrichments.

## `maxRequestRetries` (type: `integer`):

Retries transient network and YouTube 429/5xx failures with exponential backoff and jitter.

## `maxSources` (type: `integer`):

Combined safety limit for start URLs and search queries. Large jobs can be split into multiple Actor runs.

## `maxPagesPerSource` (type: `integer`):

Bounds search, channel, playlist, hashtag, and comment pagination for each URL or query. Prevents repeated continuation pages from running indefinitely.

## `maxRequestsPerSource` (type: `integer`):

Counts YouTube requests and retries for each URL or query, including enrichment and transcript requests.

## `requestTimeoutMs` (type: `integer`):

Timeout for each YouTube HTTP attempt. A timed-out request can be retried up to the configured retry limit.

## `proxyConfiguration` (type: `object`):

Optional Apify or custom proxy. Leave disabled for the cheapest HTTP-first path; enable when your traffic is throttled.

## `debug` (type: `boolean`):

Log parser diagnostics. Inputs, cookies, tokens, and full proxy URLs are never logged.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.youtube.com/@Apify"
    }
  ],
  "searchQueries": [],
  "searchResultTypes": [
    "video",
    "short"
  ],
  "maxResultsPerSource": 50,
  "maxTotalResults": 1000,
  "channelSections": [
    "videos",
    "shorts",
    "live"
  ],
  "includeChannelRecord": true,
  "includeComments": false,
  "maxCommentsPerVideo": 20,
  "commentsSortBy": "TOP_COMMENTS",
  "includeTranscript": false,
  "enrichExpandedVideos": false,
  "searchSort": "relevance",
  "uploadDate": "all",
  "duration": "all",
  "language": "en",
  "country": "US",
  "maxConcurrency": 4,
  "maxRequestRetries": 3,
  "maxSources": 200,
  "maxPagesPerSource": 100,
  "maxRequestsPerSource": 500,
  "requestTimeoutMs": 20000,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "debug": false
}
```

# Actor output Schema

## `results` (type: `string`):

Videos, Shorts, streams, channels, playlists, comments, no-result notices, and isolated error records in the default dataset.

## `summary` (type: `string`):

Source outcomes, record and duplicate counts, restart recovery, request/retry telemetry, latency, result limits, and billing-limit status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.youtube.com/@Apify"
        }
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("smart_albatross/youtube-scraper-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://www.youtube.com/@Apify" }],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("smart_albatross/youtube-scraper-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.youtube.com/@Apify"
    }
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call smart_albatross/youtube-scraper-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,smart_albatross/youtube-scraper-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JgyL6mjIQj6dQcK5D/builds/LFWipmq0Pg4D7XSih/openapi.json
