# YouTube Music Song Scraper (`w3crawler/youtube-music-song-scraper`) Actor

Search YouTube Music songs, browse Music playlists and pages, and extract rich returned song and player metadata with pagination.

- **URL**: https://apify.com/w3crawler/youtube-music-song-scraper.md
- **Developed by:** [w3crawler](https://apify.com/w3crawler) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.99 / 1,000 songs

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube Music Song Scraper

Collects song/video records from public YouTube and YouTube Music pages. Search uses the public YouTube results page; direct watch and playlist/browse URLs are also supported. The actor parses page-embedded data and never stores signed media or caption URLs.

### Dataset fields

Each record can include:

- stable video ID, title, description, accessibility text, duration and seconds, displayed/parsed views, publication text, release date/year, explicit/live/upcoming flags, badges, thumbnail variants with dimensions, and public YouTube/YouTube Music URLs
- artist objects with names, IDs, and public channel links; album name, Music browse ID, and public album URL when the page exposes them
- source type/query/URL, rank/page, public extraction provenance, locale, initial result coverage, continuation visibility, request limits, and scrape time
- optional safe public watch-page metadata: description length, keywords, tags, category, playability status, caption-language summaries, and audio/video format counts. Caption base URLs and signed stream URLs are intentionally omitted.

### Input

Provide at least one search query or public YouTube/YouTube Music watch, playlist, or browse URL.

| Field | Default | Description |
|---|---:|---|
| `searchQueries` | — | Song/video search terms. |
| `startUrls` | — | Public watch, playlist, or browse URLs. |
| `maxItems` | 20 | Maximum unique records, 1–100. |
| `maxPages` | 3 | Maximum public search pages per source, 1–10. |
| `maxRetries` | 1 | Bounded retries per public request, 0–3. |
| `requestTimeoutSecs` | 60 | Per-request timeout, 15–180 seconds. |
| `requestDelayMs` | 250 | Pacing delay between requests, 0–5000 ms. |
| `includePlayerDetails` | `true` | Attempt optional public watch-page enrichment. |
| `includeDiagnostics` | `true` | Write bounded four-field diagnostics. |
| `enableProxyFallback` | `true` | Try one configured Apify Proxy after a direct request fails. |
| `proxyConfiguration` | — | Optional account-authorized Apify Proxy or credential-free HTTP/SOCKS URLs. |
| `countryCode` | `US` | Public-page country context. |
| `languageCode` | `en` | Public-page language context. |

Example:

```json
{
  "searchQueries": ["Adele Hello"],
  "startUrls": ["https://music.youtube.com/watch?v=rYEDA3JcQqw"],
  "maxItems": 10,
  "maxPages": 2,
  "includePlayerDetails": true,
  "proxyConfiguration": { "useApifyProxy": false }
}
```

### Reliability and verification

Requests use fixed ordinary public-page headers, bounded retries, pacing, and optional configured Apify Proxy routing. There is no CAPTCHA/login bypass, request-identity spoofing, browser automation, or credential storage. Optional player enrichment failures preserve the base song record and emit a bounded diagnostic. Continuation markers are reported, but opaque continuation commands are not replayed.

Verify locally with `npm test`, `npx apify validate-schema`, `npx apify run --purge --input-file INPUT.json`, and `node validate-datasets.js`. Deploy with `npx apify actors push <actor-id> --version 2.0 --build-tag latest`, then call the cloud Actor with the bounded QA inputs and compare IDs, required fields, duplicates, and field coverage.

# Changelog

This Actor's version history is a separate document: https://apify.com/w3crawler/youtube-music-song-scraper/changelog.md

# Actor input Schema

## `searchQueries` (type: `array`):

Song or music-video search terms parsed from public YouTube results pages.

## `startUrls` (type: `array`):

Public YouTube or YouTube Music watch, playlist, or browse URLs.

## `maxItems` (type: `integer`):

Maximum unique song records saved across the run (1–100).

## `maxPages` (type: `integer`):

Bounded public search-page depth per query (1–10).

## `maxRetries` (type: `integer`):

Bounded retries after a public page request fails (0–3).

## `requestTimeoutSecs` (type: `integer`):

Timeout for each public page request (15–180 seconds).

## `requestDelayMs` (type: `integer`):

Bounded pacing delay between public page requests (0–5000 ms).

## `countryCode` (type: `string`):

Two-letter country context sent to public pages.

## `languageCode` (type: `string`):

Language context sent to public pages.

## `includePlayerDetails` (type: `boolean`):

Attempt optional public watch-page metadata for each song.

## `includeDiagnostics` (type: `boolean`):

Write bounded four-field access or enrichment diagnostics.

## `enableProxyFallback` (type: `boolean`):

Try one configured Apify Proxy request after a direct public-page failure.

## `proxyConfiguration` (type: `object`):

Optional account-authorized Apify Proxy or credential-free HTTP/SOCKS URLs.

## Actor input object example

```json
{
  "searchQueries": [
    "Adele Hello"
  ],
  "maxItems": 20,
  "maxPages": 3,
  "maxRetries": 1,
  "requestTimeoutSecs": 60,
  "requestDelayMs": 250,
  "countryCode": "US",
  "languageCode": "en",
  "includePlayerDetails": true,
  "includeDiagnostics": true,
  "enableProxyFallback": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `keyValueStore` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "Adele Hello"
    ],
    "maxItems": 20,
    "maxPages": 3,
    "maxRetries": 1,
    "requestTimeoutSecs": 60,
    "requestDelayMs": 250,
    "countryCode": "US",
    "languageCode": "en",
    "includePlayerDetails": true,
    "includeDiagnostics": true,
    "enableProxyFallback": true,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("w3crawler/youtube-music-song-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["Adele Hello"],
    "maxItems": 20,
    "maxPages": 3,
    "maxRetries": 1,
    "requestTimeoutSecs": 60,
    "requestDelayMs": 250,
    "countryCode": "US",
    "languageCode": "en",
    "includePlayerDetails": True,
    "includeDiagnostics": True,
    "enableProxyFallback": True,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("w3crawler/youtube-music-song-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "Adele Hello"
  ],
  "maxItems": 20,
  "maxPages": 3,
  "maxRetries": 1,
  "requestTimeoutSecs": 60,
  "requestDelayMs": 250,
  "countryCode": "US",
  "languageCode": "en",
  "includePlayerDetails": true,
  "includeDiagnostics": true,
  "enableProxyFallback": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call w3crawler/youtube-music-song-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,w3crawler/youtube-music-song-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SrgSsGZleBrgdvgFj/builds/0nDyCCHD9ZPxTwf32/openapi.json
