# Spotify Scraper - Artists, Playlists & Podcasts (`valev-lab/spotify-scraper`) Actor

Scrape Spotify artists, tracks, albums, playlists, and podcasts from URLs or search. Get monthly listeners, play counts, discography, and change monitoring — no API key or login.

- **URL**: https://apify.com/valev-lab/spotify-scraper.md
- **Developed by:** [Daniel Valev](https://apify.com/valev-lab) (community)
- **Categories:** Social media, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.60 / 1,000 item scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Spotify Scraper do?

**Spotify Scraper** extracts **public** catalog data from [Spotify](https://open.spotify.com) — artists, albums, tracks, playlists, podcast shows, and episodes — from **URLs** or **keyword search**. It returns structured JSON with **monthly listeners**, **play counts**, **followers**, **world rank**, biographies, discography, and playlist tracks **beyond the usual 100-item cap**.

No Spotify API key, no login, and no `sp_dc` cookies. Anonymous web-player GraphQL over HTTP; Playwright starts only if auth or persisted-query hashes need a one-shot browser bootstrap.

Use it as a **Spotify API alternative** for metrics the official Web API does not expose (monthly listeners, world rank, stream play counts).

### Why scrape Spotify?

Spotify is the largest music streaming catalog many teams track for:

- **Music analytics** — monthly listeners, play counts, follower trends
- **Playlist research** — editorial and user playlists past 100 tracks
- **Artist monitoring** — world rank, top tracks, related artists, scheduled deltas
- **Podcast discovery** — search **shows** and **episodes** by keyword (not URL-only)
- **Product / RAG datasets** — clean JSON for apps, agents, and MCP workflows

Apify schedules, webhooks, dataset exports (JSON/CSV/Excel), the REST API, and MCP make recurring runs straightforward.

### What data can Spotify Scraper extract?

| Type | Key fields |
| --- | --- |
| **Artist** | `monthlyListeners`, `followers`, `worldRank`, `topCities`, biography, top tracks with play counts, related artists, discography counts |
| **Album** | track list with **per-track play counts**, label, release date, `totalPlayCount` |
| **Track** | `playCount`, artists, album art, duration, preview URL when Spotify exposes one |
| **Playlist** | owner, followers/saves, description, **paginated** `tracks` up to `maxPlaylistTracks` (use `0` for all) |
| **Show** | publisher, description, topics, episodes (up to `maxEpisodesPerShow`) |
| **Episode** | parent show link, release date, duration, explicit/playable/video flags, preview URL |
| **Search result** | typed hit (`resultType`) from tracks / artists / albums / playlists / **shows** / **episodes** |

Optional `monitoringKey` writes snapshots to a named key-value store (`spotify-monitor-<key>`) and adds `previousSnapshotAt` + `deltas` on later runs — ideal for **Schedules**.

### How to scrape Spotify

1. Open **Spotify Scraper** in Apify Console.
2. Choose **Mode**: `urls` or `search`.
3. **URLs** — paste `open.spotify.com` links (incl. `/intl-xx/`), `spotify:type:id` URIs, or `spotify.link` short links.
4. **Search** — add `searchTerms`, set `searchType` (tracks, artists, albums, playlists, shows, episodes), and `maxResults`.
5. Optional: enable `includeDiscography`, raise `maxPlaylistTracks`, or set `monitoringKey` for change tracking.
6. Keep Apify Proxy enabled (automatic group; the Actor can escalate to RESIDENTIAL on blocks).
7. Click **Start**, then download the **Dataset** (JSON, CSV, Excel, …).

#### URL mode example

```json
{
  "mode": "urls",
  "urls": [
    "https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02",
    "https://open.spotify.com/playlist/37i9dQZF1DXcBWIGoYBM5M",
    "spotify:track:5XeFesFbtLpXzIVDNQP22n",
    "https://open.spotify.com/show/2MAi0BvDc6GTFvKFPXnkCL"
  ],
  "maxPlaylistTracks": 500
}
```

#### Search mode example (podcasts)

```json
{
  "mode": "search",
  "searchTerms": ["true crime", "lex fridman"],
  "searchType": "shows",
  "maxResults": 20
}
```

#### Monitor artist metrics over time

```json
{
  "mode": "urls",
  "urls": ["https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02"],
  "monitoringKey": "taylor-weekly"
}
```

Schedule the same input daily. Later runs attach `deltas` (e.g. monthlyListeners, followers, playlist track adds/removes).

#### Full discography (charged per album)

```json
{
  "mode": "urls",
  "urls": ["https://open.spotify.com/artist/7Ln80lUS6He07XvHI8qqHH"],
  "includeDiscography": true,
  "discographyTypes": ["albums"]
}
```

Each album record is a separate billed `item`. Prefer a narrow `discographyTypes` list for large catalogs.

### How much does it cost to scrape Spotify?

Pay-per-event pricing. You pay for events, not a separate compute invoice for this Actor’s PPE setup.

| Event | Free plan | Description |
| --- | ---: | --- |
| **Actor start** (`apify-actor-start`) | $0.00005 | Synthetic start (memory-scaled); keeps startup compute competitive |
| **Item scraped** (`item`) | **$0.004** | One top-level dataset record |

Volume discounts (USD per `item`):

| Plan | Per item | Per 1,000 items |
| --- | ---: | ---: |
| Free | $0.0040 | $4.00 |
| Bronze | $0.0035 | $3.50 |
| Silver | $0.0030 | $3.00 |
| Gold / Platinum / Diamond | $0.0026 | $2.60 |

**Not charged:** nested arrays (tracks inside an album, cities on an artist), failed URLs, empty search shelves.

#### Cost examples (Free tier, ignore tiny start fee)

| Scenario | Items | ≈ cost |
| --- | ---: | ---: |
| Health-check default (5 URLs) | 5 | ~$0.02 |
| Search 20 shows | 20 | ~$0.08 |
| 1 artist + discography of 8 albums | 9 | ~$0.036 |
| Playlist capped at 120 tracks (1 record) | 1 | ~$0.004 |

Set **Maximum charge per run** in Console if you want a hard stop.

### Input

See the **Input** tab for the full schema. Important fields:

| Field | Notes |
| --- | --- |
| `mode` | `urls` or `search` |
| `urls` | Artist, album, track, playlist, show, episode — open.spotify, URI, or spotify.link |
| `searchTerms` / `searchType` / `maxResults` | Search shelf; shows & episodes supported |
| `enrichSearchResults` | Fetch full details per hit (each hit still + each enrichment = charged items) |
| `maxPlaylistTracks` | Cap playlist pagination; `0` = all |
| `includeDiscography` / `discographyTypes` | Expand artist releases as album items |
| `maxEpisodesPerShow` | Episodes attached on show records; `0` skips listing |
| `monitoringKey` | Named KV snapshots + `deltas` |
| `proxyConfiguration` | Prefer Apify Proxy (automatic) |

Empty input runs a deterministic **5-URL health check** (artist, track, album, show, episode).

### Output

Results go to the default dataset. Download JSON, CSV, Excel, HTML, XML, or RSS, or read via the dataset API / MCP.

#### Artist example

```json
{
  "type": "artist",
  "id": "06HL4z0CvFAxyc27GXpf02",
  "uri": "spotify:artist:06HL4z0CvFAxyc27GXpf02",
  "url": "https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02",
  "name": "Taylor Swift",
  "monthlyListeners": 102570645,
  "followers": 140000000,
  "worldRank": 1,
  "topCities": [{ "city": "Jakarta", "country": "ID", "listeners": 1200000 }],
  "topTracks": [{ "id": "1BxfuPKGuaTgP7aM0Bbdwr", "name": "Cruel Summer", "playCount": 3600000000 }],
  "scrapedAt": "2026-09-28T09:16:00.000Z",
  "dataSource": "graphql",
  "isPartial": false
}
```

#### Track example

```json
{
  "type": "track",
  "id": "5XeFesFbtLpXzIVDNQP22n",
  "name": "I Wanna Be Yours",
  "playCount": 2500000000,
  "durationMs": 183066,
  "durationFormatted": "3:03",
  "artists": [{ "id": "7Ln80lUS6He07XvHI8qqHH", "name": "Arctic Monkeys" }],
  "url": "https://open.spotify.com/track/5XeFesFbtLpXzIVDNQP22n",
  "scrapedAt": "2026-09-28T09:16:00.000Z",
  "dataSource": "graphql"
}
```

#### Search / monitoring notes

- Search hits use `type: "searchResult"` with `resultType` and `searchTerm`.
- With `monitoringKey`, numeric entities may include `previousSnapshotAt` and `deltas`.

Each run also writes **`RUN_STATS`** to the default key-value store (`authPath`, request/byte counts, failures) — not billed.

### Tips

- Prefer **URL mode** when you need monthly listeners or play counts; plain search hits are lighter metadata unless `enrichSearchResults` is on.
- Cap **`maxPlaylistTracks`** and **`discographyTypes`** to control cost on huge catalogs.
- Use **`monitoringKey` + Schedule** instead of downloading full history every day — deltas are cheaper to reason about.
- Mix artists, playlists, and podcast URLs in one run when useful.

### Limitations

- **Public anonymous data only** — no private playlists, no user library, no login cookies, no full audio download, no lyrics.
- Market-restricted or removed catalog entries may be empty or partial (`isPartial` / embed fallback).
- Discography can emit **many** album items (each charged).
- Official Spotify Web API OAuth apps are out of scope; this Actor targets public web-player metrics.

### FAQ

#### Do I need a Spotify developer app or API key?

No. The Actor uses anonymous web-player access — the same class of public metadata a logged-out browser can see.

#### Can I get Spotify monthly listeners and play counts?

Yes for **artists** (monthly listeners, world rank) and for **tracks/albums** when scraped in URL mode (and on discography albums). Search-only hits may omit play counts unless you enrich them.

#### Can I scrape Spotify podcasts by keyword?

Yes. Set `mode: "search"` and `searchType` to `shows` or `episodes`. Many scrapers only accept show/episode URLs.

#### How do I scrape a playlist with more than 100 tracks?

Set `maxPlaylistTracks` above 100 (or `0` for all). The Actor paginates playlist contents instead of stopping at the embed-sized window.

#### How do I track artist growth over time?

Set `monitoringKey`, schedule the run, and read `deltas` / `previousSnapshotAt` on subsequent outputs. Snapshots live in `spotify-monitor-<key>` on your account.

#### Is this legal?

The Actor only reads **public** Spotify pages/APIs available without login. You are responsible for complying with [Spotify’s Terms](https://www.spotify.com/legal/end-user-agreement/), applicable law (including GDPR if personal data appears in public bios or names), and for having a lawful basis for your use case. Do not use it to download copyrighted audio.

#### The run failed or returned partial data

Check `RUN_STATS.failures`, try Apify Proxy / RESIDENTIAL, and confirm the URL is public in a normal browser. Open an issue on the Actor’s **Issues** tab with the run ID if it persists.

### Integrations, API, and MCP

- **Schedules** — daily artist or playlist monitors
- **Webhooks / Zapier / Make / Sheets** — push new dataset items downstream
- **Apify API** — `POST` runs with the same JSON as Console
- **MCP** — connect [Apify MCP](https://mcp.apify.com) and allow `valev-lab/spotify-scraper` for agent workflows

JavaScript:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('valev-lab/spotify-scraper').call({
  mode: 'urls',
  urls: ['https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0]?.name, items[0]?.monthlyListeners);
```

Python:

```python
from apify_client import ApifyClient

client = ApifyClient()
run = client.actor("valev-lab/spotify-scraper").call(run_input={
    "mode": "search",
    "searchTerms": ["arctic monkeys"],
    "searchType": "artists",
    "maxResults": 5,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["type"], item.get("name"))
```

### Development (repository)

```bash
npm install
npm run build
npm test
cd actors/spotify-scraper && apify validate-schema && apify run --purge < /dev/null
```

# Changelog

This Actor's version history is a separate document: https://apify.com/valev-lab/spotify-scraper/changelog.md

# Actor input Schema

## `mode` (type: `string`):

Choose URLs to scrape specific Spotify links, or Search to find tracks, artists, albums, playlists, shows, or episodes by keyword.

## `urls` (type: `array`):

Spotify URLs to scrape: open.spotify.com (including /intl-xx/), spotify:type:id URIs, or spotify.link short links. Supported types: artist, album, track, playlist, show, episode.

## `searchTerms` (type: `array`):

Keywords to search on Spotify when mode is Search. Pair with searchType (tracks, artists, albums, playlists, shows, episodes).

## `searchType` (type: `string`):

Result shelf to collect in search mode. Shows and episodes are supported.

## `maxResults` (type: `integer`):

Upper bound of search hits to return per term.

## `enrichSearchResults` (type: `boolean`):

When enabled, fetch full details for each search hit. Each enriched record is charged as a normal item.

## `maxPlaylistTracks` (type: `integer`):

Maximum playlist tracks to paginate. Use 0 for all available tracks.

## `includeDiscography` (type: `boolean`):

For artists, also emit album records for every release (with per-track play counts). Each album is charged separately.

## `discographyTypes` (type: `array`):

Which release shelves to expand when includeDiscography is enabled.

## `maxEpisodesPerShow` (type: `integer`):

How many episodes to attach on show records. Use 0 to skip episode listing.

## `monitoringKey` (type: `string`):

Optional key that enables change tracking across runs via a named key-value store (spotify-monitor-<key>).

## `locale` (type: `string`):

Locale hint passed to Spotify where supported (for example en, de).

## `maxConcurrency` (type: `integer`):

Parallel entity / search-term workers.

## `proxyConfiguration` (type: `object`):

Proxy settings for outbound requests. Prefer Apify Proxy (automatic group). The Actor can escalate to RESIDENTIAL on blocks; if proxy access fails it continues without a proxy.

## Actor input object example

```json
{
  "mode": "urls",
  "urls": [
    "https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02",
    "https://open.spotify.com/track/5XeFesFbtLpXzIVDNQP22n"
  ],
  "searchType": "tracks",
  "maxResults": 50,
  "enrichSearchResults": false,
  "maxPlaylistTracks": 1000,
  "includeDiscography": false,
  "discographyTypes": [
    "albums",
    "singles",
    "compilations"
  ],
  "maxEpisodesPerShow": 50,
  "locale": "en",
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `runStats` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02",
        "https://open.spotify.com/track/5XeFesFbtLpXzIVDNQP22n"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("valev-lab/spotify-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02",
        "https://open.spotify.com/track/5XeFesFbtLpXzIVDNQP22n",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("valev-lab/spotify-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://open.spotify.com/artist/06HL4z0CvFAxyc27GXpf02",
    "https://open.spotify.com/track/5XeFesFbtLpXzIVDNQP22n"
  ]
}' |
apify call valev-lab/spotify-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,valev-lab/spotify-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HzMvAd4Q1kPftQHXl/builds/A0iccezRggvzjErxX/openapi.json
