# Spotify Podcast Episode Extractor (`fanndev/spotify-podcast-episode-extractor`) Actor

Export every episode of any Spotify podcast with title, full description, duration, release date and share link - plus chapter timestamps, sponsor domains and likely guest names mined out of the descriptions. Includes show rating and review count. No Spotify account or API key needed.

- **URL**: https://apify.com/fanndev/spotify-podcast-episode-extractor.md
- **Developed by:** [Faisal Ahdan naufal](https://apify.com/fanndev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Spotify Podcast Episode Extractor

Export every episode of any Spotify podcast — title, **full description**,
duration, release date, share link — and then get the things the description
actually contains: **chapter timestamps**, **sponsor domains**, and the **guest
names** in the title.

Built for media researchers, PR teams hunting interview slots, and content
analysts tracking what podcasts are covering. A podcast's guest list, chapter
breakdown and advertiser roster are all buried in one free-text blob; this
turns that blob into columns.

No Spotify account, no API key, no browser.

***

### What you get per episode

| Field | Notes |
| --- | --- |
| `episodeName`, **`description`**, `releaseDate`, `releasedAt` | Spotify publishes release times to the minute |
| `duration`, `durationMinutes` | `1:12:45` and `73` |
| **`guestHints`** | Names read out of the title — a heuristic, honestly named |
| **`chapters`**, `chapterCount` | Every `12:34 Topic` line, with second offsets |
| **`sponsorDomains`** | Advertiser domains, after stripping the show's own links |
| `links`, `mentionedHandles` | Everything else the description points at |
| `hasVideo`, `hasTranscript`, `isPaywalled`, `explicit` | What kind of episode it is |
| `episodeUrl`, `shareUrl`, `coverArt`, `audioPreviewUrl` | Links and media |

Plus one **`SHOW`** row per podcast: publisher, description, episode count,
**average rating and review count**, topics and cover art.

***

### Quick start

```json
{
  "showUrls": ["https://open.spotify.com/show/4rOoJ6Egrf8K2IrywzwOMk"],
  "showNames": ["Huberman Lab"],
  "releasedAfter": "2026-01-01",
  "maxEpisodesPerShow": 500,
  "exportFormats": ["csv"]
}
```

Podcasts can be given as links **or** plain names. Individual episodes go in
`episodeUrls`.

***

### The date window is also a cost control

Episodes arrive newest first, so `releasedAfter` does not just filter — it
**stops paging** the moment the catalogue goes older than your date. Pulling
September's episodes from a 2 753-episode show costs three requests, not
twenty-eight.

***

### About the mined fields

`mineDescriptions` costs **no extra requests**; it reads text already in the
response. What it does is labelled honestly:

- **`chapters`** — timestamp lines in the description. A timestamp later than
  the episode's own duration is discarded, because it is a date or a price that
  looks like one.
- **`sponsorDomains`** — link domains with the show's own distribution and
  social platforms removed. In an ad-supported podcast these are the
  advertisers; in one that is not, they are just links.
- **`guestHints`** / **`episodeNumberHint`** — derived from the title. Nothing
  on Spotify marks who a guest is, so these are guesses from title convention.
  They are named `…Hints` for that reason and should not be treated as a field
  Spotify publishes.

***

### Cost

One request for the show, plus one per 100 episodes. A 100-episode pull is
three requests including the token. Looking a show up by name adds one; each
individual episode URL costs one.

***

### Limits worth knowing

- **Transcript text is not public.** `hasTranscript` tells you Spotify has one;
  the words are not exposed to an unauthenticated client and are not fetched.
- **Paywalled episodes appear in the list** with `isPaywalled: true`, but their
  audio is not accessible.
- **`episodeCount` comes from the episode query, not the show query** —
  Spotify's show metadata returns a single sample episode and no total.
- **Descriptions vary enormously.** Some publishers write full chapter lists;
  others leave the description empty. `descriptionLength` and `chapterCount`
  make that visible rather than silently returning nothing.

***

### Output

Every run writes to the Apify dataset. Set `exportFormats` to also drop a
ready-made `spotify-podcast-episodes.csv`, `.xlsx`, `.json` or `.ndjson` into
the run's key-value store. Chapters and links are nested, so they are kept out
of the spreadsheet columns and stay in the JSON exports.

See [CRAWLING\_METHOD.md](CRAWLING_METHOD.md) for how the data is obtained.

# Actor input Schema

## `showUrls` (type: `array`):

Spotify shows to read. Share links, /intl-xx/ localised links, spotify:show: URIs and bare show ids all work.

## `showNames` (type: `array`):

Podcasts to look up by name instead of by link, one per line. Each name costs one extra search request and the run log prints which show it matched, so check the log when a name is ambiguous.

## `episodeUrls` (type: `array`):

Specific episodes to read, for when you already know which ones you want. One request each.

## `startUrls` (type: `array`):

Show and episode links mixed together, for callers that keep one list.

## `maxEpisodesPerShow` (type: `integer`):

Stop after this many episodes per show. Spotify serves 100 episodes per request, so a 500-episode back catalogue costs five.

## `releasedAfter` (type: `string`):

YYYY-MM-DD. Episodes arrive newest first, so this also stops paging as soon as the catalogue goes older than the date - reading recent episodes of a long-running show costs a couple of requests rather than hundreds.

## `releasedBefore` (type: `string`):

YYYY-MM-DD. Combine with releasedAfter to pull one month, quarter or season.

## `mineDescriptions` (type: `boolean`):

Parse each description for chapter timestamps, links, sponsor domains and @handles, and derive likely guest names from the episode title. Costs no extra requests - it reads text that is already in the response. Guest names are labelled guestHints because nothing on Spotify marks who the guest is.

## `includeShowRecord` (type: `boolean`):

Emit one SHOW record per podcast with publisher, description, episode count, average rating, review count and cover art, alongside the EPISODE records.

## `exportFormats` (type: `array`):

Besides the dataset, write ready-made files into this run's key-value store.

## `proxyConfiguration` (type: `object`):

Off by default and genuinely optional: the endpoint this actor reads has no WAF and no IP block. Turn it on only for very large back-catalogue sweeps.

## Actor input object example

```json
{
  "showUrls": [
    "https://open.spotify.com/show/4rOoJ6Egrf8K2IrywzwOMk"
  ],
  "maxEpisodesPerShow": 100,
  "mineDescriptions": true,
  "includeShowRecord": true,
  "exportFormats": [],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Every episode, each show summary and any error records from this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "showUrls": [
        "https://open.spotify.com/show/4rOoJ6Egrf8K2IrywzwOMk"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fanndev/spotify-podcast-episode-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "showUrls": ["https://open.spotify.com/show/4rOoJ6Egrf8K2IrywzwOMk"] }

# Run the Actor and wait for it to finish
run = client.actor("fanndev/spotify-podcast-episode-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "showUrls": [
    "https://open.spotify.com/show/4rOoJ6Egrf8K2IrywzwOMk"
  ]
}' |
apify call fanndev/spotify-podcast-episode-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fanndev/spotify-podcast-episode-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1W7mgFIlhWtFamjiy/builds/2F1R5bipKO371Lah6/openapi.json
