# Podcast Transcripts from RSS: Text, Markdown & RAG Chunks (`quietbyte/podcast-transcripts`) Actor

Get publisher-provided podcast transcripts straight from RSS feeds (Podcasting 2.0 transcript tags): VTT, SRT, JSON, HTML or text turned into clean text, timestamped markdown and RAG-ready chunks. No audio, no AI guesses. Pay only for transcripts delivered; monitoring mode returns new episodes only.

- **URL**: https://apify.com/quietbyte/podcast-transcripts.md
- **Developed by:** [Quietbyte](https://apify.com/quietbyte) (community)
- **Categories:** AI, Videos, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$4.00 / 1,000 podcast transcripts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Podcast Transcripts from RSS: text, markdown and RAG chunks

Get the transcripts that podcast publishers already put in their RSS feeds, cleaned up and ready for search, notes, research or an LLM. No audio is downloaded and nothing is guessed by AI: you get the publisher's own transcript, converted to plain text, timestamped markdown and RAG-sized chunks.

**The question it answers: "Which of these shows publish a transcript, and can I have it as clean text?"**

Built for people who build search, summaries, knowledge bases, research tools and AI agents on top of podcasts.

### How it works

Many podcast hosts (Buzzsprout, Transistor, Omny, Acast, Podbean and others) add a `<podcast:transcript>` link to each episode in the RSS feed, following the open [Podcasting 2.0 namespace](https://github.com/Podcastindex-org/podcast-namespace/blob/main/docs/1.0.md). The file behind the link can be JSON, WebVTT, SRT, HTML or plain text. This Actor reads the feed, picks the best format for each episode (JSON, then VTT, SRT, HTML, text), parses it, and returns one clean record per episode.

Not every show publishes a transcript. In a sample of Apple's US top 100 feeds taken on 2026-10-09, **10 of 98 reachable feeds** included transcript links. Episodes without a transcript link are skipped and **cost nothing**.

### What you can put in

| Field | What it does |
|---|---|
| `feedUrls` | Podcast RSS feed URLs (the feed address, not the Apple or Spotify page). Up to 200. |
| `maxTranscriptsPerFeed` | The newest N episodes that have a transcript, per feed. Default 5. |
| `publishedWithinDays` | Only episodes from the last N days. |
| `titleKeywords` | Keep episodes whose title has any of these words. |
| `preferredLanguage` | Try this transcript language first when a feed offers several (for example `en`, `pt-BR`). |
| `includeMarkdown` / `includeChunks` / `includeSegments` | Choose the output shapes. Chunks and markdown are on by default. |
| `chunkWords` | Words per chunk, 50 to 2000 (default 250). |
| `maxResults` | Stop after N transcripts in total. |
| `onlyNewEpisodes` + `monitorName` | Monitoring: return only episodes not returned before. |

```json
{
  "feedUrls": ["https://rss.buzzsprout.com/2632213.rss"],
  "maxTranscriptsPerFeed": 3,
  "preferredLanguage": "en"
}
```

### What you get

One record per episode. A real (shortened) record:

```json
{
  "id": "5ce7e6f008c70017",
  "feedUrl": "https://rss.buzzsprout.com/2632213.rss",
  "podcastTitle": "On The Line with Jim Pyne",
  "language": "en-us",
  "episodeTitle": "Frank Beamer: Building Virginia Tech, Beamer Ball & a Legacy of Leadership | On the Line with Jim Pyne",
  "publishedAt": "2026-10-06T11:00:00Z",
  "durationSeconds": 1858,
  "transcriptUrl": "https://www.buzzsprout.com/2632213/19918175/transcript.json",
  "transcriptFormat": "json",
  "hasTimestamps": true,
  "hasSpeakers": true,
  "wordCount": 6112,
  "text": "SPEAKER_00: Welcome to On the Line with Jim Pine. I'm Jim Pine. …",
  "markdown": "[00:00:00] **SPEAKER_00** Welcome to On the Line with Jim Pine. …",
  "chunks": [{"index": 0, "start": 0.4, "end": 76.32, "startTimestamp": "00:00:00",
              "speakers": ["SPEAKER_00", "SPEAKER_01"], "wordCount": 241, "text": "Welcome to On the Line …"}],
  "attribution": "Transcript supplied by the podcast publisher in its RSS feed (…). The show owns it. …",
  "scrapedAt": "2026-10-09T03:58:49Z"
}
```

Field notes:

- `transcriptFormat` is the file type the publisher provided. `hasTimestamps` is false for HTML and plain-text transcripts, which carry no timing.
- `hasSpeakers` is true when the file names speakers (WebVTT voice tags, JSON `speaker`, or `Name:` lines). Speaker names are whatever the publisher wrote, such as `SPEAKER_00`.
- Word-by-word JSON transcripts are joined into sentence-sized segments.
- `episodeUrl` is the episode page when the feed lists one, otherwise `null`; `audioUrl` and `podcastUrl` are always included when the feed has them.
- The run summary (key-value store `SUMMARY`) lists, per feed, how many episodes were seen, how many carried a transcript link, how many were delivered and why a feed failed.

### Daily monitoring

Turn on `onlyNewEpisodes`, give it a `monitorName` and schedule the Actor. Each run looks at the newest `maxTranscriptsPerFeed` transcript episodes of each feed and returns only those you haven't received before, so you pay only for new transcripts. It does not walk back into the archive.

### Pricing

**$0.004 per transcript delivered** (about $4 per 1,000; event `transcript`). Everything is included: text, markdown, chunks and segments. Episodes without a transcript, failed downloads and already-seen episodes cost nothing. The run stops cleanly at your spending limit.

### Using these transcripts

The transcript belongs to the show. If you republish it, credit the show and link to the episode (the `attribution` field says so). Please check the show's own terms before reusing transcripts commercially.

### Data sources and compliance

- Public podcast RSS feeds and the transcript files the publisher links from them. No logins, no proxies, no browser.
- No audio is downloaded or transcribed, and nothing is taken from YouTube or other platforms that forbid automated access.
- Emails that appear inside transcripts are replaced with `[email removed]`.
- The Actor identifies itself with a clear User-Agent, sends at most one request per second to a host, refuses non-public addresses, and caps feed and transcript size.

### Limits

- Feeds are read up to the newest 1,000 episodes and up to 40 MB.
- Transcript files above 2 MB are skipped.
- Only the Podcasting 2.0 `podcast:transcript` tag is read. Shows that publish transcripts only on their website, or not at all, return nothing.

### FAQ

**Can it transcribe audio when there is no transcript?** Not in this version. It returns only what the publisher provides, which keeps it cheap and accurate.

**How do I find a show's RSS feed?** Most hosts show it as "RSS feed" on the show page. Directories such as Podcast Index list feed URLs.

**Why did my feed return nothing?** Check the `SUMMARY` record. If `withTranscriptTag` is 0, the show doesn't publish transcripts in its feed.

**Can I use it from AI agents or MCP?** Yes. Apify exposes every Actor through its API and MCP server, and the output is plain JSON.

### Related Actors

More tools from the same developer:

- [ATS Jobs API with Visa Sponsorship Tags](https://apify.com/quietbyte/ats-jobs-visa): Greenhouse, Lever, Ashby, SmartRecruiters and Recruitee boards with sentence-level visa tags.
- [Workday Jobs API with H-1B & UK Visa Sponsor Evidence](https://apify.com/quietbyte/workday-jobs-visa): any Workday careers site, with visa tags backed by official H-1B and UK sponsor records.
- [Remote Jobs by Country](https://apify.com/quietbyte/remote-jobs-by-country): four remote boards in one feed, filtered to jobs open to your country.

# Actor input Schema

## `feedUrls` (type: `array`):

RSS feed URLs of the podcasts (the feed address, not the Apple or Spotify page). Only episodes whose feed carries a transcript link are returned. Up to 200 feeds.

## `maxTranscriptsPerFeed` (type: `integer`):

The newest N episodes that have a transcript, per feed.

## `publishedWithinDays` (type: `integer`):

Only episodes published in the last N days. 0 means any date.

## `titleKeywords` (type: `array`):

Keep episodes whose title contains any of these words or phrases (whole words, case-insensitive). Empty keeps all.

## `preferredLanguage` (type: `string`):

When a feed offers transcripts in several languages, try this one first, e.g. en or pt-BR. Leave empty to use the best format.

## `includeMarkdown` (type: `boolean`):

Timestamped markdown with speaker names, ready for notes or LLM prompts.

## `includeChunks` (type: `boolean`):

Chunks of about the chunk size below, split on sentence-turn edges, each with its start time and speakers.

## `chunkWords` (type: `integer`):

Target words per chunk (50 to 2000). About 0.75 words per token.

## `includeSegments` (type: `boolean`):

The transcript as timed segments (start, end, speaker, text). Larger output.

## `maxResults` (type: `integer`):

Stop after this many transcripts in total. 0 means no limit.

## `onlyNewEpisodes` (type: `boolean`):

Return only episodes not returned by earlier runs with the same monitor name. Schedule the Actor daily and pay only for new transcripts.

## `monitorName` (type: `string`):

Name that groups runs which share the same memory of seen episodes. Leave empty to derive it from the feeds.

## Actor input object example

```json
{
  "feedUrls": [
    "https://rss.buzzsprout.com/2632213.rss"
  ],
  "maxTranscriptsPerFeed": 3,
  "publishedWithinDays": 0,
  "titleKeywords": [],
  "preferredLanguage": "",
  "includeMarkdown": true,
  "includeChunks": true,
  "chunkWords": 250,
  "includeSegments": false,
  "maxResults": 0,
  "onlyNewEpisodes": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feedUrls": [
        "https://rss.buzzsprout.com/2632213.rss"
    ],
    "maxTranscriptsPerFeed": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("quietbyte/podcast-transcripts").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "feedUrls": ["https://rss.buzzsprout.com/2632213.rss"],
    "maxTranscriptsPerFeed": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("quietbyte/podcast-transcripts").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feedUrls": [
    "https://rss.buzzsprout.com/2632213.rss"
  ],
  "maxTranscriptsPerFeed": 3
}' |
apify call quietbyte/podcast-transcripts --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,quietbyte/podcast-transcripts"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/faub0zsabvrlP6Tl8/builds/Cteo2wYZh5Tn8ccFV/openapi.json
