Podcast Transcripts from RSS: Text, Markdown & RAG Chunks avatar

Podcast Transcripts from RSS: Text, Markdown & RAG Chunks

Pricing

$4.00 / 1,000 podcast transcripts

Go to Apify Store
Podcast Transcripts from RSS: Text, Markdown & RAG Chunks

Podcast Transcripts from RSS: Text, Markdown & RAG Chunks

Get publisher-provided podcast transcripts straight from RSS feeds (Podcasting 2.0 transcript tags): VTT, SRT, JSON, HTML or text turned into clean text, timestamped markdown and RAG-ready chunks. No audio, no AI guesses. Pay only for transcripts delivered; monitoring mode returns new episodes only.

Pricing

$4.00 / 1,000 podcast transcripts

Rating

0.0

(0)

Developer

Quietbyte

Quietbyte

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Podcast Transcripts from RSS: text, markdown and RAG chunks

Get the transcripts that podcast publishers already put in their RSS feeds, cleaned up and ready for search, notes, research or an LLM. No audio is downloaded and nothing is guessed by AI: you get the publisher's own transcript, converted to plain text, timestamped markdown and RAG-sized chunks.

The question it answers: "Which of these shows publish a transcript, and can I have it as clean text?"

Built for people who build search, summaries, knowledge bases, research tools and AI agents on top of podcasts.

How it works

Many podcast hosts (Buzzsprout, Transistor, Omny, Acast, Podbean and others) add a <podcast:transcript> link to each episode in the RSS feed, following the open Podcasting 2.0 namespace. The file behind the link can be JSON, WebVTT, SRT, HTML or plain text. This Actor reads the feed, picks the best format for each episode (JSON, then VTT, SRT, HTML, text), parses it, and returns one clean record per episode.

Not every show publishes a transcript. In a sample of Apple's US top 100 feeds taken on 2026-10-09, 10 of 98 reachable feeds included transcript links. Episodes without a transcript link are skipped and cost nothing.

What you can put in

FieldWhat it does
feedUrlsPodcast RSS feed URLs (the feed address, not the Apple or Spotify page). Up to 200.
maxTranscriptsPerFeedThe newest N episodes that have a transcript, per feed. Default 5.
publishedWithinDaysOnly episodes from the last N days.
titleKeywordsKeep episodes whose title has any of these words.
preferredLanguageTry this transcript language first when a feed offers several (for example en, pt-BR).
includeMarkdown / includeChunks / includeSegmentsChoose the output shapes. Chunks and markdown are on by default.
chunkWordsWords per chunk, 50 to 2000 (default 250).
maxResultsStop after N transcripts in total.
onlyNewEpisodes + monitorNameMonitoring: return only episodes not returned before.
{
"feedUrls": ["https://rss.buzzsprout.com/2632213.rss"],
"maxTranscriptsPerFeed": 3,
"preferredLanguage": "en"
}

What you get

One record per episode. A real (shortened) record:

{
"id": "5ce7e6f008c70017",
"feedUrl": "https://rss.buzzsprout.com/2632213.rss",
"podcastTitle": "On The Line with Jim Pyne",
"language": "en-us",
"episodeTitle": "Frank Beamer: Building Virginia Tech, Beamer Ball & a Legacy of Leadership | On the Line with Jim Pyne",
"publishedAt": "2026-10-06T11:00:00Z",
"durationSeconds": 1858,
"transcriptUrl": "https://www.buzzsprout.com/2632213/19918175/transcript.json",
"transcriptFormat": "json",
"hasTimestamps": true,
"hasSpeakers": true,
"wordCount": 6112,
"text": "SPEAKER_00: Welcome to On the Line with Jim Pine. I'm Jim Pine. …",
"markdown": "[00:00:00] **SPEAKER_00** Welcome to On the Line with Jim Pine. …",
"chunks": [{"index": 0, "start": 0.4, "end": 76.32, "startTimestamp": "00:00:00",
"speakers": ["SPEAKER_00", "SPEAKER_01"], "wordCount": 241, "text": "Welcome to On the Line …"}],
"attribution": "Transcript supplied by the podcast publisher in its RSS feed (…). The show owns it. …",
"scrapedAt": "2026-10-09T03:58:49Z"
}

Field notes:

  • transcriptFormat is the file type the publisher provided. hasTimestamps is false for HTML and plain-text transcripts, which carry no timing.
  • hasSpeakers is true when the file names speakers (WebVTT voice tags, JSON speaker, or Name: lines). Speaker names are whatever the publisher wrote, such as SPEAKER_00.
  • Word-by-word JSON transcripts are joined into sentence-sized segments.
  • episodeUrl is the episode page when the feed lists one, otherwise null; audioUrl and podcastUrl are always included when the feed has them.
  • The run summary (key-value store SUMMARY) lists, per feed, how many episodes were seen, how many carried a transcript link, how many were delivered and why a feed failed.

Daily monitoring

Turn on onlyNewEpisodes, give it a monitorName and schedule the Actor. Each run looks at the newest maxTranscriptsPerFeed transcript episodes of each feed and returns only those you haven't received before, so you pay only for new transcripts. It does not walk back into the archive.

Pricing

$0.004 per transcript delivered (about $4 per 1,000; event transcript). Everything is included: text, markdown, chunks and segments. Episodes without a transcript, failed downloads and already-seen episodes cost nothing. The run stops cleanly at your spending limit.

Using these transcripts

The transcript belongs to the show. If you republish it, credit the show and link to the episode (the attribution field says so). Please check the show's own terms before reusing transcripts commercially.

Data sources and compliance

  • Public podcast RSS feeds and the transcript files the publisher links from them. No logins, no proxies, no browser.
  • No audio is downloaded or transcribed, and nothing is taken from YouTube or other platforms that forbid automated access.
  • Emails that appear inside transcripts are replaced with [email removed].
  • The Actor identifies itself with a clear User-Agent, sends at most one request per second to a host, refuses non-public addresses, and caps feed and transcript size.

Limits

  • Feeds are read up to the newest 1,000 episodes and up to 40 MB.
  • Transcript files above 2 MB are skipped.
  • Only the Podcasting 2.0 podcast:transcript tag is read. Shows that publish transcripts only on their website, or not at all, return nothing.

FAQ

Can it transcribe audio when there is no transcript? Not in this version. It returns only what the publisher provides, which keeps it cheap and accurate.

How do I find a show's RSS feed? Most hosts show it as "RSS feed" on the show page. Directories such as Podcast Index list feed URLs.

Why did my feed return nothing? Check the SUMMARY record. If withTranscriptTag is 0, the show doesn't publish transcripts in its feed.

Can I use it from AI agents or MCP? Yes. Apify exposes every Actor through its API and MCP server, and the output is plain JSON.

More tools from the same developer: