Podcast RSS Scraper & Transcriber avatar

Podcast RSS Scraper & Transcriber

Pricing

from $5.00 / 1,000 episode transcripts

Go to Apify Store
Podcast RSS Scraper & Transcriber

Podcast RSS Scraper & Transcriber

Scrape any podcast RSS feed for episodes with metadata, durations and chapters. Get FREE publisher transcripts (Podcasting 2.0) or fall back to Whisper with your own key. Bulk-import subscriptions via OPML, filter by date/duration/title, and monitor shows for new episodes.

Pricing

from $5.00 / 1,000 episode transcripts

Rating

0.0

(0)

Developer

XiaoZhi DataTools

XiaoZhi DataTools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Podcast RSS Scraper & Transcriber — Episodes, Free Transcripts, SRT to Text

Scrape any podcast from its RSS feed or discover shows via Apple Podcasts search — titles, descriptions, dates, durations, audio URLs, chapters, and transcripts. Pure HTTP: no browser, no proxies, no CAPTCHAs. Fast and cheap.

Why this Actor

  • Free transcripts first. Many shows publish official transcripts with every episode (Podcasting 2.0 <podcast:transcript>). This Actor fetches them at zero cost — no API key, no per-minute ASR charges. Plain text, SRT and VTT are preserved as published.
  • Whisper fallback. Episodes without a publisher transcript can be transcribed with the OpenAI Whisper API using your own key — you pay OpenAI directly, nothing is marked up here.
  • Rich metadata. Show author, categories, language, explicit flag, artwork; per-episode season/episode numbers, duration, GUID, chapters URL.
  • Filters. Date range, min/max duration, and keyword matching on episode titles.
  • Monitor mode. Run on a schedule with newEpisodesOnly and get just the fresh episodes since the last run — perfect for tracking shows.
  • OPML bulk import. Point it at podcast-app subscription exports and scrape dozens of shows in one run.
  • Transcript formats & diagnostics. Choose text/srt/vtt output per episode; transcriptStatus tells you exactly why an episode has no transcript.
  • Fast. Concurrent fetching with no headless browser in the loop.

Input

FieldWhat it does
feedUrlsDirect RSS feed URLs, one per line
opmlUrlsOPML subscription lists (URLs or pasted OPML XML) to bulk-import feeds — monitor many shows at once
searchTermsKeywords resolved to show feeds via the Apple Podcasts (iTunes) Search API — no key needed
countryiTunes storefront for search (US, GB, DE, JP…)
maxResultsPerSearch / maxEpisodesPerFeedCaps per search term and per feed
publishedAfter / publishedBeforeDate window, YYYY-MM-DD
minDurationSeconds / maxDurationSecondsSkip trailers, teasers, or over-long episodes
titleContainsKeyword filter on titles, `
usePublisherTranscriptsFetch free official transcripts (default: on)
transcriptFormatsWhich transcript formats to emit per episode: text, srt, vtt (default: text + srt)
transcriptLanguageISO-639-1 hint (e.g. en) passed to Whisper for faster, more accurate transcription
transcribe + openaiApiKeyWhisper fallback for episodes without a publisher transcript
newEpisodesOnlyMonitor mode — only episodes unseen by previous runs
requestConcurrencyParallel fetches (default 5)
{
"searchTerms": ["artificial intelligence"],
"maxResultsPerSearch": 5,
"maxEpisodesPerFeed": 20,
"publishedAfter": "2025-01-01",
"minDurationSeconds": 300,
"usePublisherTranscripts": true,
"transcribe": false
}

Output

One dataset item per episode:

  • podcastTitle, podcastAuthor, podcastCategories, podcastLanguage, podcastExplicit, podcastFeedUrl, podcastLink, podcastImage
  • episodeTitle, episodeDescription, publishedAt, durationSeconds, seasonNumber, episodeNumber, audioUrl, episodeUrl, episodeGuid, episodeImage, chaptersUrl
  • transcriptUrl, transcriptType, transcriptSource (publisher | whisper), transcriptText, transcriptSrt, transcriptVtt
  • transcriptStatus — per-episode transcript diagnostics: found, no_publisher_tag, download_failed, empty_transcript, format_fallback, whisper_failed, not_attempted. The run log also prints a status breakdown so you can see transcript coverage at a glance.

OPML bulk import

Podcast apps (Overcast, Pocket Casts, Apple Podcasts…) export subscriptions as OPML. Paste OPML URLs — or raw OPML XML — into opmlUrls to scrape all subscribed shows in one run. Combined with newEpisodesOnly on a schedule, this becomes a full monitoring pipeline for dozens of shows.

Cost

  • Metadata-only and publisher-transcript runs cost almost nothing — a few HTTP requests per feed, no proxy spend.
  • Whisper transcription downloads at most ~1 MB/minute of audio and caps uploads at 20 MB per episode (Whisper's limit is 25 MB). maxTranscribeEpisodes bounds paid transcriptions per run. You pay OpenAI directly for Whisper usage.

Use cases

  • Build podcast search / recommendation datasets
  • Content research: mine transcripts for topics, quotes, and trends
  • Monitor shows for new episodes on a schedule
  • Feed RAG pipelines and AI agents with episode text
  • Lead generation: find shows and hosts in any niche

Run locally

pip install -r requirements.txt
apify run

Changelog

  • 0.6 — Fixed publishedBefore excluding the whole end day ("on or before this date" now really includes that day); publisher transcript downloads are throttled by requestConcurrency instead of all firing at once; Whisper fallback now transcribes concurrently; fixed invalid SRT timestamps like 00:00:05,1000 when JSON transcript segments round up to a full second; invalid publishedAfter/publishedBefore values now log a warning instead of being silently ignored; hardened HTTP client setup against malformed NO_PROXY environment values.
  • 0.5 — Added output and dataset schemas for the Store listing.