Podcast Episode Scraper: RSS, Audio URLs & Transcripts
Pricing
from $0.60 / 1,000 episodes
Podcast Episode Scraper: RSS, Audio URLs & Transcripts
Scrape podcast episodes from any RSS feed or by show name: title, description, duration, audio URL and publish date, plus the full transcript for shows that publish one. No API key, pay per episode.
Pricing
from $0.60 / 1,000 episodes
Rating
0.0
(0)
Developer
The Mine Works
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors.
Give it a podcast's RSS feed, or just the show's name, and get every episode back as a clean row: title, show notes, publish date, duration in seconds, episode and season numbers, the direct audio file URL, and the full transcript text for shows that publish one. Show names are looked up in Apple's public podcast directory, so you do not need to find the feed yourself. No API key, no login and no browser.
Why choose this actor?
- A whole back catalogue in seconds. In our test of the live build on 2 October, two shows (one found by name, one by feed URL) returned 200 episodes in 24 seconds. The Vergecast's feed alone listed 1,087 episodes; you can take up to 2,000 per show.
- Real transcripts, where the publisher provides them. With transcripts on, 9 of 10 Podnews Weekly Review episodes came back with the full transcript as plain text, 38,000 to 119,000 characters each, in about 30 seconds. Every row carries
has_transcript, so you always know which episodes have one. Nothing is machine guessed. - Pay only for episodes delivered. From $0.60 per 1,000 episodes on Gold, transcript or not. A feed that will not load, a show name with no match and the run's report rows cost nothing beyond the $0.005 start fee.
Part of The Mine Works Social media and video family: Threads Scraper, Reddit Scraper, Threads Search Scraper, Instagram Profile Scraper, Instagram Followers & Following, Reddit Search Scraper.
Try it in one minute
Paste this into the JSON tab of the input page and press Start:
{"feedUrls": ["https://feeds.buzzsprout.com/1538779.rss"],"maxEpisodesPerShow": 10,"includeTranscripts": true}
You get the 10 latest Podnews Weekly Review episodes, most of them with their transcripts, in under a minute.
Give the shows in either or both of two ways: feedUrls, the RSS feed of each show (any podcast host), or searchTerms, show names that the actor looks up in Apple's podcast directory, taking up to 3 matching feeds per name. maxEpisodesPerShow caps each show (newest first in the feeds we tested), and includeTranscripts switches the transcript downloads on or off.
Apify's free plan includes $5 of credit every month, which covers about 4,800 episodes at this actor's Free plan price ($0.001 an episode plus the $0.005 start fee, in runs of 200 episodes).
Copy to your AI assistant
themineworks/podcast-episode-scraper on Apify. Reads podcast RSS feeds (given as URLs, or found by show name in Apple's podcast directory) and returns one row per episode with show title and author, episode title, show notes, publish date, duration in seconds, episode and season numbers, audio file URL, and transcript text when the feed publishes one. Call ApifyClient("TOKEN").actor("themineworks/podcast-episode-scraper").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Give feedUrls (string[] of RSS URLs) and/or searchTerms (string[] of show names; up to 3 matching feeds each). Optional: maxEpisodesPerShow (1 to 2000, default 50), includeTranscripts (default true; one extra download per episode that lists a transcript). Each row has has_transcript (true or false). Rows with _type "summary" or "info" are never billed. Full spec: GET https://api.apify.com/v2/acts/themineworks~podcast-episode-scraper/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations?fpr=ymnoit&utm_source=apify-readme&utm_medium=referral
Key features
- Up to 17 fields per episode:
show_title,show_author,feed_url,episode_title,description,guid,published_at,duration_seconds,episode_number,season,explicit,audio_url,episode_url,has_transcript,transcript_url,transcriptandscraped_at. - Any podcast host. Every public podcast has an RSS feed, and the actor reads the feed itself rather than a player app. Our October tests read feeds hosted on Buzzsprout, Megaphone and a show's own server.
- Find shows by name.
searchTermsasks Apple's public podcast directory for each name and scrapes up to 3 matching feeds; a feed found twice (by name and by URL) is read once. - Transcripts converted to text. When a feed lists a transcript file (the Podcasting 2.0
podcast:transcripttag), the actor downloads it and turns WebVTT, SRT or JSON captions into readable text, dropping cue numbers and timing lines. Files are capped at 400,000 bytes. - Clean numbers. Duration comes back in seconds whether the feed wrote
3431or00:57:11, and HTML is stripped from show notes (first 5,000 characters).
How to use it
Basic: one show by its feed
{"feedUrls": ["https://feeds.megaphone.fm/vergecast"],"maxEpisodesPerShow": 100,"includeTranscripts": false}
Metadata only: one feed download, 100 rows, a few seconds.
Find shows by name
{"searchTerms": ["Podnews Daily", "The Vergecast"],"maxEpisodesPerShow": 50}
Each name is looked up in Apple's podcast directory and up to 3 matching feeds are read. A broad name can match other shows too, so check show_title and feed_url in the results, or use feedUrls when you need exactly one show.
Transcripts for a knowledge base or RAG index
{"feedUrls": ["https://podnews.net/rss", "https://feeds.buzzsprout.com/1538779.rss"],"maxEpisodesPerShow": 200,"includeTranscripts": true}
Keep rows where has_transcript is true, chunk the transcript text, and index it with episode_title, published_at and audio_url as metadata. Transcripts keep the speaker names and timestamps the publisher wrote ("Sam Sethi 0:38 ..."), which helps citations.
New episode alerts across a set of shows
{"feedUrls": ["https://feeds.megaphone.fm/vergecast","https://podnews.net/rss"],"maxEpisodesPerShow": 5,"includeTranscripts": false}
Save it as a task and schedule it daily (for example 0 8 * * *). Each run returns each show's latest 5 episodes; compare guid with your stored list to act only on new ones. The actor keeps no memory between runs, so each run is charged for every episode it returns; keep maxEpisodesPerShow small for alerts.
Audio for your own transcription
{"searchTerms": ["The Daily"],"maxEpisodesPerShow": 20,"includeTranscripts": false}
For shows that publish no transcript, send audio_url to the speech to text service of your choice. duration_seconds lets you estimate that cost before you start.
Input parameters
| Parameter | Type | Default | What it does |
|---|---|---|---|
searchTerms | array of strings | none (form prefill: Podnews Daily) | Show names. Each is looked up in Apple's public podcast directory and up to 3 matching feeds are read. |
feedUrls | array of strings | none (form prefill: the Podnews Weekly Review feed) | Podcast RSS feed URLs, for shows you already know or shows not in Apple's directory. |
maxEpisodesPerShow | integer (1 to 2,000) | 50 (form prefill: 15) | Most episodes to take from each show, in feed order. |
includeTranscripts | boolean | true | Download and convert each transcript the feed lists. One extra download per such episode; turn off for a fast metadata only pass. |
Give at least one of searchTerms or feedUrls; with neither, the run returns only a summary row and charges only the start fee. "Form prefill" values fill the Console form for you but are not defaults.
Run options. The default memory is 512 MB and the default timeout is 3,600 seconds. A metadata only run of a few hundred episodes takes well under a minute. With transcripts on, allow a few seconds per transcript; a large back catalogue with transcripts can take many minutes, so test with a small maxEpisodesPerShow first.
What data do you get?
One row per episode. Values a feed does not publish are left out of that row rather than sent empty.
The show: show_title, show_author, feed_url.
The episode: episode_title, description (show notes, HTML removed, first 5,000 characters), guid (the feed's unique ID for the episode), published_at (the date exactly as the feed writes it, such as Fri, 11 Sep 2026 13:00:00 +1000), duration_seconds, episode_number, season, explicit (as the feed writes it, such as false), episode_url (the episode's web page, when the feed has one).
Audio: audio_url, the direct link to the episode's audio file. Many hosts route it through an analytics prefix (for example op3.dev or podtrac.com); it still downloads the file.
Transcript: has_transcript (true when transcript text was captured), transcript_url (the transcript file the feed lists, present even when transcripts are switched off) and transcript (the text).
About transcripts, honestly
A transcript exists only when the publisher adds one to the feed, and that is still a minority of podcasts. No scraper can invent one. What we saw:
| Show | Transcripts listed in the feed |
|---|---|
| Podnews Weekly Review (Buzzsprout) | 99 of the latest 100 episodes (2 October 2026) |
| Podnews Daily | every one of the latest 10 (2 October 2026) |
| The Vergecast | none of the latest 100 (2 October 2026) |
| The Daily, Late Night Linux | none (our August 2026 check) |
Transcripts come back as the publisher wrote them, which is often machine made: expect speaker names with timestamps, the odd misheard name, and some HTML entities such as ' for an apostrophe. If a listed transcript file cannot be downloaded, the episode is still delivered, with has_transcript: false and the transcript_url it should have come from, so you can retry it yourself. For shows that publish nothing, send audio_url to a speech to text service.
Each run ends with a _type: "summary" row (shows_processed, episodes_delivered, with_transcript, charged_for) and a _type: "info" row with a short message. Neither is an episode and neither is ever charged. Skip rows that have a _type field when you load episodes.
Stable fields for automations
These 11 fields were present in every episode row we sampled (220 rows from two local runs of the live build on 2 October 2026, across three shows):
| Field | What it holds |
|---|---|
show_title | Show name |
show_author | Show author or publisher |
feed_url | The RSS feed read |
episode_title | Episode title, HTML entities decoded |
description | Show notes as plain text |
guid | The feed's episode ID; the key for deduplication across runs |
published_at | Publish date as the feed writes it |
duration_seconds | Length in seconds, number |
audio_url | Direct audio file URL |
has_transcript | true when transcript text is in the row |
scraped_at | ISO timestamp when the row was captured |
episode_number, season, explicit, episode_url and transcript_url are present when the feed publishes them. We will not rename these fields. New fields may be added over time; existing ones keep their names.
Output examples
Real rows from our tests of the live build on 2 October 2026, with show notes and transcripts trimmed.
An episode with a transcript (Podnews Weekly Review):
{"show_title": "Podnews Weekly Review","show_author": "James Cridland and Sam Sethi","feed_url": "https://feeds.buzzsprout.com/1538779.rss","episode_title": "Spotify add a \"Skip Ahead\" button, and recognise women podcasters","description": "Send James and Sam a message or voicemail\nWe talk with Spotify’s first Equal Podcast Ambassador, Morgan Absher, about building Two Hot Takes into a video-first community…","guid": "Buzzsprout-19612589","published_at": "Fri, 07 Aug 2026 16:00:00 +1000","duration_seconds": 5284,"episode_number": "32","season": "4","explicit": "false","audio_url": "https://op3.dev/e/www.buzzsprout.com/1538779/episodes/19612589-spotify-add-a-skip-ahead-button-and-recognise-women-podcasters.mp3","has_transcript": true,"transcript_url": "https://www.buzzsprout.com/1538779/19612589/transcript","transcript": "Jordan Blair 0:00 I just upload one file to Buzzsprout and it distributes to all the podcast directories, regardless of if they are video or audio, which is great. James Cridland 0:08 Jordan Blair from Buzz Sprout on the company's enhanced video service plus…","scraped_at": "2026-10-02T15:33:14.696Z"}
An episode from a show that publishes no transcript (The Vergecast, found with searchTerms):
{"show_title": "The Vergecast","show_author": "The Verge","feed_url": "https://feeds.megaphone.fm/vergecast","episode_title": "Dots get up in Muse's business","description": "Is the key to turning around public opinion on AI Super Intelligence just to make it cuter? Nilay Patel, Jake Kastrenakes and David Imel talk about the latest adorable agents…","guid": "cb171a08-c3c7-11f0-8340-c7672bd394c1","published_at": "Fri, 02 Oct 2026 14:44:00 -0000","duration_seconds": 5300,"audio_url": "https://www.podtrac.com/pts/redirect.mp3/pdst.fm/e/pscrb.fm/rss/p/mgln.ai/e/257/traffic.megaphone.fm/VMP7792375063.mp3","has_transcript": false,"scraped_at": "2026-10-02T16:10:35.881Z"}
Pricing
Pay per event: you pay for each episode delivered to your dataset, plus a small start fee per run. An episode costs the same with or without a transcript.
| Event | Free | Bronze | Silver | Gold and above |
|---|---|---|---|---|
episode-scraped, per episode | $0.001 | $0.0009 | $0.00075 | $0.0006 |
episode-scraped, per 1,000 episodes | $1.00 | $0.90 | $0.75 | $0.60 |
apify-actor-start, per run | $0.005 per GB of run memory, minimum one event | same | same | same |
The start fee, exactly. Apify's apify-actor-start event is charged once when a run starts, at $0.005 for each GB of memory the run uses, with a minimum of one event. This actor runs on 512 MB by default, so a default run pays one event: $0.005. Our daily check on 2 October shows apify-actor-start: 1 and episode-scraped: 3.
Worked examples. 200 episodes on the Free plan: $0.20 plus $0.005. A 1,087 episode back catalogue on Gold: $0.65 plus $0.005. Five shows checked daily at 5 episodes each for 30 days (750 episodes) on Gold: $0.45 plus $0.15 of start fees.
Never charged: a feed that fails to load, a show name with no match in Apple's directory, items with neither a title nor an audio file, transcript downloads (included in the episode price), and the summary and info rows. A run that delivers nothing pays only the start fee.
There is no scheduled price change for this actor. The Pricing tab on this page always shows the rate for your plan; if it and this table ever differ, the Pricing tab is right.
FAQ
What does it read? Podcast RSS feeds, the public XML files every podcast publishes so that apps like Apple Podcasts, Spotify and Pocket Casts can list its episodes, plus Apple's public podcast search to turn a show name into its feed. Both are open to anyone without signing in.
How many episodes can I get?
Up to 2,000 per show (maxEpisodesPerShow), as many shows as you list, and as far back as the feed goes. Many feeds carry the full back catalogue: The Vergecast's listed 1,087 episodes. Some hosts publish only the latest few hundred.
Why does an episode have no transcript?
Usually because the publisher did not add one to the feed, so transcript_url is absent too. If transcript_url is present but has_transcript is false, the transcript file could not be downloaded during the run; the episode is still delivered and you can fetch the file yourself.
Can it get Spotify exclusive shows? No. Shows without a public RSS feed cannot be read this way.
How fresh is the data? Every run reads the feeds live; nothing is cached. A new episode appears as soon as the show's host adds it to the feed.
Do I need an Apple or Spotify account, an API key or a proxy? No. You need only an Apify account. Feeds and Apple's directory search are public, and the actor fetches them directly.
Why is my run slower with transcripts on?
Each transcript is a separate file download, and some are hundreds of kilobytes. Turn includeTranscripts off for a fast metadata only pass; transcript_url still tells you which episodes have one.
Can I run it on a schedule?
Yes. Save your input as a task, then in Apify Console go to Schedules, Create new, and pick a time or a cron expression such as 0 8 * * *. Compare guid with your stored episodes to process only new ones.
How do I export the data?
From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API. Drop rows that have a _type field if you want episodes only. Long transcripts are easiest to handle as JSON.
Can I use it from Claude, ChatGPT or another AI assistant?
- Connector URL:
https://mcp.apify.com/?tools=themineworks/podcast-episode-scraper. - Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
- ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
- Cursor or VS Code: add it as an HTTP MCP server with that URL.
- Claude Code:
claude mcp add -t http podcast-episode-scraper "https://mcp.apify.com/?tools=themineworks/podcast-episode-scraper".
Is it legal to scrape podcast feeds? Podcast feeds are published so that anyone's app can read them, and the actor reads only those public feeds and Apple's public search. Show notes, transcripts and audio are still the publishers' copyrighted work, so you are responsible for how you use them, including copyright, each show's terms, and data protection laws such as GDPR and CCPA for any personal data they contain. This is general information, not legal advice.
Integrations
- Google Sheets: export a run to a sheet, or use Apify's Google Sheets integration to add each scheduled run's episodes.
- Make, Zapier and n8n: use the Apify app or node to start a run and send new episodes or transcripts to Slack, Notion or a vector database.
- Webhooks: have Apify call your URL when a run succeeds, then read the dataset.
- API and client libraries: start runs and read datasets from Python, JavaScript or any HTTP client. The "Copy to your AI assistant" block above has the exact call.
- MCP clients: Claude, ChatGPT, Cursor, VS Code and Claude Code can call the actor as a tool through
https://mcp.apify.com.
Building a text corpus? Pair it with YouTube Transcript Scraper for video and Substack Scraper for newsletters.
More from The Mine Works
Social media and video
- Threads Scraper
- Reddit Scraper
- Threads Search Scraper
- Instagram Profile Scraper
- Instagram Followers & Following
- Reddit Search Scraper
- Twitter / X Scraper
- YouTube Transcript
- Xiaohongshu (RED) Scraper
- Telegram Channel Scraper
- Telegram Channel Finder
- Pinterest Profile Scraper
Leads and business directories
Marketing, SEO and reviews
Real estate
Science, health and government data
Jobs and hiring
Company and business data
E-commerce and marketplaces
Food and local services
Developer and AI tools
More tools
Support
Found a feed that will not parse, or need a field we do not return yet? Open an issue on the Issues tab of this page with the feed URL and the run ID, and we will reply there. To ask for a new source, email dmineworks@gmail.com. A guide for this actor also lives at themineworks.com.
Podcast Episode Scraper turns any podcast feed or show name into episode rows with audio links and, where published, full transcripts, at $1 or less per 1,000 episodes.

