Podcast to Highlights
Pricing
from $100.00 / 1,000 podcast processeds
Podcast to Highlights
Turn a YouTube podcast (or any audio URL) into a timestamped transcript and the 5-8 moments most likely to work as short clips. Captions first, speech-to-text on demand, optional LLM rescoring with your own key.
Pricing
from $100.00 / 1,000 podcast processeds
Rating
0.0
(0)
Developer
shortlane
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
20 hours ago
Last modified
Categories
Share
Give it a YouTube podcast (or a direct audio/video URL) and get back a timestamped transcript plus the 5-8 moments most likely to work as short clips: start and end times, a suggested title, a hook score and the reasons behind it. Built for clippers, editors and agencies who cut Shorts, Reels and TikToks from long conversations.
How it works
- Captions first. If the video has YouTube captions (manual or automatic), the transcript is built from them in seconds and you pay only the per-podcast fee.
- Speech-to-text when needed. No captions, a direct audio file, or
transcribe_speechon: the audio is transcribed with a speech model and charged per minute. - Highlight detection. Sentences are grouped into 15-60 s windows and scored for hooks (questions, bold claims, concrete numbers, story markers, pacing). Overlapping windows are removed and the top N are returned in timeline order.
- Optional LLM rescoring. Add your own OpenAI or Anthropic key to get sharper titles and a second opinion on the score. The key is sent only to that provider and never stored or logged.
Output
One dataset record per highlight:
{"rank": 1,"title": "80% of your time goes to manual work","start": 812.4,"end": 846.9,"duration_s": 34.5,"score": 88.0,"reasons": ["hook at the start", "concrete number", "ideal length"],"text": "did you know that eighty percent of your time goes to manual work? ...","source_url": "https://www.youtube.com/watch?v=...","source_title": "A conversation with ...","transcript_source": "youtube-captions"}
The full transcript (segments and words with timestamps) is saved in the run's key-value store as transcript.json; the OUTPUT record summarizes the run (duration, transcript source, minutes charged, LLM status).
Input
| Field | Description | Default |
|---|---|---|
url | YouTube video URL or direct audio/video URL | required |
highlights | Number of moments to return (3-15) | 6 |
language | Spoken language code (en, es, ...) | auto |
transcribe_speech | Force speech-to-text even if captions exist | false |
llm_provider / llm_api_key | Optional rescoring with your own key | — |
max_duration_hours | Reject longer sources before processing | 4 |
Pricing
- $0.10 per podcast processed (captions route).
- + $0.02 per minute of audio when speech-to-text is used (a 60-minute episode without captions costs $1.30 in total).
- Nothing is charged if the URL is invalid, the source is too long, or processing fails.
Limits
-
YouTube and proxies. YouTube blocks most datacenter IPs ("Sign in to confirm you're not a bot"). For YouTube URLs enable Proxy for YouTube with Apify residential proxies (billed to your Apify account, a few KB per run on the captions route) or pass your own proxy. Direct audio/video URLs need no proxy. If YouTube still blocks the request the run fails and nothing is charged.
-
Sources up to 4 hours (configurable up to 8).
-
YouTube captions are used as published; automatic captions may contain recognition errors. Speech-to-text uses a compact model tuned for speed; timestamps are accurate to about half a second.
-
Only public YouTube videos and publicly reachable files. No downloads of video are made in the captions route.
Pairs well with
Content Rewards Campaign Finder: find paid clipping campaigns, then run this Actor on their reference material.
Questions: shortlane.media@gmail.com