YouTube Subtitles & Timed Transcript Extractor avatar

YouTube Subtitles & Timed Transcript Extractor

Pricing

from $1.40 / 1,000 youtube video transcripts

Go to Apify Store
YouTube Subtitles & Timed Transcript Extractor

YouTube Subtitles & Timed Transcript Extractor

Extracts full spoken transcripts and timed subtitle segments from any public YouTube video without downloading video files or running slow GPU models. Ideal for AI video summarizers and RAG.

Pricing

from $1.40 / 1,000 youtube video transcripts

Rating

0.0

(0)

Developer

David Sandor

David Sandor

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

YouTube Subtitles & Timed Transcript Extractor πŸš€

Extracts full spoken transcripts and timed subtitle segments from any public YouTube video without downloading video files or running slow GPU models. Ideal for AI video summarizers and RAG.

🌟 20+ Enterprise Enhancements (v2.0)

  • RAG & LLM Ready: Pre-computed OpenAI token counts and chunked embeddings.
  • Smart Keyword Filters: Include or exclude records by targeted keyword lists.
  • Sentiment Scoring: Built-in lexical sentiment rating on text contents.
  • Noise & Tracking Scrubber: Removes tracking query parameters and boilerplate banners.
  • Pay-Per-Event (PPE): Ultra-cost-effective pricing per extracted item.
  • Zero Cold Start: Sub-second execution with automated fallback guarantees.

πŸ’» Integration Examples

Node.js (Apify Client)

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('youtube-transcript-extractor').call({
// Pass customized inputs here
enableRagEnrichment: true
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log('Extracted Items:', items);

Python

from apify_client import ApifyClient
client = ApifyClient('YOUR_APIFY_TOKEN')
run = client.actor('youtube-transcript-extractor').call(run_input={ 'enableRagEnrichment': True })
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item)

πŸ“„ Output Schema

Returns structured JSON, token counts, and RAG vector chunks.