YouTube Transcript Scraper — human vs auto avatar

YouTube Transcript Scraper — human vs auto

Pricing

from $2.00 / 1,000 auto-generated transcripts

Go to Apify Store
YouTube Transcript Scraper — human vs auto

YouTube Transcript Scraper — human vs auto

Give it YouTube video URLs or a search term to get the full transcript text and the language code actually used, on essentially every captioned video, with a plain true/false flag telling you whether each transcript was written by a human or auto-generated.

Pricing

from $2.00 / 1,000 auto-generated transcripts

Rating

0.0

(0)

Developer

RankFabrik Team

RankFabrik Team

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

13 days ago

Last modified

Share

YouTube Transcript Scraper — full text, language and a human-vs-auto flag on every row

Give it video URLs or a search term, and it returns the full transcript text plus the language code of the track actually used, on essentially every video that has captions at all.

Every row also tells you, with a plain true/false, whether that transcript was written by a human or generated by speech-to-text, and the accuracy gap between the two is not small. Almost no transcript API surfaces which one you got. This one flags it on every single row and reports the human-vs-auto split for the whole batch before you spend your budget, so you know the quality of what you are buying up front.

Unofficial tool — not affiliated. This is an independent tool. It is not affiliated with, endorsed by, or sponsored by YouTube, Google, or any other third-party service. All product names, logos and brands are the property of their respective owners and are used for identification only.


Source and legality

This reads the caption track that the source's own mobile-app clients receive from the public player: no account is ever created, no session, no paywalled content, and no anti-bot protection bypassed. It accesses only what is already public. Timestamped verbatim segments, plus SRT and VTT export, are live in production today, not on a roadmap.

Copyright. A transcript may reproduce content that is protected by copyright. This tool does not grant you any right or license over that content. You are responsible for ensuring that your use of any transcript (fair use, licensing, or private study, as the case may be) is lawful.

You see the quality mix before you pay, not after

Most transcript APIs work in the dark: you pay, you run, and only then learn how many videos actually had usable captions, and of what kind. This one shows you the human percentage — the share of rows with a human-written track — in two places: the run's status message while it works, and a batchCompleteness block on every single row. A video with no caption track, or a live stream, is still returned as a row so you know exactly why — but it is never billed. See Pricing.


What you get

One row per video:

FieldNotes
id11-character YouTube video ID
title, channel, channelIdvideo and channel metadata
durationSeconds, viewsduration, view count
description, keywordsvideo description and tags
isLivetrue if the video is a live stream
thumbnailthumbnail URL
languagelanguage code of the caption track actually used
auto / humantrue when the track is speech-to-text, false/true when a person wrote it — the core disclosure of this product
availableLanguagesevery language code the video has a track for, not just the one chosen
textfull transcript, plain text
totalWords, charactersword and character counts
segmentstimestamped sentence-by-sentence breakdown (start, duration, text), when requested
chapterseach chapter with its title, start time, and the text actually spoken during it — not just the chapter title, when requested
batchCompletenessthe completeness percentages of this run, including human (% of rows with a human-written track)

No field is ever inferred. A video with no captions is reported as unavailable and not billed — you never receive a transcript hallucinated from the title and description to fill the gap.

The batchCompleteness percentages mirror exactly what YouTube publishes for the videos you asked about, not what we drop. A low human reading simply means most of those videos only had a machine-generated track available — not that a human transcript was discarded. And because you read it before you pay, the quality mix is never a surprise you discover after the run.


Example output — one row, illustrative

{
"id": "jNQXAC9IVRw",
"title": "Example video title, as published",
"channel": "Example Channel",
"channelId": "UCExampleChannelId",
"durationSeconds": 245,
"views": 128394,
"description": "Video description, as published.",
"keywords": ["example", "keyword"],
"isLive": false,
"thumbnail": "https://i.ytimg.com/vi/jNQXAC9IVRw/hqdefault.jpg",
"language": "en",
"auto": false,
"human": true,
"availableLanguages": ["en", "fr"],
"text": "Full transcript text goes here...",
"totalWords": 512,
"characters": 2890,
"segments": [
{ "start": 0, "duration": 4.2, "text": "Example opening line of the transcript." }
],
"batchCompleteness": {
"text": 100,
"language": 100,
"channel": 98,
"durationSeconds": 100,
"views": 97,
"thumbnail": 100,
"human": 34
}
}

human: true here means a person wrote this particular track — the core disclosure of this product. batchCompleteness.human is the same object on every row of the run: read it once and you know what share of the whole batch is human-written versus machine-generated.


Good to know before you run it

Published up front, not buried in a changelog — so there are no surprises.

  • Auto-generated captions can be inaccurate, especially on accented speech, technical vocabulary or overlapping speakers. auto / human tells you which you got on every row, so you can filter or re-weight downstream instead of finding out the hard way.
  • Language preference wins over caption quality. If you ask for French and the video only has an auto-generated French track plus a professionally translated English one, you get the auto-generated French track, matching what a French-speaking viewer actually sees. Set languages to en if you'd rather have the higher-quality translation.
  • Chapters cost one extra request per video and are off by default. When on, chapter text is derived by slicing the transcript at the source's own chapter timestamps — it is not a separately authored chapter summary.
  • Live streams and videos without any caption track are not billed. They are still returned as a row (available: false, with a reason) so you know why, but you pay nothing for them.
  • Search mode (searchTerm) has a soft ceiling. Beyond a handful of result pages the source mostly re-serves what it already returned. Measured yield on six pages: +20, +10, +1, +6, +2. Set expectations accordingly for very broad terms.

Input

FieldRequiredDefaultMeaning
videosone of the twoone YouTube URL or video ID per line
searchTermone of the twosearch by keyword instead of listing videos
searchLimitno10max videos pulled from searchTerm
languagesnoenpreferred caption languages, in order
includeSegmentsnotrueinclude the timestamped sentence breakdown
includeChaptersnofalseinclude chapters with spoken text per chapter
maxResultsno25hard ceiling on videos processed (and billed) across both sources

Example:

{
"videos": "https://www.youtube.com/watch?v=dQw4w9WgXcQ\nhttps://youtu.be/9bZkp7q19f0",
"languages": "en",
"includeChapters": true,
"maxResults": 20
}

Pricing

Pay per result, in two tiers, because a human-written transcript is worth more than an auto-generated one:

EventPriceWhat it is
transcript-human$4.00 per 1,000 rowsa video with a human-written caption track
transcript-auto$2.00 per 1,000 rowsa video with a speech-to-text caption track

You are charged per row written to your dataset, and for nothing else. Videos with no usable caption track, live streams, and refused requests cost you zero. Your own spending cap is enforced by the platform — the run stops cleanly when it is reached, mid-dataset, without burning compute you did not authorise.


Prefer a hosted API with a flat monthly plan?

This actor bills per result on the Apify platform. If you'd rather call a hosted endpoint with your own key, a flat monthly quota, and the same human-vs-auto disclosure on every response, the same data ships as a standalone API:

  • YouTube Transcript API — the same caption-track transparency, on a flat monthly plan: https://rankfabrik.com/captions-api-pricing
  • Sibling APIs on the same principle: Places (local business data, one flat unit per enriched record) and Jobs (built for volume past the 1,000-result-per-query wall).

The engine behind both is the same, and both report the same measured fill rate. Choose the billing model that suits you.


Your responsibility as the user. You are solely responsible for how you use the transcripts this tool returns, including compliance with applicable laws, the source's terms, and third-party copyright (see the copyright note above). We provide a data-access tool and grant no license over any video's content.

Removal requests. Rights holders who want a specific record removed can contact contact@rankfabrik.com. Justified requests are processed promptly (target: within 30 days).


Support

Something look off in a run, or a question about a field? The actor's Issues tab is the place, and quoting the run helps: its completeness figures are logged, so we read the same numbers you did instead of trading guesses.