TikTok Transcript Scraper avatar

TikTok Transcript Scraper

Pricing

from $0.90 / 1,000 results

Go to Apify Store
TikTok Transcript Scraper

TikTok Transcript Scraper

Get plain-text transcripts of TikTok videos by URL, taken from TikTok's own subtitles. One row per video that has subtitles; videos without subtitles return no transcript, and there is no AI speech-to-text.

Pricing

from $0.90 / 1,000 results

Rating

0.0

(0)

Developer

Tokfluence Tiktok API

Tokfluence Tiktok API

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Give it TikTok video URLs and it returns the plain-text transcript of each video that has TikTok subtitles, one row per video. Rows use the field names of Clockworks' TikTok transcript row where we have the data; the transcript text itself sits under tokfluence.transcript (see below for what that means for an existing pipeline).

Read this first

  • Only videos with TikTok subtitles have a transcript. Some TikTok videos carry none. Those have no transcript here, and no run can change that.
  • There is no AI transcription. We do not run speech-to-text on the audio. A video without TikTok subtitles gives no row, not an empty or guessed transcript, and you are not charged for it. The run's status message names those videos.
  • Plain text, one language. The transcript is one block of text taken from one of the video's subtitle tracks. Segments, timestamps and subtitle files (VTT/SRT) are not kept, so the row has no timing fields.

What it returns

One row per video that has a transcript. Filled fields:

  • id: the TikTok video id
  • webVideoUrl: the video's URL
  • submittedVideoUrl: the URL as you gave it in the input, so you can match rows back to your list

Fields Tokfluence adds, under a tokfluence object:

  • tokfluence.transcript: the transcript as plain text
  • tokfluence.language: the subtitle language as TikTok labels it, for example eng-US
  • tokfluence.source: where the text came from; today always TIKTOK_AUTO_CAPTION, TikTok's own captions
  • tokfluence.summary: a short summary of the transcript written by a language model, when Tokfluence has made one, otherwise null

If you are moving a pipeline from another transcript actor: Clockworks' transcript row does not carry the text in the row. It links to files, in videoMeta.subtitleLinks[].downloadLink (timed subtitles) and videoMeta.transcriptionLink (AI transcripts). We have no such files, so both are null here, and your pipeline should read tokfluence.transcript instead of downloading a link.

Input

  • postURLs: TikTok video URLs in the full https://www.tiktok.com/@user/video/<id> form. Short vm.tiktok.com links are not accepted and are named in the status message. Up to 50 per run.
  • maxResults: stops once this many rows are in the dataset.
  • mode, under Advanced: see below.

There is no option to transcribe videos without subtitles, because we do not do that.

Fresh or fast: the Mode setting

Under Advanced, mode trades freshness for speed:

  • auto (default): serves the transcript Tokfluence already stores, and scrapes TikTok live for videos we do not hold or hold without a transcript. Most runs want this.
  • database: fastest. Never scrapes, so any video we have not captured, or captured without a transcript, is missing from the results.
  • live: always scrapes TikTok now, and slower. Tokfluence keeps the first transcript it captures for a video, so for a video we already hold a transcript for, live returns the same text as auto; auto already goes live for videos we hold without one.

Live requests run on Tokfluence's shared scraping workers, one video after another, and wait behind other live requests, so a live run can take a while.

If a live scrape outlasts the run's timeout, the run fails with the request id. The scrape keeps going and is stored when it finishes, so a database run shortly afterwards returns those transcripts without scraping again.

Fields that can be null

This actor never fills a gap itself: when we do not have a field it is null, not 0, false or an empty string. Always null in this actor:

  • videoMeta.subtitleLinks and videoMeta.transcriptionLink: we store plain text, not subtitle or transcript files
  • text: in Clockworks' row this is the video caption, not the transcript; the transcript response does not include the caption
  • mediaUrls: we do not download video files
  • url, error, errorCode, invalidUrls: videos we could not serve are reported in the run's status message and log rather than as error rows

Often null:

  • tokfluence.summary: only some transcripts have a summary

Zero or short results

When a run returns fewer rows than the videos you gave it, it still succeeds, but it logs a warning and sets the run's status message to the likely cause, naming the videos:

  • videos with no TikTok subtitles, so no transcript exists (not charged); in database mode, videos we hold without a transcript yet
  • URLs we could not find, or could not read (for example short links)
  • database mode, which never scrapes

The actor uses one Tokfluence account for every run. If that account is out of API credits, or Tokfluence's scrape service is down, the run fails with a message saying so. You are charged only for rows already in the dataset.

What Tokfluence keeps from a run

Tokfluence logs every request the actor makes: the URLs, the mode, where each row was served from, and a reference made of a one-way hash of your Apify user id plus the run id. Every TikTok handle in your URLs and in the returned rows is recorded as a candidate for Tokfluence's creator discovery. Every video scraped live is stored in Tokfluence like any other scrape, and later runs, anyone's, can be served that stored copy.

Memory

The actor defaults to 256 MB and allows up to 512 MB. It makes one API call, maps the rows and writes them to the dataset in batches of 100.

Pricing

PLACEHOLDER (Daniel): pay per result, one result event per dataset row. Price to be set in Console at publish.

Only videos with a transcript become rows, so only they are charged. If your plan's remaining budget covers fewer rows than you asked for, the actor collects up to that ceiling, says so in the log, and stops cleanly rather than returning a silently short list.

Notes

Transcripts come from subtitles TikTok publishes on public videos, captured and stored by Tokfluence. You are the data controller for anything you export; follow GDPR and any local rules that apply to you. More at Tokfluence.