TikTok Video Scraper avatar

TikTok Video Scraper

Pricing

from $0.23 / 1,000 results

Go to Apify Store
TikTok Video Scraper

TikTok Video Scraper

Scrape TikTok videos by URL: caption, views, likes, comments, shares, saves, sound, hashtags and author, in the same field names as clockworks/tiktok-video-scraper, served from Tokfluence's database when scraped in the last 24 hours and scraped live otherwise.

Pricing

from $0.23 / 1,000 results

Rating

0.0

(0)

Developer

Tokfluence Tiktok API

Tokfluence Tiktok API

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Give it TikTok video URLs and it returns one row per video: caption, views, likes, comments, shares, saves, sound, hashtags, mentions and the author's counts. Rows use the same field names as clockworks/tiktok-video-scraper. The fields we cannot fill are null, and we write no error rows, so check the lists below before you point an existing pipeline at it.

What it returns

One Clockworks video row per URL. These fields are filled when the stored video has them:

  • the video: id, text, textLanguage, createTime, createTimeISO, webVideoUrl, isAd
  • counters at the time we scraped it: playCount, diggCount, commentCount, shareCount, collectCount, repostCount
  • hashtags (id and name), mentions, detailedMentions (id, handle, profile URL, secUid), effectStickers (id and name)
  • musicMeta: sound id, name, author and whether it is original
  • videoMeta: width, height, duration, definition, format, cover, and TikTok's subtitle links where the video has subtitles
  • authorMeta: the author's id, handle, profile URL, and their followers, following, likes, video and friend counts as the video page showed them when we scraped it
  • locationMeta when the video is geotagged; hasTikTokShopProduct when we have checked the video for a Shop product tag
  • input and submittedVideoUrl: the URL you gave, as you typed it

Fields Tokfluence adds that Clockworks does not have sit under a tokfluence object:

  • tokfluence.scraped_at: when Tokfluence scraped this video, so you can see how fresh the counters are
  • tokfluence.video_engagement_rate: this video's engagement as a percentage (7.17 means 7.17%): (likes + comments + shares) / views
  • tokfluence.is_branded and tokfluence.ad_disclosure: Tokfluence's own branded-content signal. It is our classification, not TikTok's flag, which is why Clockworks' isSponsored stays null.
  • tokfluence.content_category, tokfluence.content_subcategory, tokfluence.topics: our content classification, null until it has run on the video

Input

  • postURLs (required): full TikTok video URLs, one per line, in the form https://www.tiktok.com/@user/video/<id>. Tracking query strings such as ?is_from_webapp=1 are fine. At most 200 per run. The same video given twice is scraped once.
  • maxResults: stop after this many videos. URLs past it are not scraped. Default 200.
  • mode (under Advanced): see below.

Short links (vm.tiktok.com/..., tiktok.com/t/...) are not supported yet. Open one in a browser and paste the full URL it lands on.

Clockworks' other video-scraper inputs are not supported, and have no effect if you pass them through the API: scrapeRelatedVideos and resultsPerPage (no related videos), shouldDownloadVideos, shouldDownloadCovers, shouldDownloadSlideshowImages and videoKvStoreIdOrName (no downloads), and downloadSubtitlesOptions (no subtitle files or speech-to-text; Tokfluence has a separate transcript actor).

Fresh or fast: the Mode setting

Under Advanced, mode trades freshness for speed:

  • auto (default): if Tokfluence scraped the video in the last 24 hours, you get that stored copy at once; anything older, or never seen, is scraped from TikTok now. Most runs want this.
  • database: fastest. Never scrapes. A video we have never scraped is missing from the results, and a video we have is as old as our last visit, with no upper bound: check tokfluence.scraped_at. We keep only the 15 most recent posts per creator in the database, so an older video of a creator we track can be missing even if we once scraped it.
  • live: scrapes every video from TikTok now, even one we scraped an hour ago. Freshest, and slowest.

Live scraping opens each video page in turn, so it takes time. Live requests run on Tokfluence's shared scraping workers and wait behind other live requests, so a live run can also wait before it starts. URLs are sent in batches of 50, one batch after another; a live run of 200 videos can take many minutes, so give it a generous run timeout.

If the first batch outlasts the run's timeout, the run fails with the request id. If a later batch does, the run keeps the videos already delivered, says in its status message which batch stopped and why, and does not send the batches after it. Either way the batch that was cut off keeps scraping and is stored when it finishes, so a database run shortly afterwards returns those videos without scraping again.

Fields that can be null

This actor never fills a gap itself: when we do not have a field it is null, not 0, false or an empty string.

Always null for this actor:

  • authorMeta.nickName, avatar, signature, bioLink, verified, privateAccount, commerceUserInfo, ttSeller, region: these come from the profile, and a video row does not include the author's profile. For the full profile, use the Tokfluence profile actor.
  • authorMeta.digg, isUnderAge18, roomId, createTime, originalAvatarUrl, followDatasetUrl: never captured.
  • isMuted, locationCreated, isSlideshow, slideshowImageLinks, originalVideoDetail: not captured.
  • isPinned: pinning belongs to a profile's grid, and this actor reads single videos.
  • isSponsored: see tokfluence.is_branded above.
  • mediaUrls, videoMeta.originalDownloadAddr, videoMeta.transcriptionLink, and downloadLink inside videoMeta.subtitleLinks: we do not download files into a key-value store.
  • commentsDatasetUrl: this actor does not scrape comments; the Tokfluence comments actor does.
  • hashtags[].title and hashtags[].cover, detailedMentions[].nickName and postUrl, effectStickers[].stickerStats, musicMeta.coverMediumUrl, musicMeta.originalCoverMediumUrl, musicMeta.playUrl, musicMeta.musicAlbum: not captured.
  • url, error, errorCode, invalidUrls, fromProfileSection, searchQuery, searchHashtag, searchMusic: see "Zero or short results" for errors; the rest belong to other Clockworks actors.

Often null:

  • videoMeta.downloadAddr and videoMeta.originalCoverUrl: often empty in the video as we stored it.
  • hasTikTokShopProduct: null means we have not checked the video, not "no product".
  • locationMeta: null when the video is not geotagged.
  • musicMeta as a whole is null for a video with no sound.

Links that expire: videoMeta.originalCoverUrl, videoMeta.downloadAddr and every subtitle tiktokLink are TikTok's signed URLs and stop working some time after the scrape. videoMeta.coverUrl is our own stored copy of the cover when we have one, and TikTok's signed cover otherwise.

Zero or short results

Unlike Clockworks, this actor does not write error rows into the dataset, so you are never charged for a video you did not get. When a run returns fewer videos than you gave URLs, it still succeeds, logs each problem and sets the run's status message to the likely cause:

  • a video TikTok says is unavailable or not found (in database mode, also one we have not stored);
  • a URL that is not a full video URL (short links included);
  • a video that could not be scraped live this time, with the error code (for example CAPTCHA_REQUIRED); run those again later;
  • database mode, which never scrapes.

The actor uses one Tokfluence account for every run. If that account is out of API credits, or Tokfluence's scrape service is down, the run fails with a message saying so. You are charged only for rows already in the dataset.

What Tokfluence keeps from a run

Tokfluence logs every request the actor makes: the URLs, the mode, where each row was served from, and a reference made of a one-way hash of your Apify user id plus the run id. Every TikTok handle and hashtag in your URLs and in the returned rows is recorded as a candidate for Tokfluence's creator discovery. Every video scraped live is stored in Tokfluence like any other scrape, and later runs, anyone's, can be served that stored copy.

Memory

The actor defaults to 256 MB and allows up to 512 MB. It makes one API call per 50 URLs, maps the rows and writes them to the dataset in batches of 100.

Pricing

PLACEHOLDER (Daniel): pay per result, one result event per dataset row. Price to be set in Console at publish.

If your plan's remaining budget covers fewer videos than you asked for, the actor collects up to that ceiling, says so in the log, and stops cleanly rather than returning a silently short list.

Notes

Data comes from public TikTok pages, scraped and stored by Tokfluence. You are the data controller for anything you export; follow GDPR and any local rules that apply to you. More at Tokfluence.