# TikTok Hashtag Scraper (`tokfluence/tiktok-hashtag-scraper`) Actor

Scrape videos posted under TikTok hashtags, one row per video with caption, views, likes, comments, shares and author, in Clockworks field names. Comes from Tokfluence's stored videos or a live crawl when TikTok allows it, so results can be short.

- **URL**: https://apify.com/tokfluence/tiktok-hashtag-scraper.md
- **Developed by:** [Tokfluence Tiktok API](https://apify.com/tokfluence) (community)
- **Categories:** Social media, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.17 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## TikTok Hashtag Scraper

Give it up to 10 TikTok hashtags and it returns the videos posted under them, one row per video. Rows use the same field names as clockworks/tiktok-hashtag-scraper (`authorMeta`, `videoMeta`, `searchHashtag`, ...). It is not a like-for-like replacement: read the next section before you rely on it.

### Read this first: what works today

**Live hashtag scraping is often blocked.** Since 2026-09-02 TikTok has walled its hashtag pages (`/tag/<name>`) for datacenter IP addresses. A live crawl needs Tokfluence's ISP proxies to get through, and we cannot promise that tier is available at any given time. When every proxy we try is blocked, that hashtag comes back empty and the run's status message says **WALLED** and names the hashtag. If the last attempt failed some other way (a captcha, a timeout), the message names that error instead. The run still succeeds and you are only charged for rows you receive.

**Stored hashtag feeds are thin.** In `database` mode (and in `auto` mode for a hashtag we crawled in the last 7 days) you get only the videos Tokfluence has already seen under that hashtag. Tokfluence keeps a hashtag's videos only from live crawls made through its scrape API, which is new, so expect few videos or none for most hashtags, not the full feed. A hashtag's stored feed grows only when it is crawled live.

So for now expect fewer videos per hashtag than TikTok shows, and some runs with none. When a run comes back short, the status message says what we know about why.

### What it returns

One video row per video, the same row shape as Clockworks' hashtag scraper. Filled when the stored video has them:

- `id`, `text`, `textLanguage`, `createTime`, `createTimeISO`, `webVideoUrl`, `isAd`
- `diggCount`, `playCount`, `commentCount`, `shareCount`, `collectCount`, `repostCount`
- `hashtags` (the caption's tags), `mentions`, `detailedMentions` and `effectStickers` (id and name)
- `musicMeta` (null when the video has no sound) and `videoMeta` (size, duration, format, cover, subtitle links)
- `authorMeta.id`, `authorMeta.name`, `authorMeta.profileUrl`, and the author's `fans`, `following`, `heart`, `video` and `friends` counts as they were when we captured the video
- `searchHashtag.name`: the hashtag this video was found under, without the `#` and in lower case
- `input`: the hashtag as you typed it

A video found under two of your hashtags appears once per hashtag, each row with its own `searchHashtag`, so it counts as two results.

Fields Tokfluence adds that Clockworks does not have sit under a `tokfluence` object:

- `tokfluence.scraped_at`: when Tokfluence last captured this video, so you can judge how fresh the counters are
- `tokfluence.is_branded` and `tokfluence.ad_disclosure`: Tokfluence's own reading of the caption for brand deals and ad labels. This is our classifier, not TikTok's sponsored flag, which is why `isSponsored` is null.

### Inputs

- `hashtags` (required): up to 10 per run, with or without the `#`. Case does not matter; `#Crochet` and `crochet` are the same hashtag and are sent once.
- `resultsPerPage`: the most videos to return **per hashtag**, 100 by default, at most 1000.
- `maxResults`: the most rows for the whole run.

### Fresh or fast: the Mode setting

Under **Advanced**, `mode` trades freshness for speed:

- `auto` (default): serves what Tokfluence already stores for a hashtag when it saw a video under that hashtag in the last 7 days, and crawls TikTok live for the rest. Most runs want this. A hashtag that goes live and is blocked comes back WALLED, even if we hold older videos for it.
- `database`: fastest. Never scrapes, so you get only the videos we have already seen under the hashtag (see above), newest sighting first, and their counters are as old as our last visit (check `tokfluence.scraped_at`). This is the way to get what we hold for a hashtag that came back WALLED.
- `live`: always crawls TikTok now. Freshest when it gets through; slower, and subject to the wall above.

A live crawl reads the top of the hashtag page, a fixed number of scrolls, not the whole feed, so a large `resultsPerPage` can come back short even when nothing is blocked. Every video a live crawl finds is stored, so a `database` run afterwards returns them without crawling again. Live crawls run on Tokfluence's shared scraping workers, one hashtag after another, and wait behind other live requests, so a live run can take a while.

If a live crawl outlasts the run's timeout, the run fails with the request id. The crawl keeps going and is stored when it finishes, so a `database` run shortly afterwards returns those videos.

### Fields that can be null

This actor never fills a gap itself: when we do not have a field it is `null`, not `0`, `false` or an empty string. Always null today:

- `searchHashtag.views`: we do not store a hashtag's total view count.
- `isMuted`, `isPinned`, `isSponsored`, `isSlideshow`, `locationCreated`: not captured. `hasTikTokShopProduct`: not filled by this actor.
- `mediaUrls`, `slideshowImageLinks`, `commentsDatasetUrl`: Clockworks fills these with files it copies; we do not copy files.
- `url`, `error`, `errorCode`, `invalidUrls`, `submittedVideoUrl`, `fromProfileSection`, `searchQuery`, `searchMusic`: they describe error rows or other Clockworks inputs. A hashtag we could not scrape is explained in the status message instead of in an error row.
- In `authorMeta`: `nickName`, `avatar`, `signature`, `bioLink`, `commerceUserInfo`, `ttSeller` and `region`, because a hashtag row carries the video, not the author's profile. Also `verified`, `privateAccount`, `digg`, `isUnderAge18`, `roomId`, `createTime`, `originalAvatarUrl` and `followDatasetUrl`.
- In `musicMeta`: `coverMediumUrl`, `originalCoverMediumUrl`, `playUrl`, `musicAlbum`.
- In `videoMeta`: `transcriptionLink`, and `downloadLink` inside `subtitleLinks`. The TikTok subtitle URL is in `tiktokLink` and is signed, so it stops working some time after the scrape.
- In `hashtags`: `title` and `cover`.
- In `detailedMentions`: `nickName` and `postUrl`.
- In `effectStickers`: `stickerStats`.

Often null:

- `hashtags[].id` and `detailedMentions[].id`: filled only when TikTok returned them with the video.
- `videoMeta.coverUrl` falls back to TikTok's own cover, which expires, when we have not stored a copy yet. `videoMeta.originalCoverUrl` and `videoMeta.downloadAddr` are often empty in the video as we stored it.
- `locationMeta`: only when the video is tagged with a place.

### Zero or short results

When a run returns nothing, or fewer rows than `maxResults`, it still succeeds, but it logs a warning and sets the run's status message to the likely cause:

- **WALLED**: TikTok blocked every proxy we tried on that hashtag's page. Switch Mode to `database` for what we already hold, or try again later.
- a hashtag that does not exist, or is not a valid hashtag;
- another error code for a hashtag (for example `CAPTCHA_REQUIRED`), which is usually worth retrying later;
- stored feeds being thin (`database` mode, or a hashtag `auto` served from storage);
- a live crawl reading only the top of the hashtag page;
- `resultsPerPage` times the number of hashtags being below `maxResults`.

The actor uses one Tokfluence account for every run. If that account is out of API credits, or Tokfluence's scrape service is down, the run fails with a message saying so. You are charged only for rows already in the dataset.

### What Tokfluence keeps from a run

Tokfluence logs every request the actor makes: the hashtags, the mode, where each row was served from, and a reference made of a one-way hash of your Apify user id plus the run id. Every hashtag you give, and every TikTok handle and hashtag in the returned rows, is recorded as a candidate for Tokfluence's creator discovery. Every video a live crawl finds is stored in Tokfluence like any other scrape, and later runs, anyone's, can be served that stored copy.

### Memory

The actor defaults to 256 MB and allows up to 512 MB. It makes one API call, maps the rows and writes them to the dataset in batches of 100.

### Pricing

PLACEHOLDER (Daniel): pay per result, one `result` event per dataset row. Price to be set in Console at publish.

If your plan's remaining budget covers fewer rows than you asked for, the actor collects up to that ceiling, says so in the log, and stops cleanly rather than returning a silently short list.

### Notes

Data comes from public TikTok pages, scraped and stored by Tokfluence. You are the data controller for anything you export; follow GDPR and any local rules that apply to you. More at [Tokfluence](https://tokfluence.com).

# Actor input Schema

## `hashtags` (type: `array`):

Up to 10 hashtags per run, with or without the #. Each video found comes back as one row, with searchHashtag naming the tag it was found under.

## `resultsPerPage` (type: `integer`):

Maximum number of videos returned for each hashtag, up to 1000. Max results below caps the whole run.

## `maxResults` (type: `integer`):

Stops once this many rows are in the dataset. Each row is one result.

## `mode` (type: `string`):

Trades freshness for speed. Auto serves what Tokfluence already stores when it is fresh enough and scrapes live for the rest. Database only is fastest and never scrapes, so anything we do not already hold is missing from the results. Live always scrapes TikTok now: freshest, and slower. For hashtags, live scraping is often blocked by TikTok (see the README), and stored hashtag feeds are still thin.

## Actor input object example

```json
{
  "hashtags": [
    "crochet"
  ],
  "resultsPerPage": 100,
  "maxResults": 100,
  "mode": "auto"
}
```

# Actor output Schema

## `results` (type: `string`):

Rows written to the run's dataset, one item per result, in Clockworks field names.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "hashtags": [
        "crochet"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tokfluence/tiktok-hashtag-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "hashtags": ["crochet"] }

# Run the Actor and wait for it to finish
run = client.actor("tokfluence/tiktok-hashtag-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "hashtags": [
    "crochet"
  ]
}' |
apify call tokfluence/tiktok-hashtag-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tokfluence/tiktok-hashtag-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6A6XbFUreH6nhlF7b/builds/FHGfESqhnZV8Dlgjj/openapi.json
