# TikTok Keywords Discovery (`solalab_digital/tiktok-keywords-discovery`) Actor

Expands seed keywords into TikTok search autocomplete suggestions with A-Z/0-9 expansion, recursion, raw TikTok signals, clustering and an HTML report.

- **URL**: https://apify.com/solalab\_digital/tiktok-keywords-discovery.md
- **Developed by:** [Sankov Vadim](https://apify.com/solalab_digital) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## TikTok Keywords Discovery: autocomplete expansion, scoring and clustering

Give it a seed phrase and it returns hundreds of real TikTok search suggestions. Each phrase carries the raw signals TikTok attaches to it, a 0-100 score, a topic cluster and intent tags. You can also add hashtag view counts and compare a run against the previous one. No login, no browser, no cookies.

You get A-Z / 0-9 expansion and recursion, raw TikTok signals, scores and clusters, multi-region merge, hashtag stats, change monitoring and an HTML report. It runs on 256 MB.

### 🚀 Quick start

1. Open the Actor and type one or more seed phrases into **Seed keywords** (for example `skincare routine`).
2. Leave the defaults for a cheap trial (suffix expansion, up to 40 requests, up to 50 rows) and press **Start**.
3. Open the **Overview** tab of the dataset, or open the `REPORT` file in the run's key-value store for a sortable HTML dashboard. Then raise `maxRequests` and `maxResults` for a full run.

### 🧭 Who it is for

- **SEO and content marketers.** Build a topic map for TikTok search, group phrases by cluster, and see which questions and "how to" queries people type.
- **TikTok creators.** Find video ideas that TikTok itself offers under a topic, with a rough strength score and the hashtag each idea maps to.
- **E-commerce sellers and agencies.** Filter phrases that carry TikTok's e-commerce intent flag, compare regions, and re-run the same list every month to see what is new.

### ✨ What it does

#### A-Z and 0-9 expansion, plus recursion

TikTok returns at most 10 suggestions per query. To get past that ceiling the Actor appends (or prepends) every letter and digit to your seed and queries each variant. Pick `suffix`, `prefix`, `both` or `none`, and choose the alphabets: Latin (26), digits (10), Cyrillic (33).

Recursion feeds the best phrases of each level back in as new seeds, up to 3 levels deep. `recursionWidth` sets how many phrases per level are reused. Recursive seeds are queried as they are, without alphabet expansion, which keeps the request count predictable.

Every run is bounded by `maxRequests` (up to 5000) and `maxResults` (up to 20000).

#### Raw TikTok signals

Each row carries the fields TikTok returns next to the phrase, unmodified:

- `ecomIntent`: TikTok's e-commerce intent flag
- `hotLevel`: TikTok's "hot" flag
- `isTimeSensitive`: marks phrases tied to a moment
- `predictCtrScore`: TikTok's own predicted click-through signal
- `lang`, `recallReason`, `cutQuery`: language, retrieval channel and TikTok's tokenization

#### Score 0-100

A transparent heuristic, not search volume:

`base = 40*F + 20*P + 20*C + 10*R + 5*E + 5*H`

- **F**: how many distinct queries returned the phrase (log scale, maxes out at 16)
- **P**: best position in the suggestion list (1st = full, 10th = zero)
- **C**: `predictCtrScore`, maxes out at 0.08
- **R**: share of your requested regions where the phrase appeared
- **E / H**: 1 if `ecomIntent` / `hotLevel` is above zero

When hashtag stats are available, 5% of the score is swapped for reach: `score = 0.95*base + 5*V`, where V is log-scaled hashtag views (maxes out at 10 billion). The result is capped at 100 and rounded.

#### Clusters and intent tags

Phrases are grouped without any ML. Each phrase gets the token it shares with the most other phrases (seed words and common stopwords are ignored, ties are broken by longer token, then alphabetically). Phrases with no shared token land in `other`. Intent tags are rule-based: `question`, `howto`, `for`, `vs`, `near_me`, `ecom`.

#### Several regions in one run

Put up to 10 region codes in `regions`. The same phrase found in different regions is merged into one row with `regions[]` and `regionCount`, and the region share feeds the score.

#### Monitoring

Turn on `compareWithPrevious` and every row gets a `monitor` block: `new` or `existing`, the previous score and the change. A `MONITOR_DIFF` file lists new phrases, changed phrases (score moved by 10 or more) and phrases that disappeared. Disappeared phrases are never added to the dataset and are not charged. The snapshot is stored in a named store and is overwritten only after a clean, uninterrupted run.

#### Hashtag stats

Switch on `enrichHashtags` and set `proxyConfiguration` to Apify Proxy with the `RESIDENTIAL` group to add `hashtag`, `hashtagViews` and `hashtagVideos`. The hashtag is the explicit `#tag` if the phrase has one. Otherwise the phrase is collapsed into a single token, the way TikTok forms hashtags: `skincare routine` becomes `#skincareroutine`. That collapsed form is a candidate, so a tag that does not exist simply returns empty stats. `maxHashtagLookups` limits how many tags are looked up (default 25, max 500).

#### Video captions through oEmbed

Add TikTok video links to `videoUrls` and the Actor returns extra rows of type `videoCaption`: caption, author, thumbnail and every hashtag found in the caption.

#### HTML report

Each run saves a self-contained `REPORT` page: sortable table, filters (text, minimum score, cluster, region, e-commerce), top-phrase cards, a cluster chart, the monitoring block and a CSV export of whatever is filtered. No external scripts.

### 🆚 Why use this Actor

| You need | What you get here |
|---|---|
| More than 10 phrases per seed | A-Z / 0-9 expansion and recursion with hard request caps |
| To judge which phrases matter | Raw TikTok signals plus a documented 0-100 score |
| To organise hundreds of phrases | Clusters and intent tags on every row |
| Every seed that led to a phrase | `seeds[]` keeps all of them, nothing is dropped as a duplicate |
| Regional comparison | `regions[]` merge in a single run |
| To track change over time | Named-store snapshots and a diff file |
| A quick visual read | Built-in HTML report |
| Low cost | Plain HTTP, no browser, 256 MB |

### ⚙️ Input

```json
{
  "keywords": ["skincare routine"],
  "regions": ["US"],
  "language": "en",
  "expandMode": "suffix",
  "alphabets": ["latin", "digits"],
  "recursionDepth": 0,
  "maxRequests": 40,
  "maxResults": 50,
  "enrichHashtags": true,
  "compareWithPrevious": false
}
```

| Field | Type | Default | Meaning |
|---|---|---|---|
| `keywords` | string\[] | `["skincare routine"]` | Seed phrases, 1 to 500, duplicates removed. Required. |
| `regions` | string\[] | `["US"]` | Two-letter region codes, 1 to 10. Each region multiplies requests. |
| `language` | select | `en` | Sent as `app_language`: en, es, fr, de, it, pt, tr, id, ja, ko, ru. |
| `includeSeedEcho` | boolean | false | Keep suggestions identical to the query that produced them. |
| `resultOrder` | select | `score` | `score`, `source` (discovery order) or `alphabetical`. |
| `expandMode` | select | `suffix` | `none`, `suffix`, `prefix`, `both`. |
| `alphabets` | select\[] | latin, digits | `latin`, `digits`, `cyrillic`. |
| `recursionDepth` | integer | 0 | 0 to 3 levels of feeding results back as seeds. |
| `recursionWidth` | integer | 10 | 1 to 50 phrases reused per level. |
| `maxRequests` | integer | 40 | Hard cap on requests per run, 1 to 5000. |
| `maxResults` | integer | 50 | Cap on unique rows, 1 to 20000. |
| `maxSuggestionsPerKeyword` | integer | empty | Cap on unique phrases attributed to one seed. |
| `minScore` | integer | 0 | Drop rows below this score. |
| `onlyEcommerce` | boolean | false | Keep only rows with a non-zero `ecomIntent`. |
| `includeTerms` / `excludeTerms` | string\[] | empty | Case-insensitive substring filters. |
| `enrichHashtags` | boolean | false | Add hashtag views and video counts. Needs residential proxy in the cloud. |
| `maxHashtagLookups` | integer | 25 | Ceiling for hashtag lookups, 1 to 500. |
| `videoUrls` | string\[] | empty | Video links to resolve into caption rows. |
| `compareWithPrevious` | boolean | false | Compare with the last snapshot for this key. |
| `monitorKey` | string | derived | Snapshot slot name. Empty means it is derived from seeds, regions and language. |
| `maxConcurrency` | integer | 8 | Parallel requests, 1 to 20. |
| `maxRequestsPerSecond` | number | 8 | Rate cap, 0.5 to 25. Halved automatically after a 429. |
| `proxyConfiguration` | proxy | off | Optional Apify Proxy. Use `RESIDENTIAL` for hashtag stats. |

**Request math.** Requests per seed = 1 + alphabet size (times two for `both`), multiplied by the number of regions. With the default Latin + digits alphabets that is 37 requests per seed per region for `suffix`. `maxRequests` always wins.

### 📦 Output

One row per unique phrase. Here is a real row from a local test run (`enrichHashtags` and `compareWithPrevious` on). In the cloud the hashtag fields need a residential proxy:

```json
{
  "resultType": "keywordSuggestion",
  "seedKeyword": "skincare routine",
  "seeds": ["skincare routine"],
  "suggestion": "skincare routine for oily skin",
  "normalizedSuggestion": "skincare routine for oily skin",
  "rank": 4,
  "sourcePlatform": "tiktok",
  "sourceSurface": "search_autocomplete",
  "suggestionType": "sug",
  "language": "en",
  "region": "US",
  "regions": ["US"],
  "regionCount": 1,
  "hits": 5,
  "depth": 0,
  "firstQuery": "skincare routine",
  "ecomIntent": 1,
  "hotLevel": 0,
  "isTimeSensitive": 0,
  "predictCtrScore": 0.018181765,
  "lang": "en",
  "recallReason": "tiktok_index_experience_decision_query|tiktok_index_active_7d_query|tiktok_orion_search_session|tiktok_experience_orion_query|tiktok_orion_query|tiktok_index_global_active_7d_query",
  "cutQuery": ["skincare", "routine", "for", "oily", "skin"],
  "hashtag": "skincareroutineforoilyskin",
  "hashtagViews": 24783339,
  "hashtagVideos": 1051,
  "isSeedEcho": false,
  "score": 59,
  "cluster": "skin",
  "intents": ["for"],
  "monitor": { "status": "existing", "scorePrev": 66, "scoreDelta": -7 },
  "hashtags": null,
  "scrapedAt": "2026-09-26T20:40:18.443203+00:00"
}
```

| Group | Fields |
|---|---|
| Phrase | `suggestion`, `normalizedSuggestion`, `rank` (best position, 1-10), `seedKeyword`, `seeds`, `firstQuery`, `depth`, `isSeedEcho` |
| Locale | `language`, `region`, `regions`, `regionCount`, `lang` |
| Signals | `ecomIntent`, `hotLevel`, `isTimeSensitive`, `predictCtrScore`, `recallReason`, `cutQuery`, `hits` |
| Analysis | `score`, `cluster`, `intents` |
| Hashtag | `hashtag`, `hashtagViews`, `hashtagVideos`, `hashtags` (tags found inside the text) |
| Monitoring | `monitor` (`status`, `scorePrev`, `scoreDelta`), null when comparison is off |
| Video rows | `videoUrl`, `caption`, `authorName`, `authorUrl`, `thumbnailUrl`, `hashtags` (`resultType: videoCaption`) |

Dataset views: **Overview**, **Raw TikTok signals**, **Hashtags**, **Video captions**. Export as JSON, CSV, Excel, XML, RSS or HTML from the dataset tab, or read it through the Apify API.

### ⚠️ Limits, stated plainly

- **Unofficial endpoint.** The Actor reads the public autocomplete request that TikTok's own search box uses. It is not an official API. TikTok can change or restrict it at any time, and the Actor may return fewer rows or fail if that happens.
- **No search volume.** There is no volume, CPC or trend data. The score is a heuristic built from the signals above.
- **Region is labelling, not local results.** The `region` value is sent as a parameter. Without a proxy in the matching country it does not guarantee results as a local user would see them. The language of the seed itself changes results far more than the region setting.
- **Hashtag stats need a residential proxy and are partial.** TikTok returns empty answers to datacenter IPs on this endpoint. In our cloud test, without a proxy and with the datacenter group, 0 of 25 tags resolved. With Apify Proxy group `RESIDENTIAL`, 13 of 25 did, and the rest stayed empty. Expect that order of magnitude, not full coverage, and about $0.001 of proxy traffic per 25 lookups. Parallel lookups can also trigger 429 responses; the Actor halves the request rate, waits for `Retry-After` when given, and retries up to 3 times. A lookup that still fails leaves `hashtagViews` and `hashtagVideos` as null instead of failing the run. The `hashtag` value is a candidate and may not exist as a real tag.
- **Expansion and recursion are capped.** `maxRequests`, `maxResults`, recursion depth (max 3) and width bound every run, so you may not see every phrase TikTok knows.
- **Rate ceiling unknown.** TikTok does not publish one. Lower `maxRequestsPerSecond` or add a proxy for very large runs.
- **Partial results are kept.** On abort, migration or timeout the Actor saves what it has collected.

### 💳 Pricing

Pay per event. See the **Pricing** tab of this Actor for the current event prices. A row is charged when it is added to the dataset, and the run stops cleanly when your spending limit is reached. Score, clusters, monitoring and hashtag stats cost nothing extra. `videoCaption` rows are dataset rows like any other.

### ❓ FAQ

**Does it give search volume?**
No. TikTok does not expose it in this endpoint. Use `score`, `hits` and the raw signals to compare phrases against each other.

**Why fewer rows than expected?**
Each query returns at most 10 phrases, and many variants overlap. Increase `maxRequests`, try `expandMode: both`, add recursion, or add the Cyrillic alphabet for Russian-language seeds.

**Are duplicates removed?**
Yes, by a normalized form (Unicode-folded, lowercase, single spaces). The row keeps every seed and region that produced the phrase.

**What does `hits` mean?**
The number of distinct queries that returned the phrase. A phrase caught by many variants is more strongly tied to the topic.

**What does `hashtagViews` null mean?**
Enrichment was off, no residential proxy was set, the candidate tag does not exist, or the lookup failed after retries.

**Is a TikTok account needed?**
No. No login, cookies or tokens.

**Can I run it on a schedule?**
Yes. Enable `compareWithPrevious`, keep the same `monitorKey` and schedule the Actor. Each run marks rows as new or existing.

**Do I need a proxy?**
Not for suggestions, scoring and clusters. Yes, a residential one, if you want hashtag stats.

### 🔌 Integrations

Send results on through the Apify API, webhooks, Zapier, Make, n8n or Google Sheets. The flat schema loads into a spreadsheet as is.

### 📝 Changelog

- **0.1** First release: A-Z / 0-9 expansion, recursion, raw signals, score, clusters, multi-region merge, monitoring, hashtag stats, oEmbed captions, HTML report.

### 🛟 Support

Something looks off or you need a field added? Open an issue from the Actor's **Issues** tab and include the run link and the input you used.

### 🔗 Related Actors

Check the author's profile on Apify Store for other keyword and social data Actors.

# Actor input Schema

## `keywords` (type: `array`):

Seed phrases to expand. Each seed is sent to TikTok search autocomplete. Up to 500 seeds; duplicates are removed.

## `regions` (type: `array`):

Two-letter region codes sent as the region parameter, e.g. US, GB, DE. Each region costs one extra request per query. Without a matching proxy this is regional labelling, not a guarantee of local results.

## `language` (type: `string`):

Value of the app\_language parameter. The language of the seed itself influences results far more than this setting.

## `includeSeedEcho` (type: `boolean`):

Keep suggestions that are identical to the query that produced them.

## `resultOrder` (type: `string`):

Order of the rows in the dataset.

## `expandMode` (type: `string`):

How each seed is expanded. 'none' sends one request per seed (10 suggestions). 'suffix' appends every letter/digit, 'prefix' prepends them, 'both' does both.

## `alphabets` (type: `array`):

Character sets used for expansion: latin (26), digits (10), cyrillic (33).

## `recursionDepth` (type: `integer`):

How many times discovered suggestions are fed back as new seeds. 0 disables recursion. Recursive seeds are queried as-is, without alphabet expansion.

## `recursionWidth` (type: `integer`):

How many top suggestions from each level become seeds for the next level.

## `maxRequests` (type: `integer`):

Hard ceiling on HTTP requests for the whole run. The main guard against a runaway bill.

## `maxResults` (type: `integer`):

Maximum number of unique suggestion rows pushed to the dataset.

## `maxSuggestionsPerKeyword` (type: `integer`):

Cap on unique phrases attributed to a single seed. Leave empty for no cap.

## `minScore` (type: `integer`):

Drop rows whose computed score is below this value.

## `onlyEcommerce` (type: `boolean`):

Keep only suggestions TikTok marks with a non-zero ecom\_intent signal.

## `includeTerms` (type: `array`):

Keep only suggestions containing at least one of these substrings (case-insensitive).

## `excludeTerms` (type: `array`):

Drop suggestions containing any of these substrings (case-insensitive).

## `enrichHashtags` (type: `boolean`):

Look up view and video counts for the hashtag each phrase maps to, using TikTok's public challenge endpoint. An explicit #tag wins; otherwise the phrase is squashed into one token the way TikTok forms hashtags ('skincare routine' becomes #skincareroutine). Costs one extra request per unique tag. In the cloud it needs Apify Proxy with the RESIDENTIAL group; even then only about half of the tags resolve, the rest stay empty.

## `maxHashtagLookups` (type: `integer`):

Ceiling on hashtag stat lookups per run.

## `videoUrls` (type: `array`):

TikTok video URLs to resolve via the public oEmbed endpoint. Returns the caption, author and thumbnail, plus every hashtag found in the caption, as extra rows of type videoCaption.

## `compareWithPrevious` (type: `boolean`):

Load the previous snapshot for this monitor key and mark each row as new or existing, with the score delta.

## `monitorKey` (type: `string`):

Name of the snapshot slot. Leave empty to derive it from the seeds, regions and language.

## `maxConcurrency` (type: `integer`):

Parallel in-flight requests.

## `maxRequestsPerSecond` (type: `number`):

Request rate cap. Automatically halved when TikTok answers with 429.

## `proxyConfiguration` (type: `object`):

Optional Apify Proxy. Needed for hashtag stats: TikTok blocks datacenter IPs on that endpoint, so use the RESIDENTIAL group when enrichHashtags is on. Also useful for country-specific results. Without a proxy the actor connects directly.

## Actor input object example

```json
{
  "keywords": [
    "skincare routine"
  ],
  "regions": [
    "US"
  ],
  "language": "en",
  "includeSeedEcho": false,
  "resultOrder": "score",
  "expandMode": "suffix",
  "alphabets": [
    "latin",
    "digits"
  ],
  "recursionDepth": 0,
  "recursionWidth": 10,
  "maxRequests": 40,
  "maxResults": 50,
  "minScore": 0,
  "onlyEcommerce": false,
  "includeTerms": [],
  "excludeTerms": [],
  "enrichHashtags": false,
  "maxHashtagLookups": 25,
  "videoUrls": [],
  "compareWithPrevious": false,
  "maxConcurrency": 8,
  "maxRequestsPerSecond": 8,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `report` (type: `string`):

Sortable and filterable table of every suggestion with clusters, score breakdown, monitoring block and CSV export.

## `suggestions` (type: `string`):

One row per unique suggestion, ordered as requested.

## `monitorDiff` (type: `string`):

New, gone and changed phrases compared with the previous snapshot. Present only when comparison is enabled.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "skincare routine"
    ],
    "regions": [
        "US"
    ],
    "maxRequests": 40,
    "maxResults": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("solalab_digital/tiktok-keywords-discovery").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["skincare routine"],
    "regions": ["US"],
    "maxRequests": 40,
    "maxResults": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("solalab_digital/tiktok-keywords-discovery").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "skincare routine"
  ],
  "regions": [
    "US"
  ],
  "maxRequests": 40,
  "maxResults": 50
}' |
apify call solalab_digital/tiktok-keywords-discovery --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,solalab_digital/tiktok-keywords-discovery"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xZCLq5kLJDVey6vMH/builds/wqNFVutu65ga3hWge/openapi.json
