TikTok Keywords Discovery avatar

TikTok Keywords Discovery

Pricing

from $0.20 / 1,000 results

Go to Apify Store
TikTok Keywords Discovery

TikTok Keywords Discovery

Expands seed keywords into TikTok search autocomplete suggestions with A-Z/0-9 expansion, recursion, raw TikTok signals, clustering and an HTML report.

Pricing

from $0.20 / 1,000 results

Rating

0.0

(0)

Developer

Sankov Vadim

Sankov Vadim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

TikTok Keywords Discovery: autocomplete expansion, scoring and clustering

Give it a seed phrase and it returns hundreds of real TikTok search suggestions. Each phrase carries the raw signals TikTok attaches to it, a 0-100 score, a topic cluster and intent tags. You can also add hashtag view counts and compare a run against the previous one. No login, no browser, no cookies.

You get A-Z / 0-9 expansion and recursion, raw TikTok signals, scores and clusters, multi-region merge, hashtag stats, change monitoring and an HTML report. It runs on 256 MB.

๐Ÿš€ Quick start

  1. Open the Actor and type one or more seed phrases into Seed keywords (for example skincare routine).
  2. Leave the defaults for a cheap trial (suffix expansion, up to 40 requests, up to 50 rows) and press Start.
  3. Open the Overview tab of the dataset, or open the REPORT file in the run's key-value store for a sortable HTML dashboard. Then raise maxRequests and maxResults for a full run.

๐Ÿงญ Who it is for

  • SEO and content marketers. Build a topic map for TikTok search, group phrases by cluster, and see which questions and "how to" queries people type.
  • TikTok creators. Find video ideas that TikTok itself offers under a topic, with a rough strength score and the hashtag each idea maps to.
  • E-commerce sellers and agencies. Filter phrases that carry TikTok's e-commerce intent flag, compare regions, and re-run the same list every month to see what is new.

โœจ What it does

A-Z and 0-9 expansion, plus recursion

TikTok returns at most 10 suggestions per query. To get past that ceiling the Actor appends (or prepends) every letter and digit to your seed and queries each variant. Pick suffix, prefix, both or none, and choose the alphabets: Latin (26), digits (10), Cyrillic (33).

Recursion feeds the best phrases of each level back in as new seeds, up to 3 levels deep. recursionWidth sets how many phrases per level are reused. Recursive seeds are queried as they are, without alphabet expansion, which keeps the request count predictable.

Every run is bounded by maxRequests (up to 5000) and maxResults (up to 20000).

Raw TikTok signals

Each row carries the fields TikTok returns next to the phrase, unmodified:

  • ecomIntent: TikTok's e-commerce intent flag
  • hotLevel: TikTok's "hot" flag
  • isTimeSensitive: marks phrases tied to a moment
  • predictCtrScore: TikTok's own predicted click-through signal
  • lang, recallReason, cutQuery: language, retrieval channel and TikTok's tokenization

Score 0-100

A transparent heuristic, not search volume:

base = 40*F + 20*P + 20*C + 10*R + 5*E + 5*H

  • F: how many distinct queries returned the phrase (log scale, maxes out at 16)
  • P: best position in the suggestion list (1st = full, 10th = zero)
  • C: predictCtrScore, maxes out at 0.08
  • R: share of your requested regions where the phrase appeared
  • E / H: 1 if ecomIntent / hotLevel is above zero

When hashtag stats are available, 5% of the score is swapped for reach: score = 0.95*base + 5*V, where V is log-scaled hashtag views (maxes out at 10 billion). The result is capped at 100 and rounded.

Clusters and intent tags

Phrases are grouped without any ML. Each phrase gets the token it shares with the most other phrases (seed words and common stopwords are ignored, ties are broken by longer token, then alphabetically). Phrases with no shared token land in other. Intent tags are rule-based: question, howto, for, vs, near_me, ecom.

Several regions in one run

Put up to 10 region codes in regions. The same phrase found in different regions is merged into one row with regions[] and regionCount, and the region share feeds the score.

Monitoring

Turn on compareWithPrevious and every row gets a monitor block: new or existing, the previous score and the change. A MONITOR_DIFF file lists new phrases, changed phrases (score moved by 10 or more) and phrases that disappeared. Disappeared phrases are never added to the dataset and are not charged. The snapshot is stored in a named store and is overwritten only after a clean, uninterrupted run.

Hashtag stats

Switch on enrichHashtags and set proxyConfiguration to Apify Proxy with the RESIDENTIAL group to add hashtag, hashtagViews and hashtagVideos. The hashtag is the explicit #tag if the phrase has one. Otherwise the phrase is collapsed into a single token, the way TikTok forms hashtags: skincare routine becomes #skincareroutine. That collapsed form is a candidate, so a tag that does not exist simply returns empty stats. maxHashtagLookups limits how many tags are looked up (default 25, max 500).

Video captions through oEmbed

Add TikTok video links to videoUrls and the Actor returns extra rows of type videoCaption: caption, author, thumbnail and every hashtag found in the caption.

HTML report

Each run saves a self-contained REPORT page: sortable table, filters (text, minimum score, cluster, region, e-commerce), top-phrase cards, a cluster chart, the monitoring block and a CSV export of whatever is filtered. No external scripts.

๐Ÿ†š Why use this Actor

You needWhat you get here
More than 10 phrases per seedA-Z / 0-9 expansion and recursion with hard request caps
To judge which phrases matterRaw TikTok signals plus a documented 0-100 score
To organise hundreds of phrasesClusters and intent tags on every row
Every seed that led to a phraseseeds[] keeps all of them, nothing is dropped as a duplicate
Regional comparisonregions[] merge in a single run
To track change over timeNamed-store snapshots and a diff file
A quick visual readBuilt-in HTML report
Low costPlain HTTP, no browser, 256 MB

โš™๏ธ Input

{
"keywords": ["skincare routine"],
"regions": ["US"],
"language": "en",
"expandMode": "suffix",
"alphabets": ["latin", "digits"],
"recursionDepth": 0,
"maxRequests": 40,
"maxResults": 50,
"enrichHashtags": true,
"compareWithPrevious": false
}
FieldTypeDefaultMeaning
keywordsstring[]["skincare routine"]Seed phrases, 1 to 500, duplicates removed. Required.
regionsstring[]["US"]Two-letter region codes, 1 to 10. Each region multiplies requests.
languageselectenSent as app_language: en, es, fr, de, it, pt, tr, id, ja, ko, ru.
includeSeedEchobooleanfalseKeep suggestions identical to the query that produced them.
resultOrderselectscorescore, source (discovery order) or alphabetical.
expandModeselectsuffixnone, suffix, prefix, both.
alphabetsselect[]latin, digitslatin, digits, cyrillic.
recursionDepthinteger00 to 3 levels of feeding results back as seeds.
recursionWidthinteger101 to 50 phrases reused per level.
maxRequestsinteger40Hard cap on requests per run, 1 to 5000.
maxResultsinteger50Cap on unique rows, 1 to 20000.
maxSuggestionsPerKeywordintegeremptyCap on unique phrases attributed to one seed.
minScoreinteger0Drop rows below this score.
onlyEcommercebooleanfalseKeep only rows with a non-zero ecomIntent.
includeTerms / excludeTermsstring[]emptyCase-insensitive substring filters.
enrichHashtagsbooleanfalseAdd hashtag views and video counts. Needs residential proxy in the cloud.
maxHashtagLookupsinteger25Ceiling for hashtag lookups, 1 to 500.
videoUrlsstring[]emptyVideo links to resolve into caption rows.
compareWithPreviousbooleanfalseCompare with the last snapshot for this key.
monitorKeystringderivedSnapshot slot name. Empty means it is derived from seeds, regions and language.
maxConcurrencyinteger8Parallel requests, 1 to 20.
maxRequestsPerSecondnumber8Rate cap, 0.5 to 25. Halved automatically after a 429.
proxyConfigurationproxyoffOptional Apify Proxy. Use RESIDENTIAL for hashtag stats.

Request math. Requests per seed = 1 + alphabet size (times two for both), multiplied by the number of regions. With the default Latin + digits alphabets that is 37 requests per seed per region for suffix. maxRequests always wins.

๐Ÿ“ฆ Output

One row per unique phrase. Here is a real row from a local test run (enrichHashtags and compareWithPrevious on). In the cloud the hashtag fields need a residential proxy:

{
"resultType": "keywordSuggestion",
"seedKeyword": "skincare routine",
"seeds": ["skincare routine"],
"suggestion": "skincare routine for oily skin",
"normalizedSuggestion": "skincare routine for oily skin",
"rank": 4,
"sourcePlatform": "tiktok",
"sourceSurface": "search_autocomplete",
"suggestionType": "sug",
"language": "en",
"region": "US",
"regions": ["US"],
"regionCount": 1,
"hits": 5,
"depth": 0,
"firstQuery": "skincare routine",
"ecomIntent": 1,
"hotLevel": 0,
"isTimeSensitive": 0,
"predictCtrScore": 0.018181765,
"lang": "en",
"recallReason": "tiktok_index_experience_decision_query|tiktok_index_active_7d_query|tiktok_orion_search_session|tiktok_experience_orion_query|tiktok_orion_query|tiktok_index_global_active_7d_query",
"cutQuery": ["skincare", "routine", "for", "oily", "skin"],
"hashtag": "skincareroutineforoilyskin",
"hashtagViews": 24783339,
"hashtagVideos": 1051,
"isSeedEcho": false,
"score": 59,
"cluster": "skin",
"intents": ["for"],
"monitor": { "status": "existing", "scorePrev": 66, "scoreDelta": -7 },
"hashtags": null,
"scrapedAt": "2026-09-26T20:40:18.443203+00:00"
}
GroupFields
Phrasesuggestion, normalizedSuggestion, rank (best position, 1-10), seedKeyword, seeds, firstQuery, depth, isSeedEcho
Localelanguage, region, regions, regionCount, lang
SignalsecomIntent, hotLevel, isTimeSensitive, predictCtrScore, recallReason, cutQuery, hits
Analysisscore, cluster, intents
Hashtaghashtag, hashtagViews, hashtagVideos, hashtags (tags found inside the text)
Monitoringmonitor (status, scorePrev, scoreDelta), null when comparison is off
Video rowsvideoUrl, caption, authorName, authorUrl, thumbnailUrl, hashtags (resultType: videoCaption)

Dataset views: Overview, Raw TikTok signals, Hashtags, Video captions. Export as JSON, CSV, Excel, XML, RSS or HTML from the dataset tab, or read it through the Apify API.

โš ๏ธ Limits, stated plainly

  • Unofficial endpoint. The Actor reads the public autocomplete request that TikTok's own search box uses. It is not an official API. TikTok can change or restrict it at any time, and the Actor may return fewer rows or fail if that happens.
  • No search volume. There is no volume, CPC or trend data. The score is a heuristic built from the signals above.
  • Region is labelling, not local results. The region value is sent as a parameter. Without a proxy in the matching country it does not guarantee results as a local user would see them. The language of the seed itself changes results far more than the region setting.
  • Hashtag stats need a residential proxy and are partial. TikTok returns empty answers to datacenter IPs on this endpoint. In our cloud test, without a proxy and with the datacenter group, 0 of 25 tags resolved. With Apify Proxy group RESIDENTIAL, 13 of 25 did, and the rest stayed empty. Expect that order of magnitude, not full coverage, and about $0.001 of proxy traffic per 25 lookups. Parallel lookups can also trigger 429 responses; the Actor halves the request rate, waits for Retry-After when given, and retries up to 3 times. A lookup that still fails leaves hashtagViews and hashtagVideos as null instead of failing the run. The hashtag value is a candidate and may not exist as a real tag.
  • Expansion and recursion are capped. maxRequests, maxResults, recursion depth (max 3) and width bound every run, so you may not see every phrase TikTok knows.
  • Rate ceiling unknown. TikTok does not publish one. Lower maxRequestsPerSecond or add a proxy for very large runs.
  • Partial results are kept. On abort, migration or timeout the Actor saves what it has collected.

๐Ÿ’ณ Pricing

Pay per event. See the Pricing tab of this Actor for the current event prices. A row is charged when it is added to the dataset, and the run stops cleanly when your spending limit is reached. Score, clusters, monitoring and hashtag stats cost nothing extra. videoCaption rows are dataset rows like any other.

โ“ FAQ

Does it give search volume? No. TikTok does not expose it in this endpoint. Use score, hits and the raw signals to compare phrases against each other.

Why fewer rows than expected? Each query returns at most 10 phrases, and many variants overlap. Increase maxRequests, try expandMode: both, add recursion, or add the Cyrillic alphabet for Russian-language seeds.

Are duplicates removed? Yes, by a normalized form (Unicode-folded, lowercase, single spaces). The row keeps every seed and region that produced the phrase.

What does hits mean? The number of distinct queries that returned the phrase. A phrase caught by many variants is more strongly tied to the topic.

What does hashtagViews null mean? Enrichment was off, no residential proxy was set, the candidate tag does not exist, or the lookup failed after retries.

Is a TikTok account needed? No. No login, cookies or tokens.

Can I run it on a schedule? Yes. Enable compareWithPrevious, keep the same monitorKey and schedule the Actor. Each run marks rows as new or existing.

Do I need a proxy? Not for suggestions, scoring and clusters. Yes, a residential one, if you want hashtag stats.

๐Ÿ”Œ Integrations

Send results on through the Apify API, webhooks, Zapier, Make, n8n or Google Sheets. The flat schema loads into a spreadsheet as is.

๐Ÿ“ Changelog

  • 0.1 First release: A-Z / 0-9 expansion, recursion, raw signals, score, clusters, multi-region merge, monitoring, hashtag stats, oEmbed captions, HTML report.

๐Ÿ›Ÿ Support

Something looks off or you need a field added? Open an issue from the Actor's Issues tab and include the run link and the input you used.

Check the author's profile on Apify Store for other keyword and social data Actors.