Threads Search Scraper avatar

Threads Search Scraper

Pricing

from $7.00 / 1,000 results

Go to Apify Store
Threads Search Scraper

Threads Search Scraper

Search public Threads posts by keyword and get post text, author, timestamp, like, reply, repost and quote counts and media URLs. Public search shows about 20-25 posts per keyword. Only-new mode remembers seen post IDs for monitoring. No Threads account or API key needed.

Pricing

from $7.00 / 1,000 results

Rating

0.0

(0)

Developer

Axiom Works

Axiom Works

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 hours ago

Last modified

Categories

Share

What does Threads Search Scraper do?

Threads Search Scraper collects public posts shown on Threads keyword search pages. Give it one or more words or phrases, and it returns the posts visible to a logged-out visitor for each search. Each result includes its search term, a stable post ID, the public post URL, text, author, creation time, engagement counts, and available media URLs. The Actor can remember post IDs between runs so that a scheduled search returns only posts it has not previously emitted for that term.

Typical uses: brand and competitor monitoring, tracking a hashtag or launch, and feeding Threads mentions into Slack or a sheet on a schedule with Only new posts turned on.

The Actor reads the public search page and its embedded data. It does not use a Threads account, cookies, the Meta Graph API, or an app review. It retrieves one logged-out Top results page for each keyword. That page normally contains roughly 20 to 25 posts, and Threads currently does not expose an additional logged-out results page for this method. Results are therefore a view of the available Top page at run time, not a complete archive or a reliable chronological feed. Run multiple relevant searches or schedule repeated runs when you need broader monitoring.

A run performs searches at a modest pace, about one page every two seconds, and retries temporary failures with a new proxy session. The default proxy setting uses a US residential Apify Proxy connection. You can change that setting in the input. Media URLs are returned as strings; the Actor does not fetch or save the media files.

What data can you get?

Each dataset item is one post. The fields below describe the data returned by the current version. Counts and optional media or reply fields reflect what the public search page provides at the time of the run. A null value means the page did not contain that value for the post.

FieldMeaning
idStable post identifier, also used to remove duplicates within a run.
sourceUrlPublic search page from which the post was read.
queryInput keyword or phrase that produced this item.
postIdThreads post identifier.
codeShort code used in the public post URL.
urlPublic Threads post URL.
textCaption text, or joined text fragments when the caption is absent.
createdAtPost creation time in ISO 8601 format.
author.usernameAccount username shown with the post.
author.fullNameDisplay name when supplied.
author.idAccount identifier supplied in the page data.
author.isVerifiedVerification flag supplied with the account.
author.profilePicUrlProfile picture URL when supplied.
likeCountLike count visible in the page data.
replyCountDirect reply count visible in the page data.
repostCountRepost count visible in the page data.
quoteCountQuote count visible in the page data.
reshareCountShare count visible in the page data. Threads leaves it out when a post has no shares yet, so the Actor reports 0 in that case.
isReplyWhether Threads marks the item as a reply.
replyToUsernameUsername being replied to, when present.
mediaTypeNumeric media type supplied by Threads: 1 = image, 2 = video, 8 = carousel, 19 = text only.
mediaKindReadable media kind derived from mediaType: text, image, video or carousel.
imageUrlsImage candidate URLs for the post and its carousel items.
videoUrlFirst video URL, when supplied.
threadIdIdentifier of the containing thread.
positionInThreadZero-based position among items in that thread. threadId and positionInThread are null for a reply shown on its own, where the containing thread is unknown.
scrapedAtTime the Actor mapped the item.

id and postId contain the same Threads post identifier. The extra id makes it straightforward to deduplicate or update records in downstream systems. A post that matches two keywords is emitted only once per run; its query and sourceUrl reflect the first matching search processed. Engagement counts may change later, so a saved item is a snapshot of the public page rather than a live count.

How to use

Enter one or more terms in queries and start the Actor. A simple term such as apify is enough for a first run. Use phrases such as web scraping for more focused searches. The Actor visits one public search page per term in the order provided. It emits posts until each term's maxItemsPerQuery limit is met, the available page ends, or the optional overall maxItems limit is reached.

For monitoring, create an Apify schedule and turn on onlyNew. Keep monitorStoreName stable across scheduled runs. The named key-value store holds up to 5,000 previously emitted post IDs per query. The Actor checks those IDs before pushing new items, then updates the store after processing the query. A different store name starts a separate history. Removing the store also resets that history. The store is per query, so changing a keyword starts a distinct seen list.

To narrow results, set maxAgeHours to exclude older posts, or set includeReplies to false to return only non-replies. These filters apply to posts already present on the public Top search page; they cannot cause Threads to reveal additional pages. An empty dataset can be a valid outcome for a narrow filter or for a monitor run with no unseen posts. If the public search data is missing or the page redirects to a login screen, the run fails with an error after bounded retries rather than reporting an empty successful search.

Input

FieldTypeBehavior
queriesArray of stringsRequired; 1 to 100 nonempty search terms.
maxItemsIntegerOverall result cap, from 1 to 10,000; when omitted, the available results determine the total.
maxItemsPerQueryIntegerPer-term cap from 1 to 100; default 50 (Threads returns about 20-25 per term).
onlyNewBooleanSkip IDs previously emitted for the same term; default false.
monitorStoreNameStringNamed key-value store for seen IDs; default threads-keyword-monitor.
maxAgeHoursIntegerOptional maximum post age in hours, 1 to 876000.
includeRepliesBooleanInclude items marked as replies; default true.
proxyConfigurationObjectApify Proxy configuration; defaults to US residential.

For example, this input searches two phrases, returns at most seven posts total, and excludes replies. If at least seven matching posts are visible, the dataset contains exactly seven items. maxItems takes priority over the per-query cap when the total limit is reached.

{
"queries": ["climate change", "web scraping"],
"maxItems": 7,
"maxItemsPerQuery": 20,
"includeReplies": false,
"onlyNew": false
}

The input schema pre-fills queries with apify for Console trials, but an API caller must supply queries. Invalid or empty terms and out-of-range numeric limits fail immediately with a clear input error. The proxy object uses the standard Apify Proxy input format. The default is intended for public search access; if your environment supports direct requests, set proxyConfiguration.useApifyProxy to false.

Output

The output is a dataset of post objects. The example below is a real item from a run of {"queries":["apify"],"maxItems":7} on September 30, 2026. Public post content and counts may change, and temporary media URLs can expire. This example has no media URL; other records can include images, video, or both.

{
"id": "3894484011931302948",
"sourceUrl": "https://www.threads.com/search?q=apify&serp_type=default",
"query": "apify",
"postId": "3894484011931302948",
"code": "DYL_JMymnwk",
"url": "https://www.threads.com/@yohanolo.gy/post/DYL_JMymnwk",
"text": "using apify in claude is a superpower",
"createdAt": "2026-05-11T05:51:33.000Z",
"author": {
"username": "yohanolo.gy",
"fullName": "Yohan",
"id": "65866699937",
"isVerified": false,
"profilePicUrl": "https://instagram.fict1-1.fna.fbcdn.net/v/t51.82787-19/626553192_17929253136195938_76617471915461927_n.jpg?stp=dst-jpg_s150x150_tt6&efg=eyJ2ZW5jb2RlX3RhZyI6InByb2ZpbGVfcGljLmRqYW5nby43NzQuYzIifQ&_nc_ht=instagram.fict1-1.fna.fbcdn.net&_nc_cat=102&_nc_oc=Q6cZ2gHfnm-jnOcrCeap5pk7PBAy-FesODSA6D3rSgm_i6tUVwtTSCG6y4mxnxnxUBLvHHc&_nc_ohc=egHkruEGbpcQ7kNvwFV3cqH&_nc_gid=5OfDi6C1Aa-7klOhzmHO3w&edm=APs17CUBAAAA&ccb=7-5&oh=00_AQO1sfVqzsu5GuojTlJ_3VgbxvUgANFmqfRMsrL_oT5JKg&oe=6AC27D11&_nc_sid=10d13b"
},
"likeCount": 2,
"replyCount": 1,
"repostCount": 0,
"quoteCount": 0,
"reshareCount": 0,
"isReply": false,
"replyToUsername": null,
"mediaType": 19,
"mediaKind": "text",
"imageUrls": [],
"videoUrl": null,
"threadId": "3894484011931302948",
"positionInThread": 0,
"scrapedAt": "2026-09-30T05:43:18.602Z"
}

The sourceUrl names the search page actually requested. The url field is the public post link derived from the username and code in the source object. The author details are returned in the nested author object. Optional values are represented by null or an empty array when absent. Timestamps are normalized to UTC ISO strings; integer engagement fields are kept numeric. The default dataset can be downloaded as JSON, CSV, or another format supported by Apify.

How much does it cost?

The Actor charges one result event for each post pushed to the dataset. Runs that return fewer posts trigger fewer result events. For the current monetary rate per event, consult the Actor's Pricing tab before starting a run. Infrastructure, proxy, and platform charges may depend on your Apify account and run settings.

The search page itself limits how many posts can be returned for a term. Setting a large maxItemsPerQuery cannot make the logged-out page provide more posts. A scheduled onlyNew run that finds nothing new pushes no result items.

Use with the API

You can use Apify integrations to export a dataset or start this Actor on a schedule. The snippets below show direct API usage. Supply your own Apify token through your application environment and keep it out of code repositories.

Python with apify-client:

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("axiomworks/threads-keyword-search-monitor").call(
run_input={"queries": ["apify"], "maxItems": 7}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["createdAt"], item["author"]["username"], item["url"])

JavaScript with apify-client:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('axiomworks/threads-keyword-search-monitor').call({
queries: ['apify'],
maxItems: 7,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) {
process.stdout.write(`${item.createdAt} ${item.url}\n`);
}

A cURL request can start a synchronous run and return dataset items as JSON:

curl -X POST \
"https://api.apify.com/v2/acts/axiomworks~threads-keyword-search-monitor/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"queries":["apify"],"maxItems":7}'

For recurring monitoring, use a schedule with the same monitorStoreName and onlyNew: true. If two runs overlap, they may both see an ID before either updates the store; avoid overlapping schedules when strict cross-run deduplication matters. Downstream consumers can also upsert on the stable id field for an additional safeguard.

Apify integrations with Zapier, Make, Google Sheets and webhooks can run this Actor on a schedule and send results on.

Use with AI agents (MCP)

AI agents can invoke the Actor through Apify MCP and receive structured items whose fields are defined in the dataset schema. Configure an MCP client with the Apify MCP server and restrict the available tool to this Actor:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com/?tools=axiomworks/threads-keyword-search-monitor"
}
}
}

A prompt can be specific about the terms and limit: “Search public Threads posts for climate change and return the first seven posts with author, date, text, and URL.” Another useful prompt is: “Run my Threads keyword monitor for web scraping and summarize any newly returned posts, citing their post URLs.” Supply the required queries input when calling through MCP; the Console prefill is only a suggestion for manual runs. When the agent summarizes a result, it should use the url or sourceUrl fields as evidence and treat the post content as user-generated data.

The public search page may contain promotional or unrelated posts. An agent should inspect each record's text, query, and date before drawing a conclusion. It should not assume the dataset is exhaustive or ordered by publication date. If the Actor reports an error because the search page changed, the agent should report that failure rather than presenting a zero result count as a factual claim about Threads.

FAQ

Does this require a Threads login or Meta app review? No. It reads the public search page available to a logged-out visitor. It does not use private account data.

Can it search the full Threads archive? No. The logged-out Top results page is one limited page per term. The Actor does not request authenticated GraphQL pages or offer a Recent sort that the public page does not expose.

Why might a run return fewer items than maxItems? There may be fewer visible Top results, duplicate posts, posts filtered by age or reply status, or posts already recorded by onlyNew. maxItems is an upper bound, not a promise of available content.

Are media files saved? No. Image, avatar, and video URLs are copied from public page data. Some media URLs are signed and can expire, so download or process them promptly if your use case requires them and you have the right to do so.

Can I reset monitoring history? Yes. Change monitorStoreName or delete its stored seen-ID records. Keep the name unchanged when you want successive runs to share the same history.

What happens if Threads changes its page? The Actor searches the embedded JSON recursively for search results. If that structure disappears, it retries and then fails clearly. A page change may require an Actor update.

The Actor retrieves publicly visible posts without logging in. Threads' robots rules disallow automated access, and its terms restrict collection. The Actor owner has chosen to offer this public-page workflow despite those restrictions. Whether a particular use is lawful depends on your location, purpose, contracts, and handling of personal information. Assess your own obligations before collecting or publishing post text, usernames, names, or profile photos. If you process data about people in a jurisdiction with privacy rules such as GDPR, apply those rules to your storage, retention, and downstream use. This description is general information, not legal advice.

This Actor is independently developed and is not affiliated with Threads or Meta. Results come from the public search page and may be incomplete, stale, or removed by the platform. Respect individual privacy and avoid treating public availability as permission for every downstream use.

Feedback

If a search fails, include the query, the approximate run time, and the error message when reporting it through the Actor's Apify feedback channel. Do not include API tokens or private data. Reports of missing fields are most useful when they identify the affected public post URL and the field that was expected. Feedback helps distinguish a temporary search failure from a change in the embedded page structure.