# Threads Search Scraper (`axiomworks/threads-keyword-search-monitor`) Actor

Search public Threads posts by keyword and get post text, author, timestamp, like, reply, repost and quote counts and media URLs. Public search shows about 20-25 posts per keyword. Only-new mode remembers seen post IDs for monitoring. No Threads account or API key needed.

- **URL**: https://apify.com/axiomworks/threads-keyword-search-monitor.md
- **Developed by:** [Axiom Works](https://apify.com/axiomworks) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $7.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Threads Search Scraper

### What does Threads Search Scraper do?

Threads Search Scraper collects public posts shown on Threads keyword search pages. Give it one or more words or phrases, and it returns the posts visible to a logged-out visitor for each search. Each result includes its search term, a stable post ID, the public post URL, text, author, creation time, engagement counts, and available media URLs. The Actor can remember post IDs between runs so that a scheduled search returns only posts it has not previously emitted for that term.

Typical uses: brand and competitor monitoring, tracking a hashtag or launch, and feeding Threads mentions into Slack or a sheet on a schedule with Only new posts turned on.

The Actor reads the public search page and its embedded data. It does not use a Threads account, cookies, the Meta Graph API, or an app review. It retrieves one logged-out Top results page for each keyword. That page normally contains roughly 20 to 25 posts, and Threads currently does not expose an additional logged-out results page for this method. Results are therefore a view of the available Top page at run time, not a complete archive or a reliable chronological feed. Run multiple relevant searches or schedule repeated runs when you need broader monitoring.

A run performs searches at a modest pace, about one page every two seconds, and retries temporary failures with a new proxy session. The default proxy setting uses a US residential Apify Proxy connection. You can change that setting in the input. Media URLs are returned as strings; the Actor does not fetch or save the media files.

### What data can you get?

Each dataset item is one post. The fields below describe the data returned by the current version. Counts and optional media or reply fields reflect what the public search page provides at the time of the run. A `null` value means the page did not contain that value for the post.

| Field | Meaning |
| --- | --- |
| `id` | Stable post identifier, also used to remove duplicates within a run. |
| `sourceUrl` | Public search page from which the post was read. |
| `query` | Input keyword or phrase that produced this item. |
| `postId` | Threads post identifier. |
| `code` | Short code used in the public post URL. |
| `url` | Public Threads post URL. |
| `text` | Caption text, or joined text fragments when the caption is absent. |
| `createdAt` | Post creation time in ISO 8601 format. |
| `author.username` | Account username shown with the post. |
| `author.fullName` | Display name when supplied. |
| `author.id` | Account identifier supplied in the page data. |
| `author.isVerified` | Verification flag supplied with the account. |
| `author.profilePicUrl` | Profile picture URL when supplied. |
| `likeCount` | Like count visible in the page data. |
| `replyCount` | Direct reply count visible in the page data. |
| `repostCount` | Repost count visible in the page data. |
| `quoteCount` | Quote count visible in the page data. |
| `reshareCount` | Share count visible in the page data. Threads leaves it out when a post has no shares yet, so the Actor reports 0 in that case. |
| `isReply` | Whether Threads marks the item as a reply. |
| `replyToUsername` | Username being replied to, when present. |
| `mediaType` | Numeric media type supplied by Threads: 1 = image, 2 = video, 8 = carousel, 19 = text only. |
| `mediaKind` | Readable media kind derived from `mediaType`: `text`, `image`, `video` or `carousel`. |
| `imageUrls` | Image candidate URLs for the post and its carousel items. |
| `videoUrl` | First video URL, when supplied. |
| `threadId` | Identifier of the containing thread. |
| `positionInThread` | Zero-based position among items in that thread. `threadId` and `positionInThread` are `null` for a reply shown on its own, where the containing thread is unknown. |
| `scrapedAt` | Time the Actor mapped the item. |

`id` and `postId` contain the same Threads post identifier. The extra `id` makes it straightforward to deduplicate or update records in downstream systems. A post that matches two keywords is emitted only once per run; its `query` and `sourceUrl` reflect the first matching search processed. Engagement counts may change later, so a saved item is a snapshot of the public page rather than a live count.

### How to use

Enter one or more terms in `queries` and start the Actor. A simple term such as `apify` is enough for a first run. Use phrases such as `web scraping` for more focused searches. The Actor visits one public search page per term in the order provided. It emits posts until each term's `maxItemsPerQuery` limit is met, the available page ends, or the optional overall `maxItems` limit is reached.

For monitoring, create an Apify schedule and turn on `onlyNew`. Keep `monitorStoreName` stable across scheduled runs. The named key-value store holds up to 5,000 previously emitted post IDs per query. The Actor checks those IDs before pushing new items, then updates the store after processing the query. A different store name starts a separate history. Removing the store also resets that history. The store is per query, so changing a keyword starts a distinct seen list.

To narrow results, set `maxAgeHours` to exclude older posts, or set `includeReplies` to `false` to return only non-replies. These filters apply to posts already present on the public Top search page; they cannot cause Threads to reveal additional pages. An empty dataset can be a valid outcome for a narrow filter or for a monitor run with no unseen posts. If the public search data is missing or the page redirects to a login screen, the run fails with an error after bounded retries rather than reporting an empty successful search.

### Input

| Field | Type | Behavior |
| --- | --- | --- |
| `queries` | Array of strings | Required; 1 to 100 nonempty search terms. |
| `maxItems` | Integer | Overall result cap, from 1 to 10,000; when omitted, the available results determine the total. |
| `maxItemsPerQuery` | Integer | Per-term cap from 1 to 100; default 50 (Threads returns about 20-25 per term). |
| `onlyNew` | Boolean | Skip IDs previously emitted for the same term; default `false`. |
| `monitorStoreName` | String | Named key-value store for seen IDs; default `threads-keyword-monitor`. |
| `maxAgeHours` | Integer | Optional maximum post age in hours, 1 to 876000. |
| `includeReplies` | Boolean | Include items marked as replies; default `true`. |
| `proxyConfiguration` | Object | Apify Proxy configuration; defaults to US residential. |

For example, this input searches two phrases, returns at most seven posts total, and excludes replies. If at least seven matching posts are visible, the dataset contains exactly seven items. `maxItems` takes priority over the per-query cap when the total limit is reached.

```json
{
  "queries": ["climate change", "web scraping"],
  "maxItems": 7,
  "maxItemsPerQuery": 20,
  "includeReplies": false,
  "onlyNew": false
}
```

The input schema pre-fills `queries` with `apify` for Console trials, but an API caller must supply `queries`. Invalid or empty terms and out-of-range numeric limits fail immediately with a clear input error. The proxy object uses the standard Apify Proxy input format. The default is intended for public search access; if your environment supports direct requests, set `proxyConfiguration.useApifyProxy` to `false`.

### Output

The output is a dataset of post objects. The example below is a real item from a run of `{"queries":["apify"],"maxItems":7}` on September 30, 2026. Public post content and counts may change, and temporary media URLs can expire. This example has no media URL; other records can include images, video, or both.

```json
{
  "id": "3894484011931302948",
  "sourceUrl": "https://www.threads.com/search?q=apify&serp_type=default",
  "query": "apify",
  "postId": "3894484011931302948",
  "code": "DYL_JMymnwk",
  "url": "https://www.threads.com/@yohanolo.gy/post/DYL_JMymnwk",
  "text": "using apify in claude is a superpower",
  "createdAt": "2026-05-11T05:51:33.000Z",
  "author": {
    "username": "yohanolo.gy",
    "fullName": "Yohan",
    "id": "65866699937",
    "isVerified": false,
    "profilePicUrl": "https://instagram.fict1-1.fna.fbcdn.net/v/t51.82787-19/626553192_17929253136195938_76617471915461927_n.jpg?stp=dst-jpg_s150x150_tt6&efg=eyJ2ZW5jb2RlX3RhZyI6InByb2ZpbGVfcGljLmRqYW5nby43NzQuYzIifQ&_nc_ht=instagram.fict1-1.fna.fbcdn.net&_nc_cat=102&_nc_oc=Q6cZ2gHfnm-jnOcrCeap5pk7PBAy-FesODSA6D3rSgm_i6tUVwtTSCG6y4mxnxnxUBLvHHc&_nc_ohc=egHkruEGbpcQ7kNvwFV3cqH&_nc_gid=5OfDi6C1Aa-7klOhzmHO3w&edm=APs17CUBAAAA&ccb=7-5&oh=00_AQO1sfVqzsu5GuojTlJ_3VgbxvUgANFmqfRMsrL_oT5JKg&oe=6AC27D11&_nc_sid=10d13b"
  },
  "likeCount": 2,
  "replyCount": 1,
  "repostCount": 0,
  "quoteCount": 0,
  "reshareCount": 0,
  "isReply": false,
  "replyToUsername": null,
  "mediaType": 19,
  "mediaKind": "text",
  "imageUrls": [],
  "videoUrl": null,
  "threadId": "3894484011931302948",
  "positionInThread": 0,
  "scrapedAt": "2026-09-30T05:43:18.602Z"
}
```

The `sourceUrl` names the search page actually requested. The `url` field is the public post link derived from the username and code in the source object. The author details are returned in the nested `author` object. Optional values are represented by `null` or an empty array when absent. Timestamps are normalized to UTC ISO strings; integer engagement fields are kept numeric. The default dataset can be downloaded as JSON, CSV, or another format supported by Apify.

### How much does it cost?

The Actor charges one `result` event for each post pushed to the dataset. Runs that return fewer posts trigger fewer result events. For the current monetary rate per event, consult the Actor's Pricing tab before starting a run. Infrastructure, proxy, and platform charges may depend on your Apify account and run settings.

The search page itself limits how many posts can be returned for a term. Setting a large `maxItemsPerQuery` cannot make the logged-out page provide more posts. A scheduled `onlyNew` run that finds nothing new pushes no result items.

### Use with the API

You can use Apify integrations to export a dataset or start this Actor on a schedule. The snippets below show direct API usage. Supply your own Apify token through your application environment and keep it out of code repositories.

Python with `apify-client`:

```python
import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("axiomworks/threads-keyword-search-monitor").call(
    run_input={"queries": ["apify"], "maxItems": 7}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["createdAt"], item["author"]["username"], item["url"])
```

JavaScript with `apify-client`:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('axiomworks/threads-keyword-search-monitor').call({
    queries: ['apify'],
    maxItems: 7,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) {
    process.stdout.write(`${item.createdAt} ${item.url}\n`);
}
```

A cURL request can start a synchronous run and return dataset items as JSON:

```bash
curl -X POST \
  "https://api.apify.com/v2/acts/axiomworks~threads-keyword-search-monitor/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"queries":["apify"],"maxItems":7}'
```

For recurring monitoring, use a schedule with the same `monitorStoreName` and `onlyNew: true`. If two runs overlap, they may both see an ID before either updates the store; avoid overlapping schedules when strict cross-run deduplication matters. Downstream consumers can also upsert on the stable `id` field for an additional safeguard.

Apify integrations with Zapier, Make, Google Sheets and webhooks can run this Actor on a schedule and send results on.

### Use with AI agents (MCP)

AI agents can invoke the Actor through Apify MCP and receive structured items whose fields are defined in the dataset schema. Configure an MCP client with the Apify MCP server and restrict the available tool to this Actor:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com/?tools=axiomworks/threads-keyword-search-monitor"
    }
  }
}
```

A prompt can be specific about the terms and limit: “Search public Threads posts for `climate change` and return the first seven posts with author, date, text, and URL.” Another useful prompt is: “Run my Threads keyword monitor for `web scraping` and summarize any newly returned posts, citing their post URLs.” Supply the required `queries` input when calling through MCP; the Console prefill is only a suggestion for manual runs. When the agent summarizes a result, it should use the `url` or `sourceUrl` fields as evidence and treat the post content as user-generated data.

The public search page may contain promotional or unrelated posts. An agent should inspect each record's `text`, `query`, and date before drawing a conclusion. It should not assume the dataset is exhaustive or ordered by publication date. If the Actor reports an error because the search page changed, the agent should report that failure rather than presenting a zero result count as a factual claim about Threads.

### FAQ

**Does this require a Threads login or Meta app review?** No. It reads the public search page available to a logged-out visitor. It does not use private account data.

**Can it search the full Threads archive?** No. The logged-out Top results page is one limited page per term. The Actor does not request authenticated GraphQL pages or offer a Recent sort that the public page does not expose.

**Why might a run return fewer items than `maxItems`?** There may be fewer visible Top results, duplicate posts, posts filtered by age or reply status, or posts already recorded by `onlyNew`. `maxItems` is an upper bound, not a promise of available content.

**Are media files saved?** No. Image, avatar, and video URLs are copied from public page data. Some media URLs are signed and can expire, so download or process them promptly if your use case requires them and you have the right to do so.

**Can I reset monitoring history?** Yes. Change `monitorStoreName` or delete its stored seen-ID records. Keep the name unchanged when you want successive runs to share the same history.

**What happens if Threads changes its page?** The Actor searches the embedded JSON recursively for search results. If that structure disappears, it retries and then fails clearly. A page change may require an Actor update.

### Is it legal to scrape Threads?

The Actor retrieves publicly visible posts without logging in. Threads' robots rules disallow automated access, and its terms restrict collection. The Actor owner has chosen to offer this public-page workflow despite those restrictions. Whether a particular use is lawful depends on your location, purpose, contracts, and handling of personal information. Assess your own obligations before collecting or publishing post text, usernames, names, or profile photos. If you process data about people in a jurisdiction with privacy rules such as GDPR, apply those rules to your storage, retention, and downstream use. This description is general information, not legal advice.

This Actor is independently developed and is not affiliated with Threads or Meta. Results come from the public search page and may be incomplete, stale, or removed by the platform. Respect individual privacy and avoid treating public availability as permission for every downstream use.

### Feedback

If a search fails, include the query, the approximate run time, and the error message when reporting it through the Actor's Apify feedback channel. Do not include API tokens or private data. Reports of missing fields are most useful when they identify the affected public post URL and the field that was expected. Feedback helps distinguish a temporary search failure from a change in the embedded page structure.

# Actor input Schema

## `queries` (type: `array`):

Keywords or phrases to search on public Threads, such as apify or web scraping. Supply 1-100 nonempty terms.

## `maxItems` (type: `integer`):

Maximum posts to return across all keywords, 1-10000. Leave empty for no overall cap.

## `maxItemsPerQuery` (type: `integer`):

Maximum posts per term, 1-100; default 50. Threads shows about 20-25 top posts per term to logged-out visitors.

## `onlyNew` (type: `boolean`):

Return only posts unseen for each term in the named monitor store; default false.

## `monitorStoreName` (type: `string`):

Name of the Apify key-value store that remembers up to 5000 seen post IDs per keyword, e.g. threads-keyword-monitor (letters, digits, hyphens; default threads-keyword-monitor).

## `maxAgeHours` (type: `integer`):

Exclude posts older than this number of hours, 1-876000; omit to include all available ages.

## `includeReplies` (type: `boolean`):

Include posts marked as replies; default true.

## `proxyConfiguration` (type: `object`):

Choose Apify Proxy for public Threads pages; defaults to US residential proxy. Disable it for direct requests.

## Actor input object example

```json
{
  "queries": [
    "apify"
  ],
  "maxItemsPerQuery": 20,
  "onlyNew": false,
  "monitorStoreName": "threads-keyword-monitor",
  "includeReplies": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

Public Threads posts

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "apify"
    ],
    "maxItemsPerQuery": 20,
    "includeReplies": true,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiomworks/threads-keyword-search-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["apify"],
    "maxItemsPerQuery": 20,
    "includeReplies": True,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("axiomworks/threads-keyword-search-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "apify"
  ],
  "maxItemsPerQuery": 20,
  "includeReplies": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call axiomworks/threads-keyword-search-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiomworks/threads-keyword-search-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/V7nT3Q19d70bJhTwI/builds/h1ECZdzZPN220kHih/openapi.json
