# Threads Keyword Search Scraper - Search Posts by Keyword (`headply/threads-keyword-search-scraper`) Actor

Search Meta Threads by keyword and export every matching public post: text, author, likes, replies, reposts and media. No login, no weekly quota.

- **URL**: https://apify.com/headply/threads-keyword-search-scraper.md
- **Developed by:** [Mayowa Ogedengbe](https://apify.com/headply) (community)
- **Categories:** Social media, Marketing, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Threads Keyword Search Scraper

Search [Threads](https://www.threads.com) (Meta) by keyword and export every matching public post as structured JSON.

**No login. No cookies. No weekly search quota.**

***

### The problem this solves

Meta's official Threads API allows roughly **500 keyword searches per rolling seven days**. Social listening burns through that in an afternoon. Brand monitoring across a handful of terms burns through it faster.

This Actor reads Threads search the way a logged-out visitor does. There is no weekly cap, nothing to authenticate, and no account pool to keep alive.

***

### What you get

For every keyword you supply, every matching public post with:

- Full post text, language, and the canonical URL
- Author handle, display name, verification status and profile picture
- Likes, replies, reposts, quotes and reshares
- Images, video and carousels, at the largest available resolution
- Link previews, and quoted posts resolved inline
- The keyword that found it, in `searchQuery`, so one run can track many terms

Threads offers two orderings, by relevance and by recency. They return overlapping but different sets. Choose **Both** for the widest coverage.

***

### Input

```json
{
  "searchQueries": ["your brand", "competitor brand", "category term"],
  "searchSort": "both",
  "maxPostsPerQuery": 200,
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
```

**Running with no input.** If you start a run without setting anything, the Actor returns a small live sample so you can see the output shape immediately. Set your own keywords, profiles or post URLs for a real run.

The same run can optionally take `profiles` and `postUrls` as well, if you want profile records or reply trees alongside your keyword results.

***

### Output

```json
{
  "type": "post",
  "id": "3981852126213720917_63055343223",
  "pk": "3981852126213720917",
  "code": "DdCYWl7GktV",
  "url": "https://www.threads.com/@zuck/post/DdCYWl7GktV",
  "text": "Mostly superintelligence and MMA takes",
  "language": "en",
  "createdAt": "2026-09-08T14:56:24.000Z",
  "likeCount": 2715,
  "replyCount": 1602,
  "repostCount": 231,
  "quoteCount": 142,
  "isQuotePost": false,
  "quotedPost": null,
  "mediaType": "text",
  "media": [],
  "author": {
    "username": "zuck",
    "fullName": "Mark Zuckerberg",
    "isVerified": true,
    "profileUrl": "https://www.threads.com/@zuck"
  },
  "authorUsername": "zuck",
  "source": "search",
  "searchQuery": "superintelligence",
  "scrapedAt": "2026-09-10T04:31:00.000Z"
}
```

Three details that matter downstream:

- **`authorUsername` is the flat copy of `author.username`.** Spreadsheet and table views cannot read nested paths, so use the flat field for CSV exports and the nested `author` object in code.
- **`pk` is a string.** Threads primary keys exceed 2^53, so pipelines that read them as JSON numbers silently corrupt the last digits and break de-duplication. This Actor never emits them as numbers.
- **Missing counters stay `null`, never `0`.** You can tell "no likes" apart from "not published by Threads".

***

### Tracking a keyword over time

Schedule the Actor with your keywords and append every run to the same dataset. Each record carries `searchQuery` and `scrapedAt`, so de-duplicating on `id` gives you a clean time series of when each post first appeared and how the conversation grew.

***

### Pricing

Pay per result. A run that returns nothing costs nothing. Search results are billed as posts scraped.

***

### Limits, stated plainly

**Depth.** Threads embeds roughly the first page of results in the page it serves, about 25 posts per keyword and ordering. Running both orderings roughly doubles coverage per term. If you need thousands of posts on one term, this version will not get you there; if you need broad coverage across many terms, it will.

**Proxies.** Threads rate-limits single IP addresses quickly. Residential proxies are strongly recommended beyond a handful of requests.

**Public data only.** Private accounts and anything behind a login are out of scope by design.

***

### Legal and compliance

This Actor reads only public Threads pages, the same ones any logged-out visitor can open. It does not log in and does not bypass an access control.

Threads posts contain personal data. If you are in the EU or UK, GDPR applies to what you collect and what you do with it, and having a lawful basis is your responsibility as the data controller.

***

### Related Actors

- **Threads Scraper** for posts, reply trees, profiles and search in one Actor
- **Threads Profile Scraper** for handles, follower counts and profile posts

***

### FAQ

**Do I need a Threads or Instagram account?**
No. Nothing here authenticates, and there is nowhere to enter credentials.

**How many keywords can I track in one run?**
As many as you like. Each is fetched separately, and concurrency is configurable.

**Why is `text` empty on a few records?**
Those are quote posts where the author added no words of their own. The post they quoted is in `quotedPost`, with its text and author.

**Does this work with threads.net links?**
Yes. Threads moved from threads.net to threads.com, and both forms are accepted, as are bare handles and bare post shortcodes. You do not need to rewrite old links.

**How is this different from the official Threads API?**
Meta's Threads API needs an app, an access token, and it caps keyword search at roughly 500 queries per rolling seven days. This Actor needs none of that and has no weekly quota, because it reads the same public pages a logged-out visitor sees.

**Can I use it from Make, Zapier, n8n or a script?**
Yes. Every Apify Actor is callable over the REST API and through Apify's integrations, and results come back as JSON, CSV or Excel.

**Does it break when Meta ships an update?**
Less often than most. Scrapers that call the internal GraphQL API depend on persisted query ids that Meta rotates without notice. This one reads what Threads server-renders into the page and recognises records by shape rather than by a fixed path.

# Actor input Schema

## `searchQueries` (type: `array`):

Keywords, phrases, brand names or hashtags to search on Threads. Each term is scraped separately and every result is tagged with the term that found it, so one run can track many terms at once.

## `searchSort` (type: `string`):

Threads ranks search results either by relevance or by recency. The two orderings return overlapping but different sets, so 'Both' gives the widest coverage at twice the request cost.

## `maxPostsPerQuery` (type: `integer`):

Upper bound on results returned for each keyword.

## `profiles` (type: `array`):

Threads handles or profile URLs to scrape in the same run, for example @zuck. Returns the profile plus its recent posts.

## `postUrls` (type: `array`):

Threads post URLs or shortcodes. Returns each post and its reply tree with conversation depth.

## `includeReplies` (type: `boolean`):

For each post URL, also scrape its replies with conversation depth. Turn off to fetch the posts alone.

## `includeProfilePosts` (type: `boolean`):

Return each account's recent posts alongside its profile record. Turn off for profile details only, which is faster and cheaper.

## `maxPostsPerProfile` (type: `integer`):

Upper bound on posts returned for each profile.

## `maxRepliesPerPost` (type: `integer`):

Upper bound on replies returned for each post. Set to 0 to skip replies entirely.

## `proxyConfiguration` (type: `object`):

Threads rate-limits single IP addresses quickly. Residential proxies are strongly recommended for anything beyond a handful of requests.

## `maxConcurrency` (type: `integer`):

Parallel requests. Raise for speed, lower if you see blocks.

## `maxRequestRetries` (type: `integer`):

How many times a failed request is retried, with a fresh proxy address each time, before it is given up on.

## Actor input object example

```json
{
  "searchQueries": [
    "openai"
  ],
  "searchSort": "default",
  "maxPostsPerQuery": 100,
  "includeReplies": true,
  "includeProfilePosts": true,
  "maxPostsPerProfile": 100,
  "maxRepliesPerPost": 200,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxConcurrency": 10,
  "maxRequestRetries": 4
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset containing all scraped data

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "openai"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("headply/threads-keyword-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["openai"] }

# Run the Actor and wait for it to finish
run = client.actor("headply/threads-keyword-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "openai"
  ]
}' |
apify call headply/threads-keyword-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,headply/threads-keyword-search-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/mrIakoaOLexeXQD4d/builds/t0E0MfD6Fifr3F4ca/openapi.json
