# Threads Scraper (`yasaslive/threads-scraper`) Actor

Scrape public Threads profiles, posts, replies and search results without login. Get text, likes, replies, reposts, media, hashtags and follower counts. Export to JSON, CSV, Excel or API.

- **URL**: https://apify.com/yasaslive/threads-scraper.md
- **Developed by:** [Eonix Pvt Ltd](https://apify.com/yasaslive) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.02 / actor start

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Threads Scraper — Posts, Profiles & Replies (Meta Threads)

Scrape public **Threads** (threads.net / threads.com) profiles, their latest posts, full reply threads and keyword search results. **No login, no cookies, no Threads account needed.** You get clean, ready-to-use data: post text, likes, replies, reposts, quotes, images, videos, hashtags, mentions, follower counts and more, exported as JSON, CSV, Excel or straight into your tools.

### What you can scrape

- 👤 **Profiles**: name, bio, follower count, verified badge, profile picture, bio links
- 📝 **A profile's posts**: text, publish date, likes, replies, reposts, quotes, images, videos, link previews, hashtags, mentions. Multi-part threads are split into individual posts and linked together.
- 💬 **Replies to any post**: the post itself plus its replies, including nested reply chains, each linked to its parent via `parentPostId`
- 🔎 **Keyword search**: the top public posts Threads shows for any search term

### Why scrape Threads instead of (or as well as) Instagram?

Every Threads account is an Instagram account: **same creators, same brands, same audience**. But while Instagram scrapers are everywhere, Threads has **roughly 10x less competition**, so the data is fresher, less picked over, and a real edge for anyone doing research or outreach. Threads is also *text-first*, so you get what people actually **say** about brands and topics, not just captions under photos.

### Use cases

- **Brand monitoring**: track what people say about your brand, products or competitors. Schedule the Actor with a search query and get new mentions every day.
- **Creator & influencer research**: compare follower counts, posting frequency and engagement (likes, replies, reposts) across creators before you reach out.
- **Lead generation**: find people and businesses talking about a problem you solve, then collect their profiles and bio links.
- **Market research & trend spotting**: see which topics, hashtags and formats get the most engagement in your niche.
- **AI agents & RAG**: feed clean, structured Threads conversations (posts + reply trees) into LLM pipelines, sentiment analysis or vector databases.
- **Community & crisis monitoring**: follow the full reply thread of a viral post and export every reply for analysis.

### Sample output

Real items from a test run with the default input (`profiles: ["zuck"]`), trimmed for readability (long image URLs shortened with `…`):

```json
[
  {
    "type": "profile",
    "userId": "63055343223",
    "username": "zuck",
    "fullName": "Mark Zuckerberg",
    "biography": "Mostly superintelligence and MMA takes",
    "profilePicUrl": "https://instagram.fcmb1-2.fna.fbcdn.net/v/t51.82787-19/825322135_17989325280103224_1252773933700107438_n.jpg?…",
    "followerCount": 5744496,
    "isVerified": true,
    "isPrivate": false,
    "externalUrl": null,
    "bioLinks": [],
    "threadsProfileUrl": "https://www.threads.com/@zuck",
    "sourceInput": "zuck",
    "scrapedVia": "http",
    "scrapedAt": "2026-09-26T19:18:05.647Z"
  },
  {
    "type": "post",
    "id": "3994109866205639942",
    "code": "Ddt7cL5EfUG",
    "url": "https://www.threads.com/@zuck/post/Ddt7cL5EfUG",
    "username": "zuck",
    "text": "Agrippa said it's time to get back to work 😎",
    "publishedAt": "2026-09-25T16:50:21.000Z",
    "likeCount": 4132,
    "replyCount": 368,
    "repostCount": 219,
    "quoteCount": 37,
    "images": ["https://instagram.fcmb1-2.fna.fbcdn.net/v/t51.82787-15/823965183_17989304049103224_2823186847710225415_n.webp?…"],
    "videos": [
      "https://instagram.fcmb1-2.fna.fbcdn.net/o1/v/t16/f2/m84/AQO9jbyD1anYYh9Ih-HQxh8Gsmaj7GE6TCSyT4NL7VnDyjGKtRTmiPMsJiPcbBQBtl4kjZmi0I46RbXZXy5pS30YZGNz4XY5audLQaE.mp4?…"
    ],
    "isReply": false,
    "parentPostId": null,
    "hashtags": [],
    "source": "profile"
  },
  {
    "type": "post",
    "id": "3992880350758034895",
    "code": "Ddpj4YZkd3P",
    "url": "https://www.threads.com/@zuck/post/Ddpj4YZkd3P",
    "username": "zuck",
    "text": "Here's everything I announced at Meta Connect today 👇",
    "publishedAt": "2026-09-24T00:07:31.000Z",
    "likeCount": 2656,
    "replyCount": 602,
    "repostCount": 161,
    "quoteCount": 34,
    "images": [],
    "videos": [],
    "isReply": false,
    "parentPostId": null,
    "hashtags": [],
    "source": "profile"
  }
]
```

#### All output fields

Every item has a `type` of either `"post"` or `"profile"`. Use the **Posts & replies** and **Profiles** tabs in the dataset view to see them separately. Missing values are always `null` (never an empty string), counts are numbers, and dates are ISO-8601 (UTC).

| Post field                                                             | Description                                                                                    |
| ---------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| `id`, `code`, `url`                                                    | Threads post ID, short code and link                                                           |
| `username`, `userId`, `userFullName`, `userIsVerified`                 | Author                                                                                         |
| `text`                                                                 | Full post text                                                                                 |
| `publishedAt`                                                          | When it was posted (ISO-8601)                                                                  |
| `likeCount`, `replyCount`, `repostCount`, `quoteCount`, `reshareCount` | Engagement                                                                                     |
| `images`, `videos`                                                     | Media URLs (highest resolution available; carousels included)                                  |
| `linkPreview`                                                          | `{ url, displayUrl, title, imageUrl }` for posts that share a link (tracking redirect removed) |
| `isReply`, `parentPostId`, `replyToUsername`                           | Reply structure: rebuild whole conversation trees with `parentPostId`                          |
| `quotedPostId`                                                         | ID of the post being quoted, if any                                                            |
| `topicTag`, `hashtags`, `mentions`                                     | Threads topic tag, `#hashtags` in the text (plus the topic tag), and `@mentions`               |
| `source`, `sourceInput`                                                | How it was found (`profile`, `post`, `reply`, `search`) and which of your inputs produced it   |
| `scrapedVia`, `scrapedAt`                                              | `http` or `browser`, and when it was scraped                                                   |

| Profile field                    | Description                       |
| -------------------------------- | --------------------------------- |
| `userId`, `username`, `fullName` | Identity                          |
| `biography`                      | Bio text                          |
| `followerCount`                  | Followers                         |
| `isVerified`, `isPrivate`        | Badges and privacy                |
| `profilePicUrl`                  | Largest available profile picture |
| `externalUrl`, `bioLinks`        | First bio link and all bio links  |
| `threadsProfileUrl`              | Link to the profile               |

### Pricing: pay only for results

This Actor uses simple **pay-per-event** pricing. You only pay for what you get:

| Event           | Price      | When                                   |
| --------------- | ---------- | -------------------------------------- |
| Actor start     | **$0.01**  | Once per run                           |
| Post scraped    | **$0.001** | Per post, reply or search result saved |
| Profile scraped | **$0.003** | Per profile saved                      |

That's **$1 per 1,000 posts**. Duplicates are never charged, and if a profile or post can't be scraped, you pay nothing for it.

**Worked example:** you want the latest 100 posts from 10 creators.
10 profiles × $0.003 + 1,000 posts × $0.001 + $0.01 start = **$1.04**.

**Stay in control of cost:** set **Max results** in the input, or set **Max total charge** in the run options. When the limit is reached, the Actor stops cleanly, keeps everything scraped so far, and records `budgetReached: true` in the run's `STATS` record.

### Input

The default input works out of the box: it scrapes the profile and latest posts of `@zuck`. Fill in **at least one** of Profiles, Post URLs, Search queries or Start URLs.

| Parameter               | Type         | Default           | What it does                                                                                                                                     |
| ----------------------- | ------------ | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| `profiles`              | list of text | `["zuck"]`        | Usernames (`zuck`, `@zuck`) or profile URLs. Returns the profile + latest posts.                                                                 |
| `postUrls`              | list of text | `[]`              | Post links, e.g. `https://www.threads.com/@zuck/post/Ddt7cL5EfUG`. Returns the post + replies. `threads.net` and short `/t/CODE` links work too. |
| `searchQueries`         | list of text | `[]`              | Keywords to search for.                                                                                                                          |
| `startUrls`             | list of URLs | `[]`              | Any mix of profile, post and search URLs (handy for bulk import from a spreadsheet).                                                             |
| `maxPostsPerProfile`    | number       | `50`              | Posts per profile. `0` = profile details only.                                                                                                   |
| `maxRepliesPerPost`     | number       | `100`             | Replies per post URL. `0` = the post only.                                                                                                       |
| `maxPostsPerSearch`     | number       | `50`              | Results per search query.                                                                                                                        |
| `maxResults`            | number       | `50`              | Cap on total items for the whole run. `0` = no cap.                                                                                              |
| `enableBrowserFallback` | yes/no       | `true`            | If Threads blocks plain requests 3 times in a row, retry with a real browser.                                                                    |
| `maxConcurrency`        | number       | `5`               | Pages fetched in parallel.                                                                                                                       |
| `proxyConfiguration`    | proxy        | Apify Residential | Threads blocks datacenter IPs; keep residential.                                                                                                 |

Example input:

```json
{
  "profiles": ["zuck", "mosseri"],
  "postUrls": ["https://www.threads.com/@zuck/post/Ddt7cL5EfUG"],
  "searchQueries": ["web scraping"],
  "maxPostsPerProfile": 100,
  "maxRepliesPerPost": 200,
  "maxResults": 1000
}
```

### How to use it from your code and tools

In the snippets below, replace `ACTOR_ID` with the ID shown at the top of this Actor's page (for example `username/threads-scraper`) and `YOUR_APIFY_TOKEN` with your token from **Apify Console → Settings → Integrations**.

#### REST API (any language)

Run the Actor and get the results in one call:

```bash
curl -X POST "https://api.apify.com/v2/acts/ACTOR_ID/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"profiles": ["zuck"], "maxPostsPerProfile": 20}'
```

(In the URL, write `ACTOR_ID` as `username~threads-scraper`, with a tilde instead of the slash.)

#### Python

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("ACTOR_ID").call(run_input={"profiles": ["zuck"], "maxPostsPerProfile": 20})

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    if item["type"] == "post":
        print(item["publishedAt"], item["likeCount"], item["text"])
```

#### Node.js

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('ACTOR_ID').call({ searchQueries: ['your brand'], maxPostsPerSearch: 50 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.filter((item) => item.type === 'post'));
```

#### Make (Integromat)

Add the **Apify → Run an Actor** module, pick this Actor, paste your input JSON, and enable **Run synchronously**. Then add **Apify → Get Dataset Items** to loop over the results into Google Sheets, Airtable, Slack, a CRM, or anywhere else.

#### n8n

Use the **Apify** node with the **Run Actor and get dataset** operation (or an HTTP Request node calling the REST endpoint above). Schedule it with a **Cron** trigger for daily brand monitoring.

#### AI agents (MCP)

Give Claude, ChatGPT, Cursor or any MCP-compatible agent direct access to Threads data through the Apify MCP server:

```
https://mcp.apify.com?actors=ACTOR_ID
```

Add it as a remote MCP server in your client, authenticate with your Apify account, and ask things like *"What are people on Threads saying about our product launch this week?"*

### FAQ

**Do I need a Threads or Instagram account?**
No. The Actor only reads public pages, exactly as a logged-out visitor sees them. Never give a scraper your account credentials.

**Why did one of my profiles fail?**
Open the run's **Storage → Key-value store → FAILED\_INPUTS** record. It lists every input that returned nothing and why. The most common reason is that the username doesn't exist or is private: Threads shows a login page for both, so the Actor can't tell them apart and reports "profile may not exist, may be private, or was blocked". You are not charged for failed inputs.

**How many search results can I get?**
Without logging in, Threads shows only its **top results** for each query, usually 10–25 posts. For deeper coverage, add several related queries, or scrape the profiles that show up in the results.

**Do image and video links expire?**
Yes. Threads media URLs are signed and expire after a few days. Download the files promptly if you need to keep them.

**What is `scrapedVia`?**
`http` means the fast, lightweight route was used; `browser` means Threads was blocking plain requests and the Actor switched to a real headless browser for those inputs. The data is identical either way.

**How do I rebuild a conversation tree?**
Every reply has `parentPostId`. Direct replies point to the post you scraped; nested replies point to the reply they answer.

**Can I schedule it?**
Yes. Use **Schedules** in Apify Console (e.g. every morning), and combine with Make, n8n or a webhook to push new posts wherever you need them.

### Limitations

- Only **public** data. Private profiles, follower/following lists, likes lists and anything else that needs a login are out of scope.
- Search returns Threads' top logged-out results only (see FAQ).
- A profile's feed lists its **own posts** (including multi-part threads), not replies it made on other people's posts.
- Threads changes its website often. If something breaks, we update the Actor, and you can report issues on the **Issues** tab.

### Legal & ethical use

This Actor collects **only publicly available data** that any logged-out visitor can see. It does not log in, bypass paywalls, or access private content. Scraping public data is generally lawful, but **you are responsible for how you use it**: comply with Threads'/Meta's Terms of Service, the GDPR, CCPA and other data-protection laws that apply to you, and don't use personal data for spam, harassment or any unlawful purpose. If you process personal data of EU residents, make sure you have a legal basis to do so. When in doubt, consult a lawyer. See also Apify's blog post [Is web scraping legal?](https://blog.apify.com/is-web-scraping-legal/)

# Changelog

This Actor's version history is a separate document: https://apify.com/yasaslive/threads-scraper/changelog.md

# Actor input Schema

## `profiles` (type: `array`):

Threads usernames to scrape, with or without the @ (e.g. "zuck" or "@zuck"). Profile URLs also work. Returns the profile plus up to "Max posts per profile" of their latest posts.

## `postUrls` (type: `array`):

Links to individual Threads posts, e.g. https://www.threads.com/@zuck/post/Ddt7cL5EfUG (threads.net links work too). Returns the post plus up to "Max replies per post" replies.

## `searchQueries` (type: `array`):

Keywords to search on Threads, e.g. "web scraping". Without login Threads shows only its top results for each query (usually 10–25 posts).

## `startUrls` (type: `array`):

Optional. Paste any mix of Threads profile, post or search URLs – they are routed automatically. Handy for bulk-importing a list from a spreadsheet.

## `maxPostsPerProfile` (type: `integer`):

Stop collecting a profile's posts after this many (a multi-part thread counts each part). Set to 0 to scrape only the profile details.

## `maxRepliesPerPost` (type: `integer`):

For each Post URL, stop after this many replies (nested replies included). Set to 0 to scrape only the post itself.

## `maxPostsPerSearch` (type: `integer`):

Stop after this many results for each search query.

## `maxResults` (type: `integer`):

Hard cap on the total number of items (profiles + posts) saved in this run – a simple way to control cost. Set to 0 for no cap (the run's max total charge still applies).

## `enableBrowserFallback` (type: `boolean`):

If Threads shows a login wall 3 times in a row, retry the blocked inputs with a headless browser. Slower and uses more compute, but rescues runs during heavy blocking. You are only charged for results either way.

## `maxConcurrency` (type: `integer`):

How many pages to fetch in parallel. Higher is faster but more likely to be rate-limited by Threads.

## `proxyConfiguration` (type: `object`):

Threads blocks datacenter IPs, so residential proxies are strongly recommended (the default).

## Actor input object example

```json
{
  "profiles": [
    "zuck"
  ],
  "postUrls": [],
  "searchQueries": [],
  "startUrls": [],
  "maxPostsPerProfile": 50,
  "maxRepliesPerPost": 100,
  "maxPostsPerSearch": 50,
  "maxResults": 50,
  "enableBrowserFallback": true,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `posts` (type: `string`):

All posts, replies and search results (type = "post").

## `profiles` (type: `string`):

Profile details (type = "profile").

## `stats` (type: `string`):

Error counters by category, route used, budgetReached flag.

## `failedInputs` (type: `string`):

Inputs that produced no data and why (e.g. profile not found, login wall).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "profiles": [
        "zuck"
    ],
    "postUrls": [],
    "searchQueries": [],
    "startUrls": [],
    "maxPostsPerProfile": 50,
    "maxRepliesPerPost": 100,
    "maxPostsPerSearch": 50,
    "maxResults": 50,
    "enableBrowserFallback": true,
    "maxConcurrency": 5,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("yasaslive/threads-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "profiles": ["zuck"],
    "postUrls": [],
    "searchQueries": [],
    "startUrls": [],
    "maxPostsPerProfile": 50,
    "maxRepliesPerPost": 100,
    "maxPostsPerSearch": 50,
    "maxResults": 50,
    "enableBrowserFallback": True,
    "maxConcurrency": 5,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("yasaslive/threads-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "profiles": [
    "zuck"
  ],
  "postUrls": [],
  "searchQueries": [],
  "startUrls": [],
  "maxPostsPerProfile": 50,
  "maxRepliesPerPost": 100,
  "maxPostsPerSearch": 50,
  "maxResults": 50,
  "enableBrowserFallback": true,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call yasaslive/threads-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,yasaslive/threads-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DSe7WopibkTWBDnTR/builds/Z07ue7DaGof1tICna/openapi.json
