# Threads Scraper - Posts, Profiles, Search & Monitor (No Login) (`valev-lab/threads-scraper`) Actor

Scrape public Meta Threads profiles, posts, and keyword search over HTTP. Monitor mode bills only for new posts. No login, cookies, or GraphQL doc\_id calls.

- **URL**: https://apify.com/valev-lab/threads-scraper.md
- **Developed by:** [Daniel Valev](https://apify.com/valev-lab) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.20 / 1,000 post extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Threads Scraper do?

**Threads Scraper** is a **Threads API alternative** for public [Meta Threads](https://www.threads.com) data. It scrapes **profiles**, **posts**, and **keyword search** results **without login**, cookies, or a browser — over plain HTTP with a Chrome TLS fingerprint.

Use it for **Threads brand monitoring**, competitor tracking, creator research, campaign listening, and AI/RAG pipelines. Input is **drop-in compatible** with [`automation-lab/threads-scraper`](https://apify.com/automation-lab/threads-scraper) for shared fields, plus **deep search**, **profile tabs**, and **monitor mode** that bills only for new posts.

#### Modes

- **profile** — username, display name, bio, followers, verification, profile picture, bio links
- **posts** — posts from selected profile tabs (`threads` / `replies` / `media` / `reposts`), merged and deduplicated
- **search** — keyword search; `searchDepth: deep` merges public surfaces
- **monitor** (`onlyNewPosts`) — remember seen `postId`s; emit and charge only new posts

### Why scrape Threads?

Threads is Meta’s growing text-first network. Teams use this Actor to:

- **Brand monitoring** — track mentions and competitor accounts on a schedule
- **Influencer / creator research** — engagement, posting cadence, media mix
- **Campaign & launch listening** — keyword search around product names and hashtags
- **Content research** — which formats and topics get likes, replies, and reposts
- **Data / AI pipelines** — structured JSON for warehouses, agents, and MCP workflows

Apify **Schedules**, **webhooks**, dataset exports (JSON/CSV/Excel), the REST API, and MCP make recurring runs straightforward.

### What data can Threads Scraper extract?

| Record | Key fields |
| --- | --- |
| **Profile** | `username`, `fullName`, `biography`, `followerCount`, `isVerified`, `isPrivate`, `profilePicUrl`, `bioLinks`, `userId`, `url` |
| **Post** | `postId`, `code`, `text`, `likeCount`, `replyCount`, `repostCount`, `quoteCount`, `shareCount`, `viewCount`, `mediaType`, `media[]`, `hashtags`, `mentions`, `urls`, `linkPreview`, reply/repost flags, `timestamp` / `date`, `url` |
| **Context** | `sourceTab` (posts mode), `searchQuery` / `searchSurface` (search mode), `isNew` (monitor mode) |

Missing values are always `null` (never omitted empty strings). **CDN media URLs expire** — download promptly if you need durable copies. Full field list: **Output** / **Dataset** tabs in Console.

### How to scrape Threads

1. Open **Threads Scraper** in [Apify Console](https://console.apify.com) (or the Store page after publish).
2. Choose **mode**: `profile`, `posts`, or `search`.
3. Add **usernames** (e.g. `zuck`, `@zuck`, or a Threads URL) and/or **searchQueries**.
4. Set **maxPosts** (start with `5`–`20`). Optionally set date bounds or profile tabs.
5. Keep **Apify Proxy → RESIDENTIAL** enabled (datacenter exits often return empty HTML shells).
6. Click **Start**. When the run finishes, open the **Dataset** and download JSON, CSV, or Excel.
7. For recurring brand listening, enable **monitor mode** (see examples below) and attach a **Schedule**.

No coding required for Console runs. Developers can use the **API** / **MCP** sections below.

### Compared to automation-lab/threads-scraper

| | This Actor (`valev-lab/threads-scraper`) | `automation-lab/threads-scraper` |
| --- | --- | --- |
| Transport | **HTTP-only** (no Playwright) | Browser for posts/search |
| Shared input | `mode`, `usernames`, `searchQueries`, `searchSort`, `maxPosts`, dates, `includeProfile` | Same core fields |
| Extras | `searchDepth`, `profileTabs`, **monitor** (`onlyNewPosts`) | — |
| Start fee (Free) | **$0.005** | **$0.02** |
| Per post / profile (Free) | **$0.0035** | **$0.005** |
| Best for | Cheaper scheduled monitors, HTTP footprint | Maximum browser-parity fields when SSR omits them |

Same sample input works on both for overlapping fields:

```json
{
  "mode": "posts",
  "usernames": ["zuck", "mosseri"],
  "maxPosts": 20,
  "includeProfile": true
}
```

### How much does it cost to scrape Threads?

Pay-per-event. Platform compute/proxy usage is included in this Actor’s Console PPE setup (unless you change that later).

| Event | Free | Bronze | Silver | Gold+ |
| --- | ---: | ---: | ---: | ---: |
| Actor start (`apify-actor-start`) | $0.005 | $0.005 | $0.005 | $0.005 |
| Post extracted (`post-extracted`) | $0.0035 | $0.0030 | $0.0026 | $0.0022 |
| Profile scraped (`profile-scraped`) | $0.0035 | $0.0030 | $0.0026 | $0.0022 |

≈ **$3.50 / 1,000 posts** on Free (down to **$2.20 / 1,000** on Gold+). Monitor skips and in-run duplicates are **not** charged.

| Scenario (Free) | ≈ event charge |
| --- | ---: |
| Health check (5 posts + 1 profile) | $0.026 |
| 20 posts + 1 profile | $0.0785 |
| Monitor run with 3 new posts | $0.0155 |

### Input

See the **Input** tab in Console for the full form. Summary:

| Field | Type | Default | Notes |
| --- | --- | --- | --- |
| `mode` | `profile` / `posts` / `search` | `posts` | Same as reference Actor |
| `usernames` | string\[] | `["zuck"]` | `zuck`, `@zuck`, or Threads URL |
| `searchQueries` | string\[] | `["artificial intelligence"]` | Search mode; `#tag` allowed |
| `searchSort` | `top` / `recent` | `top` | Preferred order (see limits) |
| `searchDepth` | `standard` / `deep` | `deep` | Deep merges public surfaces |
| `maxPosts` | 1–200 | `20` | Ceiling per username/query |
| `postedAfter` / `postedBefore` | ISO / RFC3339 | — | Inclusive / exclusive |
| `includeProfile` | boolean | `true` | Posts mode |
| `profileTabs` | tab\[] | all four | Posts mode |
| `onlyNewPosts` | boolean | `false` | Monitor |
| `monitorStateName` | string | `threads-monitor-state` | Named KV store |
| `monitorFirstRun` | `emit` / `baseline` | `emit` | First-run behavior |
| `proxyConfiguration` | proxy | Apify `RESIDENTIAL` | Prefer residential |

#### Posts example

```json
{
  "mode": "posts",
  "usernames": ["zuck"],
  "maxPosts": 5,
  "includeProfile": true,
  "profileTabs": ["threads", "replies"]
}
```

#### Deep search example

```json
{
  "mode": "search",
  "searchQueries": ["ai agents"],
  "searchSort": "top",
  "searchDepth": "deep",
  "maxPosts": 20
}
```

#### Date window example

```json
{
  "mode": "posts",
  "usernames": ["zuck"],
  "maxPosts": 100,
  "postedAfter": "2026-05-01T00:00:00Z",
  "postedBefore": "2026-06-01T00:00:00Z",
  "includeProfile": false
}
```

#### Monitor mode (important)

Monitor remembers `postId`s in a named key-value store (`monitorStateName`) and only emits/charges posts it has not seen. An **empty dataset after a successful run is often correct**, not a failure.

| `monitorFirstRun` | First run for that store key | Later runs |
| --- | --- | --- |
| `baseline` | Records existing IDs, **0 dataset items**, no post charges | Emits only new posts (`isNew: true`) |
| `emit` | Emits current posts as new and charges them | Emits only posts not seen yet |

**Baseline (first scheduled run)** — expect an empty dataset:

```json
{
  "mode": "posts",
  "usernames": ["zuck"],
  "maxPosts": 5,
  "includeProfile": false,
  "profileTabs": ["threads"],
  "onlyNewPosts": true,
  "monitorStateName": "acme-threads-monitor",
  "monitorFirstRun": "baseline"
}
```

Expected: log `Monitor baseline … charged nothing`, `postsSaved: 0`, `monitorSkipped` ≈ 5. Only the Actor start event is charged.

**Same store, later run (`emit`)** — empty dataset means nothing new:

```json
{
  "mode": "posts",
  "usernames": ["zuck"],
  "maxPosts": 5,
  "includeProfile": false,
  "profileTabs": ["threads"],
  "onlyNewPosts": true,
  "monitorStateName": "acme-threads-monitor",
  "monitorFirstRun": "emit"
}
```

**Demo emit on a fresh store** (items on the first run):

```json
{
  "mode": "posts",
  "usernames": ["zuck"],
  "maxPosts": 5,
  "includeProfile": false,
  "profileTabs": ["threads"],
  "onlyNewPosts": true,
  "monitorStateName": "acme-threads-monitor-demo",
  "monitorFirstRun": "emit"
}
```

**Search + monitor** (brand listening):

```json
{
  "mode": "search",
  "searchQueries": ["Acme launch"],
  "searchSort": "recent",
  "searchDepth": "standard",
  "maxPosts": 50,
  "onlyNewPosts": true,
  "monitorStateName": "acme-threads-monitor",
  "monitorFirstRun": "baseline"
}
```

Verify with key-value store record `RUN_SUMMARY` (`monitorSkipped`, `postsSaved`).

### Output

You can download the dataset as **JSON, CSV, Excel, or HTML**. Each item is either a profile or a post.

#### Profile example

```json
{
  "type": "profile",
  "username": "zuck",
  "fullName": "Mark Zuckerberg",
  "biography": "…",
  "followerCount": 5745053,
  "isVerified": true,
  "isPrivate": false,
  "profilePicUrl": "https://…",
  "bioLinks": [],
  "userId": "63055343223",
  "url": "https://www.threads.com/@zuck",
  "scrapedAt": "2026-09-28T18:00:00.000Z"
}
```

#### Post example

```json
{
  "type": "post",
  "postId": "3996155940894885511",
  "code": "Dd1MqfcG0aH",
  "username": "zuck",
  "text": "…",
  "likeCount": 1332,
  "replyCount": 1740,
  "repostCount": 93,
  "quoteCount": 100,
  "shareCount": 112,
  "viewCount": null,
  "mediaType": "text",
  "media": [],
  "hashtags": [],
  "mentions": [],
  "urls": [],
  "isNew": null,
  "timestamp": 1790598932,
  "date": "2026-09-28T…",
  "url": "https://www.threads.com/@zuck/post/Dd1MqfcG0aH",
  "scrapedAt": "2026-09-28T18:00:00.000Z"
}
```

Run summary: key-value store → `RUN_SUMMARY`.

### Honest limits (logged-out SSR)

- **No pagination** beyond the first HTML response Threads embeds
- Profile tabs typically **~4–10** posts each; merged tabs often **~14–25** unique
- Search surface often **~8–19** posts; deep merge can reach **~40–50** unique
- `searchSort` is **best-effort** — logged-out redirects often drop `filter=recent|top`
- No reply-thread expansion, followers list, or likers list
- `viewCount` is usually `null` in SSR
- Prefer **RESIDENTIAL** proxy — some exits return HTTP 200 empty shells

### Scheduling a monitor

1. Create an Actor Task with `onlyNewPosts: true` and `monitorFirstRun: "baseline"`.
2. Confirm the baseline run has **0 dataset items** and non-zero `monitorSkipped` in `RUN_SUMMARY`.
3. Switch the task to `monitorFirstRun: "emit"` (or use `emit` from the start if you want the first run charged).
4. Attach an Apify **Schedule**. Keep the same `monitorStateName`.
5. Optional: webhook → Slack / Sheets / your API when the dataset has items. Empty monitor runs (nothing new) are normal.

### Integrations, API, and MCP

- **Console** — run, schedule, and export datasets without code
- **Schedules** — hourly/daily brand or competitor monitors
- **Webhooks** — notify Slack, Make, Zapier, n8n, or your API when a run succeeds (and has items)
- **Integrations** — push datasets to Google Sheets, Amazon S3, and other Apify integrations
- **API / clients** — Node.js, Python, cURL (below)
- **MCP** — call the Actor from an Apify MCP–connected agent with the same JSON input

#### Node.js

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('valev-lab/threads-scraper').call({
  mode: 'posts',
  usernames: ['zuck'],
  maxPosts: 5,
  includeProfile: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

#### Python

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("valev-lab/threads-scraper").call(run_input={
    "mode": "posts",
    "usernames": ["zuck"],
    "maxPosts": 5,
    "includeProfile": True,
})
print(client.dataset(run["defaultDatasetId"]).list_items().items)
```

#### cURL

```bash
curl "https://api.apify.com/v2/acts/valev-lab~threads-scraper/runs?token=YOUR_API_TOKEN" \
  -X POST -H "Content-Type: application/json" \
  -d '{"mode":"posts","usernames":["zuck"],"maxPosts":5}'
```

### FAQ

#### Does Threads Scraper need a Threads login?

No. It only reads **public logged-out** pages. No cookies, session IDs, or CAPTCHA solving.

#### Why is my monitor dataset empty?

Usually that means the run **succeeded** and there was nothing new to bill. With `monitorFirstRun: "baseline"`, the first run always records IDs and writes **0** items. Later `emit` runs also write 0 items when every post was already seen — check `RUN_SUMMARY.monitorSkipped`.

#### Why did I get fewer posts than `maxPosts`?

Logged-out Threads does not expose reliable pagination. `maxPosts` is a **ceiling**. When the public HTML payload is exhausted, the run stops with `no_more_public_posts` (a warning, not a hard error).

#### Do I need residential proxy?

**Yes, prefer RESIDENTIAL.** Some datacenter or direct exits return HTTP 200 with an empty HTML shell and no post JSON. The Actor retries shells, but residential is the reliable default.

#### Why are media URLs broken later?

Threads/Instagram CDN links **expire**. Download or re-scrape when you need durable media.

#### Is view count always available?

Often **no** in logged-out SSR. `viewCount` is `null` when Threads does not embed it; likes/replies/reposts/quotes are usually present.

#### Is this the same as automation-lab/threads-scraper?

Input overlaps for drop-in tasks. This Actor is HTTP-only, typically cheaper on Free, and adds monitor / deep search / profile tabs. Field coverage can differ where browser-only metrics are missing from SSR. See the comparison table above.

#### Where do I report issues?

Use the Actor **Issues** tab on Apify, or contact the developer from the Store page.

### Legal

Public logged-out data only. You are responsible for complying with Meta’s terms, applicable law, and GDPR. This Actor does **not** extract emails or phone numbers from bios. Do not scrape personal data without a legitimate reason; if unsure, consult your lawyers.

### Changelog

#### v0.1

- Initial Actor: profile / posts / search modes
- Deep search, profile tabs, and monitor mode (baseline / emit)
- Store-oriented README: how-to, field table, comparison, integrations, FAQ
- Impit HTTP transport; PPE events `apify-actor-start`, `post-extracted`, `profile-scraped`

# Actor input Schema

## `mode` (type: `string`):

Choose what to scrape: user profiles only, user posts with engagement data, or search results by keyword.

## `usernames` (type: `array`):

Threads usernames for profile or posts mode. Accepts zuck, @zuck, or https://www.threads.com/@zuck (threads.net URLs too). Tracking params are stripped.

## `searchQueries` (type: `array`):

Keywords to search for on Threads (search mode only). #tag queries are allowed. Empty or >100-character entries are skipped.

## `searchSort` (type: `string`):

Preferred ordering: top (relevance) or recent (freshest first). Logged-out SSR often ignores filter=recent|top and returns the default ranking — the Actor still requests the matching surface and documents this limit.

## `searchDepth` (type: `string`):

standard = one public search surface matching searchSort. deep = merge default, recent/top attempt, users, and tags surfaces, then deduplicate.

## `maxPosts` (type: `integer`):

Ceiling on posts per username or search query (1–200). Logged-out pages usually expose fewer than this; the run stops with no\_more\_public\_posts when the public SSR payload is exhausted.

## `postedAfter` (type: `string`):

Optional ISO date (YYYY-MM-DD) or RFC3339 timestamp. Include posts at or after this time (inclusive).

## `postedBefore` (type: `string`):

Optional ISO date (YYYY-MM-DD) or RFC3339 timestamp. Include posts before this time (exclusive).

## `includeProfile` (type: `boolean`):

In posts mode, also emit a profile dataset item before posts.

## `profileTabs` (type: `array`):

Which profile tabs to fetch and merge in posts mode.

## `onlyNewPosts` (type: `boolean`):

When true, skip posts already seen in the named monitor key-value store. Skipped posts are not emitted and not charged.

## `monitorStateName` (type: `string`):

Named key-value store for monitor seen-IDs. Use different names to keep separate monitors.

## `monitorFirstRun` (type: `string`):

emit = charge and deliver new posts on the first run. baseline = record existing IDs only (emit and charge nothing) on the first run for each monitor key.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings. Prefer RESIDENTIAL: some datacenter / no-proxy paths return a valid HTTP 200 HTML shell (~255 KB) with no embedded post/user JSON. Residential with a browser-grade TLS fingerprint usually returns the full SSR payload.

## Actor input object example

```json
{
  "mode": "posts",
  "usernames": [
    "zuck"
  ],
  "searchQueries": [
    "artificial intelligence"
  ],
  "searchSort": "top",
  "searchDepth": "deep",
  "maxPosts": 5,
  "includeProfile": true,
  "profileTabs": [
    "threads",
    "replies",
    "media",
    "reposts"
  ],
  "onlyNewPosts": false,
  "monitorStateName": "threads-monitor-state",
  "monitorFirstRun": "emit",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "posts",
    "usernames": [
        "zuck"
    ],
    "searchQueries": [
        "artificial intelligence"
    ],
    "searchSort": "top",
    "searchDepth": "deep",
    "maxPosts": 5,
    "includeProfile": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("valev-lab/threads-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "posts",
    "usernames": ["zuck"],
    "searchQueries": ["artificial intelligence"],
    "searchSort": "top",
    "searchDepth": "deep",
    "maxPosts": 5,
    "includeProfile": True,
}

# Run the Actor and wait for it to finish
run = client.actor("valev-lab/threads-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "posts",
  "usernames": [
    "zuck"
  ],
  "searchQueries": [
    "artificial intelligence"
  ],
  "searchSort": "top",
  "searchDepth": "deep",
  "maxPosts": 5,
  "includeProfile": true
}' |
apify call valev-lab/threads-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,valev-lab/threads-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rlgTwHHHvxdDoVlHd/builds/2hUsHnMpLDdw1b19m/openapi.json
