# Threads Scraper Pro (`celebrated-quadraphonic/threads-scraper-pro`) Actor

Threads Scraper Pro - 高级Threads数据抓取工具，支持监控、情感分析、高级搜索

- **URL**: https://apify.com/celebrated-quadraphonic/threads-scraper-pro.md
- **Developed by:** [XiaoZhi DataTools](https://apify.com/celebrated-quadraphonic) (community)
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Threads Scraper Pro

高级Threads数据抓取工具，支持监控、情感分析、高级搜索。

### 功能特性

#### 核心功能

- **帖子抓取**：抓取用户帖子，包含文本、互动数据、媒体URL
- **资料抓取**：获取用户资料，包含粉丝数、关注数、简介等
- **关键词搜索**：搜索包含特定关键词的帖子
- **标签搜索**：搜索特定标签下的帖子
- **监控功能**：实时监控用户/话题，支持通知（开发中）

#### 高级功能

- **时间过滤**：按时间范围过滤帖子
- **互动数据过滤**：按点赞、回复、转发数过滤
- **排序功能**：按时间、热度、互动数据排序
- **批量处理**：支持批量抓取多个用户/话题
- **增量更新**：仅抓取新内容，避免重复
- **情感分析**：帖子情绪分析（开发中）
- **AI集成**：支持MCP协议，可与AI代理集成

#### 技术特性

- **反爬策略**：浏览器自动化、代理轮换、请求频率控制
- **多格式导出**：JSON、CSV、Excel
- **错误重试**：自动重试失败请求
- **并发控制**：可配置并发数

### 使用方法

#### 输入参数

| 参数 | 类型 | 描述 | 默认值 |
|------|------|------|--------|
| `mode` | string | 抓取模式：posts/profile/search/hashtag/monitor | posts |
| `targets` | array | 目标列表（用户名、关键词、标签） | 必填 |
| `maxResults` | integer | 每个目标的最大结果数 | 100 |
| `timeRange` | object | 时间范围过滤 | null |
| `engagementFilter` | object | 互动数据过滤 | null |
| `sortBy` | string | 排序方式：recent/popular/engagement | recent |
| `includeMedia` | boolean | 是否包含媒体URL | true |
| `includeReplies` | boolean | 是否包含回复 | false |
| `concurrency` | integer | 并发数 | 3 |
| `proxyConfiguration` | object | 代理配置 | 使用Apify代理 |

#### 示例输入

##### 抓取用户帖子

```json
{
  "mode": "posts",
  "targets": ["username1", "username2"],
  "maxResults": 50,
  "engagementFilter": {
    "minLikes": 10,
    "minReplies": 5
  }
}
```

##### 搜索关键词

```json
{
  "mode": "search",
  "targets": ["AI", "machine learning"],
  "maxResults": 100,
  "timeRange": {
    "from": "2026-01-01T00:00:00Z",
    "to": "2026-12-31T23:59:59Z"
  },
  "sortBy": "popular"
}
```

##### 搜索标签

```json
{
  "mode": "hashtag",
  "targets": ["#AI", "#tech"],
  "maxResults": 200
}
```

### 输出格式

#### 帖子数据

```json
{
  "id": "1234567890",
  "text": "帖子内容",
  "timestamp": "2026-01-01T12:00:00Z",
  "likeCount": 100,
  "replyCount": 50,
  "repostCount": 25,
  "quoteCount": 10,
  "media": [
    {
      "type": "image",
      "url": "https://..."
    }
  ],
  "author": {
    "id": "user123",
    "username": "username",
    "fullName": "Full Name",
    "followersCount": 1000,
    "isVerified": false
  }
}
```

#### 用户资料数据

```json
{
  "id": "user123",
  "username": "username",
  "fullName": "Full Name",
  "biography": "用户简介",
  "followersCount": 1000,
  "followingCount": 500,
  "postsCount": 200,
  "isVerified": false,
  "profilePicUrl": "https://...",
  "externalUrl": "https://..."
}
```

### 定价

- **启动事件**：$0.02/次
- **数据项**：$0.0015/条
- **无月费**：按使用量计费

### 技术架构

#### 反爬策略

- 浏览器自动化（Playwright）
- 代理轮换（住宅代理）
- 请求频率控制
- User-Agent轮换
- Cookie管理

#### 数据存储

- 结构化JSON输出
- 支持多种导出格式
- 数据压缩与分页
- 增量更新支持

#### 性能优化

- 并发请求控制
- 内存优化
- 缓存机制
- 断点续传

### 开发计划

#### 阶段1：MVP（已完成）

- \[x] 基础帖子/资料抓取
- \[x] 关键词搜索
- \[x] 基本过滤器
- \[x] JSON/CSV导出

#### 阶段2：增强（进行中）

- \[ ] 高级搜索（时间、互动过滤）
- \[ ] 批量处理
- \[ ] 增量更新
- \[ ] 错误重试

#### 阶段3：智能（计划中）

- \[ ] 监控功能
- \[ ] AI情感分析
- \[ ] 实时通知
- \[ ] 关联分析

#### 阶段4：优化（计划中）

- \[ ] 性能优化
- \[ ] 文档完善
- \[ ] 用户反馈收集

### 许可证

MIT License

# Actor input Schema

## `mode` (type: `string`):

Choose what to scrape from Threads

## `targets` (type: `array`):

Usernames, keywords, or hashtags to scrape (one per line)

## `maxResults` (type: `integer`):

Maximum number of results to return per target

## `minLikes` (type: `integer`):

Only return posts with at least this many likes (0 = no filter)

## `minReplies` (type: `integer`):

Only return posts with at least this many replies (0 = no filter)

## `sortBy` (type: `string`):

How to sort the results

## `includeMedia` (type: `boolean`):

Extract image and video URLs from posts

## `includeReplies` (type: `boolean`):

Also scrape reply threads under each post

## Actor input object example

```json
{
  "mode": "posts",
  "maxResults": 100,
  "minLikes": 0,
  "minReplies": 0,
  "sortBy": "recent",
  "includeMedia": true,
  "includeReplies": false
}
```

# Actor output Schema

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("celebrated-quadraphonic/threads-scraper-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("celebrated-quadraphonic/threads-scraper-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call celebrated-quadraphonic/threads-scraper-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,celebrated-quadraphonic/threads-scraper-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dIbOkz0s7fisuhyNY/builds/uT7eE8goIs3lSrvW7/openapi.json
