# Facebook 公开小组帖子与完整照片集抓取 (`spbotdel/facebook-group-posts-cn`) Actor

抓取 Facebook 公开小组的最新与历史帖子，输出机器可读的 JSON 数据：正文、发布时间、作者、互动数据、稳定帖子链接，以及全部可获取的照片链接（含隐藏的 +N 照片集）。支持 Apify API、MCP、日常监控与历史回填。

- **URL**: https://apify.com/spbotdel/facebook-group-posts-cn.md
- **Developed by:** [Sergei Belostotskii](https://apify.com/spbotdel) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.49 / 1,000 facebook 小组帖子

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Facebook 公开小组帖子与完整照片集抓取

[![Run on Apify](https://img.shields.io/badge/Run%20on-Apify-2f7df6)](https://apify.com/spbotdel/facebook-group-posts-cn)
[![AI agents](https://img.shields.io/badge/AI%20agents-MCP%20ready-6f42c1)](https://docs.apify.com/integrations/mcp)
[![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)

抓取 **Facebook 公开小组的最新与历史帖子**，输出干净的机器可读 JSON：正文、Facebook 发布时间、作者、互动数据、稳定帖子链接，以及**全部可获取的照片链接**（包括藏在 `+N` 缩略图后面的照片）。

> \*\*一个计费结果 = 一个小组帖子。\*\*展开的照片链接包含在同一条数据、同一个按帖计费的价格里，不按照片收费。不需要你提供 Facebook 账号或 Cookie。

| 适合 | 不适合 |
| --- | --- |
| 公开小组、最新帖监控、历史回填、图片多的帖子、AI agent、MCP、API 与定时任务 | 私密/需登录的小组、公共主页或个人主页、Marketplace 搜索、评论展开、视频下载 |

### 为什么需要它：丢照片问题

Facebook 经常只显示几张预览图加一个 `+N` 角标，剩下的照片藏在单独的照片集里。只看预览的采集器会返回一条看起来正常的帖子，却悄悄丢掉大部分商品、房源、活动或证据照片。

本 Actor 把照片完整度当作帖子的一部分：

1. 采集动态，保留时间戳、作者与互动数据。
2. 识别照片集 token 与可疑的 `+N` 布局。
3. 展开可恢复的照片集（含备用链接）。
4. 输出照片质量字段，而不是对不完整的相册谎报成功。

### 你能拿到什么数据

每条结果对应一个帖子：`source_post_id`、`source_url`（稳定链接）、`created_at`（Facebook 发布时间）、`raw_text`（正文）、`author`（ID/姓名/主页链接）、`stats`（点赞/评论数）、`media`（全部照片对象：原图、缩略图、宽高、OCR 文字、来源）。

### 快速开始

1. 填入公开小组链接（支持数字 ID 与个性化链接），`maxPostsPerGroup` 填想要的条数。
2. 保持 **New posts first（最新优先）** 与 **Expand all photos（展开全部照片）** 开启。
3. 日常监控：传入 `knownPostIds`（已入库的帖子 ID）或 `sinceDate`，遇到已知帖子即停止，不重复花钱。
4. 历史回填：拿上一次 `SUMMARY.pointer.nextCursor`，用 `startCursor` 继续向更早翻页。

### 定价

**每 1000 个帖子 $2.99**（Free/Bronze 计划），另加每次运行 `$0.00005` 的启动事件。
付费 Apify 计划享受 Store 折扣：Silver `$2.69`，Gold/Platinum/Diamond `$2.49`/1000 帖。

| 数量 | 费用 |
| ---: | ---: |
| 20 帖 | `$0.0598` |
| 100 帖 | `$0.2990` |
| 1000 帖 | `$2.9900` |

- 一个 dataset 条目 = 一个公开小组帖子；
- 所有找回的照片链接都含在同一条结果里；
- 不按照片数量收费；
- 以 [Actor 页面](https://apify.com/spbotdel/facebook-group-posts-cn)的价格卡为准。

### 常见问题

#### 需要 Facebook 账号或 Cookie 吗？

不需要。只采集公开小组可见的内容。若 Facebook 临时弹出登录墙，Actor 会在 `SUMMARY` 中如实报告，而不是假装小组是空的。

#### 中文小组支持吗？

支持。正文、作者名中的中文（简体/繁体）会原样输出，照片集展开不受语言影响。

#### 能采评论或私密小组吗？

评论展开与私密小组不在本 Actor 范围内。只需要公开帖子与照片——选它。

### 限制

公开小组可见什么，就采什么：被删除、被设限、过期的媒体拿不到；视频仅返回可恢复的链接，不保证下载；Facebook 改版可能导致短暂波动，SUMMARY 会给出诊断字段。

### 反馈

英文原版：[facebook-group-posts-all-photos-scraper](https://apify.com/spbotdel/facebook-group-posts-all-photos-scraper)。遇到问题请附 run ID 与 `SUMMARY`，不要贴 Cookie 或 token。

# Changelog

This Actor's version history is a separate document: https://apify.com/spbotdel/facebook-group-posts-cn/changelog.md

# Actor input Schema

## `groupUrls` (type: `array`):

必填。Facebook 公开小组的链接、自定义链接或数字 ID。请只填写小组链接，不要填写帖子链接、Marketplace 链接或私密/需登录的小组。一次运行最多处理 maxGroupsPerRun 个小组。

## `maxGroupsPerRun` (type: `integer`):

一次运行中最多处理的小组数量。结果和计费按小组倍增。多余的小组链接将被跳过，并列在 SUMMARY.skippedGroupUrls 中。

## `maxPostsPerGroup` (type: `integer`):

每个小组返回的最大帖子结果数。每个帖子为一个计费结果：10 个小组 x 1,000 个帖子约为 10,000 个计费结果。深度历史回填一次最多取 1,000 条，保存 SUMMARY.pointer.nextCursor，再用 startCursor 继续。

## `sortMode` (type: `string`):

小组动态的排序方式。「最新优先」最适合采集最新帖子与日常监控。

## `expandAllPhotos` (type: `boolean`):

尽力还原藏在 Facebook +N 缩略图后面的照片链接。建议保持开启以获得完整数据；仅在快速预览时关闭。

## `includeRawPayload` (type: `boolean`):

在每条结果中附带解析后的 Facebook 原始数据。便于调试，但会让输出大很多。

## `onlyPostsNewerThan` (type: `string`):

规范日期下边界（含）的 \[onlyPostsNewerThan, onlyPostsOlderThan) 窗口。返回在此 Facebook 时间戳或之后发布的帖子。使用“最新优先”排序时，Actor 在遇到更旧帖子后停止。最适合增量监控；与 onlyPostsOlderThan 配合可精确收集时间段。

## `onlyPostsOlderThan` (type: `string`):

规范日期上边界（不含）的 \[onlyPostsNewerThan, onlyPostsOlderThan) 窗口。跳过在此 Facebook 时间戳或之后发布的帖子，并继续向更早历史推进。与 onlyPostsNewerThan 结合可收集特定时间段。

## `knownPostIds` (type: `array`):

你的系统里已存的帖子 ID。配合「最新优先」排序，遇到其中之一即停止，最适合每日监控。 状态路由：每日最新监控 -> knownPostIds；较旧帖子回填 -> 从 SUMMARY.pointer.nextCursor 获取 startCursor；多组回填 -> startCursorsByGroup；已存储的流水线状态 -> checkpoint。切勿将回填游标传入最新运行。

## `checkpoint` (type: `object`):

上一次 SUMMARY 中的 checkpoint 对象。在你的系统保存运行状态时，用于增量监控或按小组续采。 状态路由：每日最新监控 -> knownPostIds；较旧帖子回填 -> 从 SUMMARY.pointer.nextCursor 获取 startCursor；多组回填 -> startCursorsByGroup；已存储的流水线状态 -> checkpoint。切勿将回填游标传入最新运行。

## `startCursor` (type: `string`):

SUMMARY.pointer.nextCursor 中的游标。仅用于回填更早的历史帖子，不要拿昨天的游标来采新帖子。 状态路由：每日最新监控 -> knownPostIds；较旧帖子回填 -> 从 SUMMARY.pointer.nextCursor 获取 startCursor；多组回填 -> startCursorsByGroup；已存储的流水线状态 -> checkpoint。切勿将回填游标传入最新运行。

## `startCursorsByGroup` (type: `object`):

可选：小组链接到 GraphQL 游标的映射，用于多小组运行中回填更早页面。优先直接传上一次的 SUMMARY/checkpoint。 状态路由：每日最新监控 -> knownPostIds；较旧帖子回填 -> 从 SUMMARY.pointer.nextCursor 获取 startCursor；多组回填 -> startCursorsByGroup；已存储的流水线状态 -> checkpoint。切勿将回填游标传入最新运行。

## `paginationMode` (type: `string`):

Ranked 快照先收集候选帖子再按发布时间本地排序；游标翻页保持动态顺序并给出 next 指针，更适合监控与回填。

## `maxCandidates` (type: `integer`):

本地排序前收集的候选帖子。保留 0 以使用自动预算（约为 maxPostsPerGroup 的 3 倍，最低 +20）。仅在 SUMMARY 报告 candidate\_limit\_reached 且帖子缺失时提高。

## `maxPages` (type: `integer`):

公共小组初始化之后的最大 GraphQL 翻页请求数。保留 0 以使用自动预算（约为 maxPostsPerGroup 的 0.85 倍，最低 20）。仅在 SUMMARY 报告页面上限且帖子缺失时提高。

## `bootstrapRetries` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 若 Facebook 返回异常页面或登录墙，用全新代理/会话重新初始化小组主页的尝试次数。

## `graphqlPageRetries` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 遇到临时性空白或过小的 GraphQL 分页响应时的重试次数。

## `proxyCountry` (type: `string`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 Apify 代理的国家代码。

## `mediaSetRetries` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 每个 Facebook media/set 请求的重试次数。

## `mediaSetRetryDelayMs` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 media/set 重试的基础等待时间。

## `mediaExpansionConcurrency` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 同时展开 Facebook media/set 页面的帖子数。数值越大完整照片采集越快，但临时性失败也可能增多。

## `mediaExpansionAdaptive` (type: `boolean`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 开启后，若近期 media/set 展开质量下降，Actor 会自动把过高的展开并发降回稳定水平。

## `mediaDeferredRetry` (type: `boolean`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 主流程结束后、输出前，仅对可疑的照片集行再补采一次。小幅增加运行时间，换来更高的 +N 覆盖率。

## `mediaDeferredRetryConcurrency` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 延迟补采阶段同时重试的可疑帖子数。

## `mediaDeferredRetryExtraRetries` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 延迟补采阶段对可疑 media/set 行的追加重试次数。

## `mediaDeferredRetryDelayMs` (type: `integer`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 延迟补采的基础等待时间。

## `debug` (type: `boolean`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 在 SUMMARY 中附带响应样本与发现过程细节。日常生产运行请保持关闭。

## `discoverBundles` (type: `boolean`):

高级。保持默认值，除非 SUMMARY 诊断明确指出此开关。 下载 Facebook JS 包以发现最新的 GraphQL doc\_id，会多耗流量；备用 doc\_id 反正都会试一遍。

## Actor input object example

```json
{
  "groupUrls": [
    "https://www.facebook.com/groups/1564552680458634/"
  ],
  "maxGroupsPerRun": 10,
  "maxPostsPerGroup": 30,
  "sortMode": "CHRONOLOGICAL",
  "expandAllPhotos": true,
  "includeRawPayload": false,
  "paginationMode": "ranked_snapshot",
  "maxCandidates": 0,
  "maxPages": 0,
  "bootstrapRetries": 4,
  "graphqlPageRetries": 2,
  "proxyCountry": "US",
  "mediaSetRetries": 1,
  "mediaSetRetryDelayMs": 800,
  "mediaExpansionConcurrency": 3,
  "mediaExpansionAdaptive": true,
  "mediaDeferredRetry": true,
  "mediaDeferredRetryConcurrency": 2,
  "mediaDeferredRetryExtraRetries": 2,
  "mediaDeferredRetryDelayMs": 1500,
  "debug": false,
  "discoverBundles": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "groupUrls": [
        "https://www.facebook.com/groups/1564552680458634/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("spbotdel/facebook-group-posts-cn").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "groupUrls": ["https://www.facebook.com/groups/1564552680458634/"] }

# Run the Actor and wait for it to finish
run = client.actor("spbotdel/facebook-group-posts-cn").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "groupUrls": [
    "https://www.facebook.com/groups/1564552680458634/"
  ]
}' |
apify call spbotdel/facebook-group-posts-cn --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,spbotdel/facebook-group-posts-cn"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DsbNRPEVDgIrX9ayl/builds/fKKImxyq2MlW90aoN/openapi.json
