# WeChat Official Account Article Scraper (`herus13/wechat-article-scraper`) Actor

Scrape WeChat Official Account (公众号) articles by keyword, article link or album — no login. Rows carry title, account\_name, published\_at, summary and an article link; full-article rows add content\_text, content\_markdown, images, author and originality. Export JSON, CSV or Excel.

- **URL**: https://apify.com/herus13/wechat-article-scraper.md
- **Developed by:** [herus13](https://apify.com/herus13) (community)
- **Categories:** Social media
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 search results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

The WeChat Official Account Article Scraper extracts public articles from WeChat Official Accounts (微信公众号). One actor, three modes: search articles by keyword, read specific article links in full, or list every article in an account's album — no login, no WeChat account.

### 中文说明

微信公众号文章抓取工具。三种模式：按关键词搜索公众号文章、按链接获取文章全文、获取公众号合集（专辑）中的全部文章。

搜索结果含标题 `title`、摘要 `summary`、公众号名称 `account_name`、发布时间 `published_at` 与永久链接；开启全文后每行还含正文 `content_text`、Markdown、图片、作者、原创标记与合集信息。

无需登录，也无需微信账号。运行自带住宅代理，结果支持导出 JSON、CSV 或 Excel。阅读数、点赞数与评论不对匿名访问开放，因此不提供。

### What you get

One dataset, one row per record, with `kind` saying what each row is. A search result row:

```json
{
  "kind": "search_result",
  "source_keyword": "咖啡",
  "page": 1,
  "position": 1,
  "total_results_estimate": 9455,
  "docid": "ab735a258a90e8e1-6bee54fcbd896b2a-38aaebf7e189d78e56f37bdd146f4cf5",
  "title": "咖啡地图 | “咖啡”遇上西雅图,好喝夜未眠",
  "summary": "不久前, 君去西雅图出差,参加咖啡展会.短短几天内,一口气逛了40多家咖啡馆……",
  "account_name": "企鹅吃喝指南",
  "account_username": "qiechihe",
  "account_verified": true,
  "published_at": "2017-06-11T11:52:21+08:00",
  "msgid": "2653009417",
  "idx": 1,
  "permanent_url": "https://mp.weixin.qq.com/s?__biz=MjM5Mzc5NTk1OQ==&mid=2653009417&idx=1&sn=147bc1880d7a81c3e16e91cb88f10377&chksm=bd448c11…",
  "cover_url": "http://mmbiz.qpic.cn/mmbiz_jpg/ibLgUdQXiaIVQkJlRvYZaRXj…/0?wx_fmt=jpeg",
  "scraped_at": "2026-09-27T17:11:39+00:00",
  "_meta": {"author": "https://apify.com/herus13", "status": "ok", "message": null}
}
```

`_meta` is run metadata, not scraped data. A target that fails appears as a row with only `_meta` (status `failed`, the reason in `message`); those rows are never charged.

Every mode writes into the same dataset, and the Output tab has one view per kind:

- `search_result` — one search hit: `title`, `summary`, `account_name`, `account_username` (the account's WeChat id), `account_verified`, `published_at`, `permanent_url`, `cover_url`, `has_video`, the hit's `page` and `position`, and `source_keyword`.
- `album_article` — one article of an album: `title`, `url` (its permanent link), `published_at`, `position` in the album, `cover_url`, plus the album's `album_title`, `album_article_count` and `album_read_count`.
- `article` — one full article: `title`, `author`, `digest`, `content_text`, `content_markdown`, `content_html`, `images` (with width and height), `video_ids`, `published_at`, the account (`account_name`, `account_username`, `account_alias`, `account_signature`), `is_original`, the region the account chose to show (`ip_region_province`), its album and `tags`, and `read_more_url` (阅读原文). A full article found by search carries the search hit in `search`; one found in an album carries the entry in `album_entry`.

Times are ISO 8601 in Beijing time. `permanent_url` is set whenever the source gives a permanent link. On search rows that is only some hits: Sogou gives most results as an encrypted `sogou_url` that expires within hours, with no permanent link behind it until the article is opened, so read those promptly or turn on `include_content`. A full article carries `biz`, `mid` and `idx` whenever the page or its link names them; `sn` and `permanent_url` need the article's permanent link, which a search hit without one does not reveal, so they stay empty there.

### What it costs

| Event | Price |
|---|---|
| Actor Start (`apify-actor-start`) | $0.00005 |
| search result (`search-result`) | $0.003 |
| album article (`album-article`) | $0.003 |
| article (`article`) | $0.015 |

### Input

| Field | Type | Required | Default | What it does |
|---|---|---|---|---|
| `mode` | string | no | `"search"` | What to collect: `search` finds Official Account articles by keyword; `article` reads specific article links in full; `album` lists every article in an account's album (合集). Each mode reads only its own fields — a field filled in for another mode is rejected, not ignored, so clear it when you switch. Default search. |
| `keywords` | array | no | — | Search terms, one per entry, for example `人工智能` or `咖啡`. Each term returns up to `max_items` articles; an article found by two terms is returned once. Required when `mode` is search; used ONLY by search — leave empty for every other mode (the form's example term must be cleared). |
| `article_urls` | array | no | — | mp.weixin.qq.com article links, one per entry: a short link (`https://mp.weixin.qq.com/s/` plus an id), a full link carrying \_\_biz, mid, idx, sn and chksm, or a search-result link (`/s?src=11&timestamp=…&signature=…`). A full link without chksm is refused, since WeChat answers it with a verification page instead of the article. Required when `mode` is article; used ONLY by article — leave empty for every other mode. |
| `album_urls` | array | no | — | WeChat album (合集) links, one per entry, such as `https://mp.weixin.qq.com/mp/appmsgalbum?__biz=…&album_id=…`. Each album returns up to `max_items` articles, newest first, with the permanent link of each. Required when `mode` is album; used ONLY by album — leave empty for every other mode. |
| `max_items` | integer | no | `20` | How many rows to return per keyword or per album. N returns at most N; 0 returns none; -1 means no limit. Search stops at WeChat search's own depth, about 94 articles per keyword, whatever the limit; an album is read up to 2,000 articles. Article mode must keep the default. Default 20. |
| `include_content` | boolean | no | `false` | Also open every search hit or album entry and return the full article instead of the summary row: body text, HTML and Markdown, images, author, account details, originality and tags. One more request per row, billed as an article. Default false. Used by search and album; article mode always returns full content and must leave it off. |
| `since` | string | no | — | Keep only articles published on or after this date, written as YYYY-MM-DD in Beijing time, for example 2026-09-01. Leave empty for no lower bound. WeChat search cannot filter by date and orders by relevance, so the filter runs on the rows fetched: a narrow window can return fewer rows than `max_items`. Used by search and album — leave empty for article mode. |
| `proxyUrls` | array | no | — | Leave empty and the run uses the residential proxy this actor ships with, included in the price of the run. To route the run through your own account instead, add one gateway URL per entry, for example http://user:pass@host:port — works with DataImpulse, Bright Data, Oxylabs, Smartproxy or any provider that issues URLs. When set, only these URLs are used. Use US exits. |

### How to run it

1. Open the actor and pick a **What to scrape** mode. Each mode reads only its own fields; clear the example keyword if you switch to article or album.
2. Fill that mode's target: keywords for search, article links for article, album links for album.
3. Turn on **Include full article content** if you want the body of every search hit or album entry, not just the summary row.
4. Run it and export the dataset as JSON, CSV or Excel.

Full articles for two keywords published since 1 September:

```json
{
  "mode": "search",
  "keywords": ["人工智能", "新能源汽车"],
  "max_items": 30,
  "include_content": true,
  "since": "2026-09-01"
}
```

Every article in an album:

```json
{
  "mode": "album",
  "album_urls": ["https://mp.weixin.qq.com/mp/appmsgalbum?__biz=MjM5MjA0MDk2MA==&action=getalbum&album_id=3144682336811220998"],
  "max_items": -1
}
```

Minimal input:

```json
{
  "mode": "search",
  "keywords": ["咖啡"],
  "max_items": 5
}
```

Available as the MCP tool `herus13--wechat-article-scraper` on mcp.apify.com and through the Apify API; send the same JSON.

### Use cases

- **Brand and product monitoring**: search a brand or product name every day and collect which Official Accounts wrote about it, with titles, summaries and publish times.
- **Content research and archiving**: keep the full text, Markdown and images of an account's articles by reading its albums, for search indexes, knowledge bases or LLM pipelines.
- **Media and KOL discovery**: find which accounts publish most on a topic and whether they are verified, then read their albums for their back catalogue.
- **Market and policy research on China**: gather what WeChat's publishers, from state media to industry newsletters, say about a sector, as plain text ready for translation or a classifier.
- **Link enrichment**: turn a list of mp.weixin.qq.com links shared in chats or spreadsheets into titled, dated, full-text rows.

### FAQ

**Do I need a WeChat account?** No. Every mode reads what an anonymous visitor sees, and the actor never logs in.

**How many results does a keyword return?** WeChat search shows an anonymous visitor ten pages per keyword, about 94 articles, and ranks them by relevance. Use several related keywords to go wider, and albums to go deeper into one account.

**Can I get read counts, likes or comments?** No. WeChat serves read, like and 在看 counts and comments only inside its own app to a signed-in user, so no anonymous scraper can return them honestly. They are left out rather than shipped as zeros.

**Can I get an account's full history?** Not directly: WeChat keeps an account's history page inside its app. Albums (合集) are public, so an account's albums list its articles back to the start, with permanent links.

**Why was my article link refused?** A full article link needs its `chksm` part; without it WeChat answers with a verification page instead of the article. Copy the whole link from the browser's address bar or use the short `mp.weixin.qq.com/s/…` link.

**Why did my run stop with "does not use"?** Each mode reads only its own fields, and a field filled in for another mode is refused rather than silently ignored. The form opens with an example keyword; clear it when you pick article or album.

**Why does `since` return fewer rows than `max_items`?** WeChat search cannot filter by date for an anonymous visitor, so the date filter runs on the results fetched. A narrow window keeps only the matching ones.

**Is a proxy included?** Yes, a residential proxy is included in the price of the run. To use your own, paste gateway URLs into **Your own proxy URLs**. Scraping at volume? Your own [DataImpulse](https://dataimpulse.com/?aff=404588\&utm_source=apify) account is cheaper per GB.

**Is scraping WeChat legal?** This actor reads only public articles, the ones any visitor can open in a browser, and never signs in. You are responsible for how you use the data and for respecting the authors' copyright; check WeChat's terms and your local rules.

### Related actors

Building a wider China social-listening pipeline? Pair this actor with:

- [Weibo Scraper](https://apify.com/herus13/weibo-scraper) — posts, profiles, comments and the realtime hot-search list from Weibo.
- [Bilibili Video Scraper](https://apify.com/herus13/bilibili-scraper) — videos, creators, comments and danmaku from China's largest video community.
- [1688 Scraper](https://apify.com/herus13/alibaba-1688-scraper) — products, suppliers and reviews from China's largest wholesale marketplace.

# Actor input Schema

## `mode` (type: `string`):

What to collect: `search` finds Official Account articles by keyword; `article` reads specific article links in full; `album` lists every article in an account's album (合集). Each mode reads only its own fields — a field filled in for another mode is rejected, not ignored, so clear it when you switch. Default search.

## `keywords` (type: `array`):

Search terms, one per entry, for example `人工智能` or `咖啡`. Each term returns up to `max_items` articles; an article found by two terms is returned once. Required when `mode` is search; used ONLY by search — leave empty for every other mode (the form's example term must be cleared).

## `article_urls` (type: `array`):

mp.weixin.qq.com article links, one per entry: a short link (`https://mp.weixin.qq.com/s/` plus an id), a full link carrying \_\_biz, mid, idx, sn and chksm, or a search-result link (`/s?src=11&timestamp=…&signature=…`). A full link without chksm is refused, since WeChat answers it with a verification page instead of the article. Required when `mode` is article; used ONLY by article — leave empty for every other mode.

## `album_urls` (type: `array`):

WeChat album (合集) links, one per entry, such as `https://mp.weixin.qq.com/mp/appmsgalbum?__biz=…&album_id=…`. Each album returns up to `max_items` articles, newest first, with the permanent link of each. Required when `mode` is album; used ONLY by album — leave empty for every other mode.

## `max_items` (type: `integer`):

How many rows to return per keyword or per album. N returns at most N; 0 returns none; -1 means no limit. Search stops at WeChat search's own depth, about 94 articles per keyword, whatever the limit; an album is read up to 2,000 articles. Article mode must keep the default. Default 20.

## `include_content` (type: `boolean`):

Also open every search hit or album entry and return the full article instead of the summary row: body text, HTML and Markdown, images, author, account details, originality and tags. One more request per row, billed as an article. Default false. Used by search and album; article mode always returns full content and must leave it off.

## `since` (type: `string`):

Keep only articles published on or after this date, written as YYYY-MM-DD in Beijing time, for example 2026-09-01. Leave empty for no lower bound. WeChat search cannot filter by date and orders by relevance, so the filter runs on the rows fetched: a narrow window can return fewer rows than `max_items`. Used by search and album — leave empty for article mode.

## `proxyUrls` (type: `array`):

Leave empty and the run uses the residential proxy this actor ships with, included in the price of the run. To route the run through your own account instead, add one gateway URL per entry, for example http://user:pass@host:port — works with DataImpulse, Bright Data, Oxylabs, Smartproxy or any provider that issues URLs. When set, only these URLs are used. Use US exits.

## Actor input object example

```json
{
  "mode": "search",
  "keywords": [
    "人工智能"
  ],
  "max_items": 20,
  "include_content": false
}
```

# Actor output Schema

## `results` (type: `string`):

One row per record. The kind field says what it is: search\_result, album\_article or article. Each row also carries `_meta` (author, status, message): run metadata, not scraped data.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "人工智能"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("herus13/wechat-article-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "keywords": ["人工智能"] }

# Run the Actor and wait for it to finish
run = client.actor("herus13/wechat-article-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "人工智能"
  ]
}' |
apify call herus13/wechat-article-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,herus13/wechat-article-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yG89ieenHi5se7ts3/builds/hqXflWW1JRl5M91wU/openapi.json
