# Changelog of Weibo Scraper — Hot Search, Posts, Comments & Creator Feeds (`zhorex/weibo-scraper`) Actor

- **URL**: https://apify.com/zhorex/weibo-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/zhorex/weibo-scraper.md

## Changelog

### 2026-10-02 — Run status fits Apify's 500 characters

No price change; rows, charges and requests are unchanged. Apify shows only the first 500 characters of a run's status message, so a long status lost its end (often the date-window note or a charge notice). A status that fits is exactly as before. A longer one now keeps, in this order: the result line, every charge-limit / not-charged / failure notice (in shorter words if needed, never dropped), the one tip, the review ask, then secondary details such as the date-window echo and purely informational notes (the key-value report pointer, options not applied, null detail fields). If the notices alone still pass 500 characters, a charge-limit notice moves right after the result line, whole, so it is never the part cut away. It is never cut mid-word, and the full status is written once to the run log.

### 2026-10-01 — Two opt-in add-ons: author profiles, follow-up comments

No price change, no new charge event: both bill the existing per-result event (`item-scraped`, $0.035) per row delivered, pushed before charged, inside the run's charge limit. Both are **off by default and not in the prefill**; with both off every mode is byte-identical to the previous commit (same requests, rows, charges, logs, status text and delta-state writes; checked offline mode by mode, `test_follow_up_and_author_profiles.py`).

- **`includeAuthorProfiles`** (`search`, `hot_timeline`, `user_posts`): one `user_profile`-shaped row per distinct post author (`mode: "author_profile"`), at most 50 authors per run, only as many as the charge limit can still bill, one author every 2 s (two requests each), at most ~3 minutes per run, stopping after 5 failures in a row. Fills a post's `authorFollowers` / `authorFollowing` where Weibo left them `null`; an author whose profile fails gets no row, no charge, and keeps `null`. In `deltaMode` only the authors of NEW posts, and profile rows never enter the delta memory.
- **`followUpComments`** (`search` + `deltaMode` only): the posts a delta run delivers wait in the `deltaStateKey` record (`followUp`: post ID and posting time, at most 10,000); a later delta run re-reads each thread once, 24 h to 7 days after the post, newest first, and delivers only comment IDs the delta memory has never seen, as comment rows marked `isFollowUp: true` (once that memory is full and forgets its oldest ids, only IDs above the highest one delivered for the post, kept in its queue entry). A due post the run's own search returns again is skipped only when it reports no comments. At most 300 threads, 400 requests (never passed inside a thread), 6 pages per thread and ~4 minutes per run, 2 s apart, 5 minutes' reserve before the timeout, 5 failures in a row stop it. Posts that leave the queue unread (7 days old, or past its 10,000 cap) are counted in the status message. A post leaves the queue only after its rows are pushed (one whose rows the charge limit cut stays). Not with `adFilter: "promo_only"`.
- `dataset_schema.json` gains one field, `isFollowUp` (boolean, nullable). Both options are named as not applied when set in a mode they do not cover.

### 2026-09-30 — Review ask, profile fields

No price change. Defaults and prefill are unchanged.

- **Fixed (`user_posts`):** an account whose profile detail came back with `sunshine_credit: null` lost its profile row (the read raised and the row was dropped with "Profile fetch failed"). The row is now delivered, and charged, like any other profile row.
- **Review ask:** briefly changed to a neutral ask on every one-off run, then reverted to the owner's rule: the log asks only when a one-off run returned rows, and the status only on one-off runs with 10+ rows that were not trimmed or held back ("Useful? A 30-sec review..."). Never on a delta run (`post_details` also skips scheduled runs). The log's issue line no longer promises a fix within 48h.

### 2026-09-26 — deltaMode: a row is no longer delivered and charged twice

No price change. Defaults, prefill and every run without `deltaMode` are unchanged (checked offline against the previous commit: same requests, rows, charges and status in every mode).

- **Fixed: past 20,000 remembered IDs, delta runs forgot random ones.** A `deltaStateKey`'s memory was saved from an unordered set and cut to 20,000 IDs, so once a stream passed 20,000 every run forgot IDs at random — including posts Weibo was still returning — and the next run delivered and charged those again. Offline, a `search` monitor already at 20,000 IDs (1,000 posts per run, 100 of them new) got 133 repeated rows in 30 runs, 4% of what it was charged for. The memory now keeps its order, a row Weibo returns again counts as recent, and only the IDs Weibo has gone longest without returning are forgotten.
- **Fixed (`complaints`):** a company list's memory was cut in delivery order, so complaints that stayed at the top of the list (recent replies) were the first forgotten past 20,000 and were delivered and charged again. They now count as recent for as long as the list shows them.
- **Fixed: the same post twice in one delta run** (an account listed twice in `userIds`) was delivered and charged twice. It is now delivered once; the status counts the second copy as a duplicate within the run, not as already seen.
- The memory holds up to 200,000 IDs per `deltaStateKey` (about 4.8 MB, half of what Apify accepts in one request). Memories saved by earlier builds load as they are; nothing is reset.
- A delta run that finds nothing new now saves which already-delivered rows Weibo still returned (it used to write nothing).

### 2026-09-25 — Post details: current text & engagement for a list of posts

No price change, no new charge event. **Nothing changes for any existing mode** — same requests, same rows, same charges, same status text (checked offline, mode by mode including `user_profile` and `complaints`, against the previous commit). No existing default or prefill changed; the input only gains the `post_details` mode value, and the `postIds` / `maxResults` descriptions mention it.

#### Added

- **New mode `post_details`: one row per post, no login, no comments.** Paste post IDs or URLs in `postIds`: the numeric ID, the short code from post links (`RjB5HwkDj`), `https://weibo.com/<uid>/<id or code>`, `https://m.weibo.cn/detail/<id>` or `https://m.weibo.cn/status/<id or code>`. Each row is the post as it stands at run time — `text`, `repostsCount`, `commentsCount`, `attitudesCount`, author fields, `createdAt` / `createdAtIso`, `source`, `images`, `videoUrl`, `isRepost` / `repostOf`, `postUrl` — with the same keys as the post row of `post_comments` (`mode: "post_detail"`), so the two join. On a Schedule, successive runs give each post's engagement over time.
  - Charged as one ordinary result (`item-scraped`, $0.035) per post **delivered**, pushed then charged in batches of 25, with the charge limit and its free sample of 10 applied once per run. `maxResults` caps the posts per run, and the status says how many it left out.
  - The short code is the numeric ID in base 62 (checked on 10 of 10 live pairs), so every entry is resolved to the numeric ID before any request: one post pasted as a code and as a number is fetched, and charged, once.
  - A post Weibo does not show (deleted, hidden, never existed — Weibo answers `{"ok": 0, "message": "该微博不存在", "error_code": 20101}`, measured today) gives no row and no charge; the status names up to 10 per reason, and every post not delivered is in the key-value store record `POST_DETAILS_REPORT`, with `retryIds`. A clean run writes no record.
  - A value Weibo does not send is `null`, never `0`, `""` or `false`: Weibo's post endpoint carries no author follower count, bio or verification reason, which the `post_comments` post row fills with `""`.
  - Long posts get their full text (`fetchFullText`, on by default) one request at a time; a post whose full text did not load keeps `isTruncated: true` and is named in the status.
  - At least 2 s between requests; after 5 failures or refusals in a row the run renews its anonymous session once, then stops and lists the rest as not fetched. New posts stop starting 5 minutes before the run's timeout.
  - `sentimentAnalysis` applies. `includeComments`, `geoRollup`, `adFilter`, `deltaMode` and the search date window do not; the status names them. `deltaMode` is ignored (one log line, no delta state read or written). With `postIds` empty, the run takes 3 of Weibo's current hot posts to show the output, as `post_comments` does.
  - In this mode the review ask (run log and status) is skipped on runs started by an Apify Schedule.
  - Entry handling: several posts pasted into one entry are one entry each; a bare code that reads as a word (`Marketing`, a header cell) or a code of the wrong length (`weibo.com/<uid>/profile`) is named as skipped instead of being decoded into a post ID nobody typed; the status names each post left out as it was pasted (URL, code or ID), and says to split the list when `maxResults` is already at its 5,000 maximum. The 3 hot posts taken when `postIds` is empty are said to be charged.
  - Posts mixing images and a video carry their images in `mix_media_info` (measured today: `pic_num` 2, `pic_infos` null); `images` is read from there instead of coming back `null`.

### 2026-09-24 — search: say when Weibo stopped serving pages

No price change. When Weibo stops answering a search before `maxResults` (server errors after 3 tries, or a refusal) and the run has no date window, the status now says which keyword stopped on which page and how many of the requested posts were delivered (the rest were not charged). Measured today: `比亚迪` answers page 1 and then HTTP 500 on page 2, while 耐克 / 泡泡玛特 / 小米汽车 page on normally; such runs used to end with only "Done! Scraped 45 items".

### 2026-09-24 — deltaMode pitch removed from run logs and status

Text only: after a one-off `search` / `user_posts` run, the run log no longer suggests turning on `deltaMode` with a daily Schedule, and the status no longer ends with "Want a recurring feed? Turn on deltaMode…". Nothing else changed — same requests, rows, charges, defaults and prefill, and `deltaMode` works exactly as before (still documented in the input and README). Checked offline against the previous commits: every mode's output is identical apart from the removed text.

The README and input-form wording changed the same way: the top banner now points at the date window (`searchDateFrom` / `searchDateTo`), the deeper search past Weibo's first result window, `user_profile` and `complaints` instead of a daily delta feed; the "pay only for new / a quiet run costs $0" lines were removed from the `deltaMode`, `includeComments` and hot\_timeline text, which now describe what delta does without selling it; the scheduling section's first example is a plain daily search, with the `deltaMode` version kept as an optional variant; and the false claim that scheduled runs append to one dataset now says each run delivers its own dataset. An unused log function that mentioned deltaMode was deleted (it was never called).

### 2026-09-24 — user\_posts: point a walled run at user\_profile

No price change, no new request, no change to rows or charges. When Weibo serves no timeline to an anonymous session (the `user_posts` login wall, already named in the status and charged $0), the status now adds one sentence: the same `userIds` give each account's profile — followers, following, post count, verification, bio, location — without a cookie in `user_profile` mode. It appears only on those walled runs; every other status is unchanged (checked offline against the previous commits).

### 2026-09-24 — Black Cat complaints

No price change. **Nothing changes for any existing mode** — same requests, same rows, same charges, same status text (checked offline, mode by mode including `user_profile`, against the previous commit). No existing default or prefill changed.

#### Added

- **New mode `complaints`: consumer complaints from Black Cat (黑猫投诉, tousu.sina.com.cn), Sina's complaint platform — no login.** New optional input `complaintCompanyIds`: company page URLs (`https://tousu.sina.com.cn/company/view/?couid=<id>`) or the bare `couid` numbers. Each company's complaint list is read in the order given, `maxResults` shared across them. Black Cat ranks a list by recent activity, not filing date (seen live), so no page of complaints already delivered ends a list; only a second ID for a list already read in the run (a page of only its complaints, and the same complaint count) is skipped, and named. Leave it empty for Black Cat's public latest-complaints feed; the status says which source the run used.
  - One row per complaint: `complaintId`, `title`, `summary` (as the list serves it; `isTruncated` when it is a preview ending in `...`), `appeal` (投诉要求), `issue` (投诉问题), `amountInvolvedCny` (涉诉金额 — public feed only, `null` on company lists), `companyName`, `companyId` (as served: on a company list Black Cat echoes the requested ID, a sub-brand's complaint in its parent's list included), `status` / `statusEn` / `statusCode` (Black Cat's own labels; an unlabelled code keeps `null` labels), `createdAt` (Beijing time) and `createdAtTimestamp` (raw), `upvoteCount`, `commentsCount`, `shareCount`, `url` (the complaint page exactly as served — it needs its `sld` token), `source` (`company` or `feed`). **No complainant name, avatar or ID** in any row.
  - Charged as one ordinary result (`item-scraped`, $0.035) per complaint **delivered**, pushed then charged in batches of 25, with the charge limit and its free sample of 10 applied once per run. A company ID Black Cat rejects (10002 参数错误) or answers with no complaints gives no row and no charge; the status names it. A complaint that turns up twice in a run (a sub-brand's complaint in its parent's list too, two IDs for one company, or a list that moved while it was read) is delivered and charged once.
  - Black Cat shows anonymous visitors the **first 500 complaints of a list** — the most recently active (page 51 answers 登录查看更多内容, measured on a company list and the feed); the run stops there and says so, and also names a list that came back empty before its own count. No keyword search: that endpoint now answers with a web page, not data.
  - The run status names up to 10 companies per reason; every list not delivered in full is in the key-value store record `COMPLAINTS_REPORT`, with `retryCompanyIds`. A clean run writes no record.
  - At least 0.5 s between requests; the run stops after 5 requests in a row fail or are refused.
  - `deltaMode` works here by complaint number, per company list (records `DELTA_COMPLAINTS_<deltaStateKey>_<couid>` in the usual delta store); each list is read to its end or the 500 depth, complaints already delivered skipped and not charged. Company IDs typed in full-width digits (２０９２…) are read as ASCII. `sentimentAnalysis` scores title + summary. `includeComments`, `geoRollup`, `adFilter` and `cookieString` do not apply, and the status says so.

### 2026-09-24

No price change. **Nothing changes for any existing mode** — same requests, same rows, same charges, same status text (checked offline, mode by mode, against the previous build).

#### Added

- **New mode `user_profile`: one row per Weibo account, no login.** Paste numeric user IDs or profile URLs (`weibo.com/u/<id>`, `weibo.com/<id>`) in `userIds`. Each row carries `screenName`, `followersCount`, `friendsCount`, `statusesCount`, `verified` / `verifiedType` / `verifiedReason`, `description`, `gender`, `location`, `avatarHd`, `coverImage`, `domain`, `profileUrl`, and from Weibo's profile detail `createdAt` (account creation), `birthday`, `company` and `sunshineCredit`.
  - Charged as one ordinary result (`item-scraped`, $0.035) per profile **delivered**. An ID Weibo does not resolve (deleted, banned, not a user), or that comes back without a screen name, gives no row and no charge; neither does an ID whose request failed. Duplicates are fetched and charged once, and a zero-padded ID (`02803301701`) is the same account as `2803301701`. `maxResults` does not apply (one row per ID).
  - The run status names up to 10 IDs per reason. The full list of IDs not delivered, and why, is saved in the run's default key-value store as the record `PROFILE_REPORT`, with `retryIds` = the IDs worth re-running. A run that delivered every ID writes no record.
  - Rows are delivered, then charged, in batches of 25 as the run goes, so a run cut off by its timeout keeps what it already delivered; the run also stops starting new IDs 5 minutes before its timeout.
  - If Weibo refuses or fails 5 IDs in a row, the run renews its anonymous session once and, if the next ID fails too, stops instead of hitting Weibo with the rest of the list. Those IDs are reported as not fetched (not charged).
  - A field Weibo does not send is `null` — not `""`, `false` or `-1`. New field `location` is the place the account shows on its profile (self-declared), not an IP location.
  - Respects the run's charge limit: IDs the limit cannot pay for are not fetched (with the usual free sample of 10), and the status says how many were left out.
  - Follower and following **lists** are not offered: Weibo serves `/ajax/friendships/friends` only to logged-in sessions (an anonymous request gets `{}`).
  - Options that work on posts (`includeComments`, `geoRollup`, `adFilter`, `sentimentAnalysis`, `deltaMode`) do nothing in this mode, and the status says so rather than applying them.

### 2026-09-23

No price change. **Nothing changes for a run that does not set the two new fields** — same requests, same rows, same charges, same status text.

#### Added

- **Optional date window in `search` mode: `searchDateFrom` / `searchDateTo`.** Both optional, ISO (`2026-09-01` or `2026-09-01T08:00:00`). A plain `searchDateFrom` is 00:00:00 that day and a plain `searchDateTo` is 23:59:59, so a date range covers whole days.
  - **Dates are Beijing time (UTC+08:00) unless you write an offset**, because that is the zone Weibo stamps every post with and the zone every delivered row's `createdAtIso` already carries. A UTC window would return rows whose own date field shows the day *after* the end date and would miss the first 8 hours of the start date. An offset you write (`Z`, `+00:00`, `+05:30`) is always honoured exactly — including on a bare date, which CPython's own `fromisoformat` mangles, so the Actor parses that shape itself. The status message and the run log print the window in both zones.
  - Applies to **`search` mode only**. Set in any other mode, the two fields filter nothing — and now the run says so in its status message and log instead of delivering and charging every row in silence.
  - Weibo's search API has no lower-bound parameter (measured 2026-09-23: `starttime` and `timescope` are ignored, `endtime` is honoured), so the Actor starts at `searchDateTo` and pages **backwards** until the posts are older than `searchDateFrom`.
  - Posts outside the window — and the rare row with an unreadable date — are dropped **before delivery**, so they are never pushed and never charged.
  - `maxResults`, the run's charge limit, the run timeout and Weibo's own page limits still cap the walk. When one of them stops the run before it reaches `searchDateFrom`, the status message says so and names the reason. It always reports the oldest and newest post actually delivered, and never claims complete coverage of a period.
  - In `deltaMode` the dates filter the rows but the deeper history walk stays off (as it always is in delta mode); the status message says the run is not full coverage of the period.
  - An unparseable date, or `searchDateFrom` at or after `searchDateTo` (an empty window), **fails the run** with an explanatory message and charges nothing, rather than quietly scraping a different period. A `searchDateTo` in the future starts from now.
  - The "we reached your start date" claim is made on how deep the walk actually got — the newest post of the last page Weibo served — and never on a single row. One out-of-order old post used to end the walk on page 1 and then have the status claim the whole period had been covered; a term Weibo never answered for used to flip a covered run to "NOT the full period". Neither can happen now, and a term that returned nothing is reported as exactly that.

### 2026-09-21

No price change.

#### Changed (check these if a pipeline depends on them)

- **`hot_search_delta` compares the whole board, whatever `maxResults` is.** With a `maxResults` below the board size (e.g. 25), a topic that slipped just below the cut used to come back as `dropped` (an extra billed row) and then as `new` with its `firstSeenAt` and peaks reset. Now `maxResults` only sets how many topics get a row, and `dropped` means the topic left the board entirely. The first run after this update compares against the old, cut snapshot, so a topic tagged `new` in that one run may have been just below your previous cut; the status says so when it happens.
- **`isTruncated` is now false on posts Weibo confirms are complete.** When the full-text lookup answers but has no longer text, the flag is cleared instead of staying true. It stays true only when the lookup failed.
- **`fetchFullText` now also applies to the post row of `post_comments`** (comment rows are unchanged).
- The run status now says why when `includeComments` added no comments or `geoRollup` added no rollup rows. The search-mode tip no longer suggests `includeComments` when the results have no replies.

### 2026-09-16

No price change. Nothing is removed from the output; the changes below are new fields and corrected values.

#### Changed (check these if a pipeline depends on them)

- **`hot_search` `rank` now follows Weibo's own board position (`realpos`).** The board carries one paid slot; it used to ship as a normal trend with a rank, pushing every topic below it down by one. That row is still delivered, now with `isPromoted: true` and `rank: null`, and `adFilter: "organic_only"` removes it (unbilled).
- **`hot_search_delta`:** the first run after this update reports `rankDelta: null` / `status: "steady"` for topics still on the board, because the saved snapshot used the old rank numbering. Runs after that diff normally.
- **`isHot`** is now true for the 热, 沸 and 爆 badges (it was 爆 only, so it was false on almost every row). Filters on `isHot` will return more topics.
- **`autoLocalize` search** now alternates pages between the Latin and Chinese brand name and dedupes as it goes, so both names contribute even when `maxResults` is small. Searches with a single term are unchanged.
- **Post `text`** (and `repostOf.text`) no longer ends in invisible zero-width spaces.
- **`geoRollup`** rows can be countries as well as provinces: known countries get an English `provinceEn` and `isOverseas: true`; 中国香港 / 中国澳门 / 中国台湾 map to Hong Kong / Macau / Taiwan.

#### Added

- `isTruncated` on posts and `repostOf`: Weibo's list endpoints send only the start of long posts; this flags those rows. The full text is not fetched.
- `createdAtIso` on posts, `repostOf` and comments (ISO 8601). `createdAt` keeps Weibo's raw format.
- Comment rows: `authorFollowers`, `authorVerified`, `replyCount`, `authorRegion`, `authorRegionEn`.
- `videoUrl` is now filled for video cards Weibo sends as `object_type: "video"`. It is a signed link that expires about 1 hour after the scrape.
- Status message says when `hot_timeline` stopped on its 240 s time budget before reaching `maxResults`, and when a `user_posts` timeline came back empty without a cookie.
