Zhihu Scraper — Hot List, Q&A & Profiles avatar

Zhihu Scraper — Hot List, Q&A & Profiles

Pricing

from $3.00 / 1,000 results

Go to Apify Store
Zhihu Scraper — Hot List, Q&A & Profiles

Zhihu Scraper — Hot List, Q&A & Profiles

Scrape zhihu.com — trending hot-list questions (热榜), full Q&A answers with text and engagement counts, and author profiles as structured data. No login or API key required. Incremental mode flags new and changed records for monitoring and AI pipelines.

Pricing

from $3.00 / 1,000 results

Rating

0.0

(0)

Developer

Black Falcon Data

Black Falcon Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

What does Zhihu Scraper do?

Zhihu Scraper extracts structured Q&A data from zhihu.com — trending hot-list questions, full question and answer text, and author profiles. Each record carries engagement metrics (upvotes, comments, thanks, and followers), with no login required.

How to use this actor

  • 👉 Register for a free Apify account — no credit card required.
  • 🎉 Just click Sign up free on Apify → and complete a quick signup.
  • 💰 A free Apify account includes $5 in monthly credits — enough to test this actor.
  • ⏳ Scrape during the free trial, with no commitment or upfront payment required.

Key features

  • 🔔 Notifications — Telegram, Slack, Discord, WhatsApp Cloud API, and generic webhook out of the box. Pair with incremental for daily new-listing alerts without pipeline glue.
  • 🔗 Paste-mode — paste any zhihu URL straight from your browser — single-listing pages, search-results URLs, or category SEO URLs. Mix freely with keyword and IDs in the same run; results dedupe by ID.
  • 📦 Compact mode — AI-agent and MCP-friendly payloads with core fields only.
  • 📌 Change classification — each record carries a changeType of NEW / UPDATED / UNCHANGED / REAPPEARED / EXPIRED. Default emits NEW + UPDATED + REAPPEARED; opt into the others with emitUnchanged / emitExpired. Repost detection flags previously-expired listings that come back.
  • 🔌 MCP connectors — export your results into Notion via Apify's MCP connectors — a clean run-summary page, no glue code. Opt-in via the App connector field; deterministic field-mapping, no AI. Built on Apify's connector framework, so more destinations open up as their catalog grows.
  • ♻️ Incremental mode — recurring runs emit and charge only for listings that are new or whose tracked content changed. First run builds the baseline; subsequent runs emit only NEW / UPDATED / REAPPEARED records (UNCHANGED + EXPIRED opt-in). Saves 80–95% on daily monitoring.
  • 🧹 Empty-field stripping — drop null, empty-string, and empty-array fields from each record before push. Smaller payloads for AI agents and dashboards that already handle missing fields gracefully.
  • 📤 Export anywhere — Download the dataset as JSON, CSV, or Excel from the Apify Console, or stream live via the Apify API and integrations (Make, Zapier, Google Sheets, n8n, …).

What data can you extract from zhihu.com?

Each result includes Core listing fields (type, id, recordId, url, title, excerpt, content, and contentLength, and more). In standard mode, all fields are always present — unavailable data points are returned as null, never omitted. In compact mode, only core fields are returned.

Input

Configure the actor through the input schema in Apify Console.

Key parameters:

  • operation — What to scrape. Hot List = today's trending questions (热榜). Questions = full question + answers for the URLs you supply. Profiles = author profile stats for the URLs you supply. (default: "hotList")
  • limit — Number of trending questions to fetch (Hot List operation). 1–100. (default: 50)
  • startUrls — For Questions: question URLs (https://www.zhihu.com/question/12345) or bare ids. For Profiles: profile URLs (https://www.zhihu.com/people/token) or bare tokens.
  • includeAnswers — Fetch each question's page for the full question detail + first-page answers (with full text and engagement). Turn off for a fast trending snapshot (Hot List metadata only). (default: true)
  • maxAnswersPerQuestion — Cap answers emitted per question. 0 = all embedded first-page answers. (default: 0)
  • compact — Output only core fields (for AI-agent / MCP workflows). (default: false)
  • excludeEmptyFields — Drop null, empty-string, and empty-array fields from each record before push. (default: false)
  • incrementalMode — Compare against previous run state and tag each record NEW / CHANGED / UNCHANGED. stateKey is optional — defaults to a stable key derived from the operation and targets. (default: false)
  • stateKey — Optional. Stable identifier for the tracked set (e.g. "zhihu-hotlist"). Leave empty to auto-generate.
  • emitUnchanged — When incremental, also emit records that haven't changed. (default: false)
  • telegramToken — Telegram bot token (from @BotFather). Required for Telegram notifications.
  • telegramChatId — Telegram chat or channel ID (e.g. "-100123456789"). Required when telegramToken is set.
  • ...and 8 more parameters

Input examples

Trending hot list (fast snapshot) — undefined

→ undefined

{
"operation": "hotList",
"limit": 20,
"includeAnswers": false
}

Hot list with full answers — undefined

→ undefined

{
"operation": "hotList",
"limit": 10,
"includeAnswers": true
}

Specific questions by URL — undefined

→ undefined

{
"operation": "questions",
"startUrls": [
"https://www.zhihu.com/question/19550225"
]
}

Author profiles by URL — undefined

→ undefined

{
"operation": "profiles",
"startUrls": [
"https://www.zhihu.com/people/zhang-jia-wei"
]
}

Track the hot list for changes (incremental) — undefined

→ undefined

{
"operation": "hotList",
"limit": 50,
"includeAnswers": false,
"incrementalMode": true
}

Output

Each run produces a dataset of structured listing records. Results can be downloaded as JSON, CSV, or Excel from the Dataset tab in Apify Console.

Example listing record

{
"type": "question",
"id": "2061371759589323800",
"recordId": "zhihu.com:question:2061371759589323800",
"url": "https://www.zhihu.com/question/2061371759589323800",
"title": "经济学家任泽平 VIP 付费会员群「暴雷」,有人听信操作建议亏损 1000多万,你如何看这种现象?",
"excerpt": "科技股持续下跌,网红经济学家任泽平站上风口浪尖。此前,任泽平持续看好科技牛行情,指出AI科技牛是康波周期量级的机会,天花板远远没有看到,是这代人最重要的时代机遇。 此前科技股走出了波澜壮阔的行情,双创指数均创下历史新高,然而近期韩国“存储双雄”跳水,引发A股连锁反应,科技股遭遇踩踏行情。 据报道,近期群名为“泽平宏观VIP群30”的付费会员群出现投资者激烈控诉。该投资者称,自己听信任泽平相关 “科...",
"contentLength": 0,
"commentCount": 0,
"answerCount": 192,
"followerCount": 439,
"hotRank": 1,
"hotHeat": "442 万热度",
"authorName": "用户",
"questionId": "2061371759589323800",
"questionTitle": "经济学家任泽平 VIP 付费会员群「暴雷」,有人听信操作建议亏损 1000多万,你如何看这种现象?",
"topics": [
"922",
"6722",
"6723",
"19800",
"166889"
],
"createdTime": "2026-07-17T00:48:45.000Z",
"fetchedAt": "2026-07-18T00:00:00.000Z"
}

Incremental fields

When incremental mode is on, each record also carries:

  • changeType — one of NEW, UPDATED, UNCHANGED, REAPPEARED, EXPIRED. Default output covers NEW / UPDATED / REAPPEARED; set emitUnchanged: true to opt into the others.

How to scrape zhihu.com

  1. Go to Zhihu Scraper in Apify Console.
  2. Configure the input.
  3. Set maxResults to control how many results you need.
  4. Click Start and wait for the run to finish.
  5. Export the dataset as JSON, CSV, or Excel.

Use cases

  • Extract listing data from zhihu.com for market research and competitive analysis.
  • Monitor new and changed listings on scheduled runs without processing the full dataset every time.
  • Feed structured data into AI agents, MCP tools, and automated pipelines using compact mode.
  • Export clean, structured data to dashboards, spreadsheets, or data warehouses.

How much does it cost to scrape zhihu.com?

Zhihu Scraper uses pay-per-event pricing. You pay a small fee when the run starts and then for each result that is actually produced.

  • Run start: $0.01 per run
  • Per result: $0.003 per listing record

Example costs:

  • 10 results: $0.04
  • 25 results: $0.085
  • 100 results: $0.31
  • 200 results: $0.61
  • 500 results: $1.51

Example: recurring monitoring savings

These examples compare full re-scrapes with incremental runs at different churn rates. Churn is the share of listings that are new or whose tracked content changed since the previous run. Actual churn depends on your query breadth, source activity, and polling frequency — the scenarios below are examples, not predictions.

Example setup: 250 results per run, daily polling (30 runs/month). Event-pricing examples scale linearly with result count.

Churn rateFull re-scrape run costIncremental run costSavings vs full re-scrapeMonthly cost after baseline
5% — stable niche query$0.76$0.05$0.71 (94%)$1.43
15% — moderate broad query$0.76$0.12$0.64 (84%)$3.67
30% — high-volume aggregator$0.76$0.24$0.53 (69%)$7.05

Full re-scrape monthly cost at daily polling: $22.80. First month with incremental costs $2.14 / $4.31 / $7.58 for the 5% / 15% / 30% scenarios because the first run builds baseline state at full cost before incremental savings apply.

Platform usage (compute and proxies) is billed separately by Apify based on actual consumption. Incremental runs consume less on result processing, though fixed per-run overhead stays the same.

FAQ

How many results can I get from zhihu.com?

The number of results depends on the search query and available listings on zhihu.com. Use the maxResults parameter to control how many results are returned per run.

Does Zhihu Scraper support recurring monitoring?

Yes. Enable incremental mode to only receive new or changed listings on subsequent runs. This is ideal for scheduled monitoring where you want to track changes over time without re-processing the full dataset.

Can I integrate Zhihu Scraper with other apps?

Yes. Zhihu Scraper works with Apify's integrations to connect with tools like Zapier, Make, Google Sheets, Slack, and more. You can also use webhooks to trigger actions when a run completes.

Can I use Zhihu Scraper with the Apify API?

Yes. You can start runs, manage inputs, and retrieve results programmatically through the Apify API. Client libraries are available for JavaScript, Python, and other languages.

Can I use Zhihu Scraper through an MCP Server?

Yes. Apify provides an MCP Server that lets AI assistants and agents call this actor directly. Use compact mode and excludeEmptyFields to keep payloads manageable for LLM context windows.

This actor extracts publicly available data from zhihu.com. Web scraping of public information is generally considered legal, but you should always review the target site's terms of service and ensure your use case complies with applicable laws and regulations, including GDPR where relevant.

Your feedback

If you have questions, need a feature, or found a bug, please open an issue on the actor's page in Apify Console. Your feedback helps us improve.

You might also like

Getting started with Apify

New to Apify? Create a free account with $5 credit — no credit card required.

  1. Sign up — $5 platform credit included
  2. Open this actor and configure your input
  3. Click Start — export results as JSON, CSV, or Excel

Need more later? See Apify pricing.