# Zhihu Scraper 2026 (`devcake/zhihu-scraper`) Actor

Collect Zhihu answers, articles, comments, trending questions, and creator profiles. Export to Excel or CSV for market research and brand monitoring.

- **URL**: https://apify.com/devcake/zhihu-scraper.md
- **Developed by:** [devcake](https://apify.com/devcake) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 standard results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Zhihu Scraper — Answers, Comments & Trends

Turn Zhihu conversations into structured data for **China market research, brand monitoring, content discovery, and creator research**. Zhihu Scraper collects answers, articles, comments, trending questions, videos, and public creator information, ready to explore or export in Excel, CSV, and JSON.

### 🔎 Discover what people are discussing on Zhihu

Search Zhihu by topic, brand, product, company, or question and bring the most useful public conversations together in one dataset. Instead of reviewing pages one by one, you can compare content, authors, engagement, and discussion themes at scale.

Zhihu Scraper can collect:

- **Search results** for keywords and phrases, with a choice of answers, articles, videos, or all available content
- **Question answers** with text, authors, publication details, and engagement
- **Articles and individual answers** from their Zhihu links
- **Comments and replies** from answers, articles, and videos
- **Zhihu Hot List** questions with their current rank and popularity labels
- **Creator profiles** with public biography, verification, follower, and publishing information when available
- **Video details** such as titles, creators, cover images, play counts, and duration when provided by Zhihu

### 🎯 Turn Chinese social media data into useful research

#### 🛍️ Market and consumer research

Study how people describe products, services, categories, and everyday problems in their own words. Zhihu answers and comment threads can reveal recurring questions, decision factors, objections, preferences, and unmet needs.

#### 🏷️ Brand monitoring

Collect public discussions around a company, product, campaign, or competitor. Compare the surrounding context, engagement, and authors behind brand mentions without reducing nuanced conversations to a single score.

#### 💡 Content research

Find questions people care about, explore the answers receiving attention, and identify themes worth explaining in articles, videos, newsletters, or social posts. The Zhihu Hot List adds a current view of topics attracting interest now.

#### 👤 Creator and expert discovery

Review public creator profiles alongside the answers and articles they appear in. Use follower counts, verification details, biographies, and engagement signals to build a more informed shortlist of relevant voices.

#### 📚 Research datasets

Build a focused Zhihu dataset around a subject, community, or discussion. Structured exports make it easier to organize qualitative research, compare sources, and prepare material for further analysis.

### 📊 Get the context behind every conversation

Zhihu data is most useful when the content and its context stay connected. Results can include:

- 📝 **Content:** titles, summaries, available answer or article text, source links, images, topics, and publication dates
- 💬 **Discussion:** comment text, replies, thread relationships, dates, likes, and reply counts
- 📈 **Engagement:** upvotes, comments, thanks, collections, views, answer counts, and other available interaction signals
- 👥 **Authors:** names, profile links, headlines, avatars, verification, biographies, follower counts, and publishing activity
- 🔥 **Trends:** Hot List rank, popularity label, question activity, topics, and collection time
- 🎬 **Videos:** titles, creator details, duration, play count, cover images, and source links

Not every Zhihu page exposes every detail. Missing values remain empty rather than being guessed, so your dataset stays faithful to the available source information.

### 🌏 Built for a clearer view of Zhihu

Zhihu combines long-form answers, professional perspectives, personal experience, and active discussion. That makes it valuable for research, but difficult to review consistently by hand. Zhihu Scraper brings these different content types into one organized collection while preserving source links for context.

The result is useful for researchers, marketers, strategists, journalists, analysts, agencies, and teams exploring Chinese social media data. No coding is required.

### ✅ What makes this Zhihu Scraper useful?

- **Broad Zhihu coverage:** search results, questions, answers, articles, comments, replies, videos, trends, and profiles
- **Conversation-level detail:** content, authors, engagement, and reply relationships stay connected
- **Research-ready exports:** review results in Apify or export them to Excel, CSV, and JSON
- **Flexible discovery:** research a topic broadly or focus on selected Zhihu links
- **Transparent results:** unavailable information stays empty, and shortened content is identified when Zhihu marks it as limited
- **Focused source collection:** the Actor gathers source material without inventing sentiment, translations, or conclusions

### ⚠️ Good to know

Zhihu controls what content is available, so some answers, articles, comments, or profile details may be missing or shortened. Search rankings, Hot List positions, and engagement counts can also change over time.

Video files and transcripts are not included. Zhihu Scraper does not translate Chinese text or assign sentiment scores. It collects the available source material so you can apply the analysis method that fits your research.

Use collected data responsibly and in line with applicable laws, platform rules, and privacy requirements.

### ❓ Frequently asked questions

#### What is Zhihu?

Zhihu is a Chinese question-and-answer and knowledge-sharing platform. People use it to publish detailed answers, articles, opinions, and videos across professional and consumer topics. Its mix of long-form content, creator profiles, comments, and engagement signals makes it a valuable source for understanding questions and discussions in the Chinese market.

#### How can social media be used for market research?

Public social media discussions can help researchers identify recurring questions, product expectations, objections, vocabulary, and emerging interests. Zhihu is especially useful when the reasoning behind an opinion matters: answers often provide more context than a short post, while comments show how other users react, agree, challenge, or add detail.

#### Can I create a Zhihu dataset for research?

Yes. You can build a Zhihu dataset from topic searches or selected public links and include available content, authors, comments, engagement, and trend information. Results can be exported to Excel, CSV, or JSON for qualitative review, categorization, comparison, or use with your preferred analysis tools.

#### Can Zhihu Scraper collect comments and replies?

Yes. It can collect comments from public Zhihu answers, articles, and videos, with the option to include replies. Each comment and reply is kept as its own record while preserving the relationship to the surrounding discussion, making it easier to follow conversations and compare audience reactions.

#### Does Zhihu Scraper analyze sentiment or translate Chinese text?

No. It returns the available Zhihu content and engagement information without assigning sentiment or translating the text. This avoids presenting automated interpretation as source data and lets you choose the language model, translation service, research framework, or review process that best fits your project.

# Actor input Schema

## `keywords` (type: `array`):

Add up to 20 topics, brands, products, or phrases. Queries run in order and share Max results; duplicate content is saved once. The unchanged example is automatically ignored when another section is configured.

## `content_type` (type: `string`):

Optionally limit search results to answers, articles, or videos.

## `max_items` (type: `integer`):

Set the total number of search results to save, from 30 to 1,000. Multiple queries run in order and share this limit.

## `urls` (type: `array`):

Add one or more question, answer, article, video, or creator-profile URLs. The Actor detects each URL type automatically and processes URLs in order. Leave empty when using another section.

## `max_content_results` (type: `integer`):

Set the total results to save across all URLs, from 30 to 1,000. Questions can return multiple answers; other URLs usually return one result.

## `hot_list` (type: `boolean`):

Turn on to collect Zhihu's current trending questions. Leave off when using another section.

## `comments_urls` (type: `array`):

Add one or more answer, article, or video URLs whose comments you want to collect. URLs are processed in order. Leave empty when using another section.

## `max_comments` (type: `integer`):

Set the total number of comment records to save across all URLs, from 15 to 1,000. You will receive up to this many records; Zhihu may return fewer when less content is available. Replies count toward this limit when Include replies is on.

## `comment_id` (type: `string`):

Legacy API field for collecting replies to one root comment.

## `include_replies` (type: `boolean`):

Also collect replies for every comment. Each comment and reply is saved as a separate result.

## `operation` (type: `string`):

Legacy API field. The form detects the operation automatically.

## `url` (type: `string`):

Legacy single-URL input. Use Zhihu URLs in the form.

## `comments_url` (type: `string`):

Legacy single comments URL. Use Content URLs in the Comments section.

## `cookie` (type: `string`):

Optional account access information for an existing setup.

## `keyword` (type: `string`):

Legacy single-query input. Use Search queries in the form.

## Actor input object example

```json
{
  "keywords": [
    "露营"
  ],
  "content_type": "all",
  "max_items": 30,
  "max_content_results": 30,
  "hot_list": false,
  "max_comments": 15,
  "include_replies": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `summary` (type: `string`):

Live collection stage, workflow, saved-result and input progress, continuation details, and safe error diagnostics.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "露营"
    ],
    "max_items": 30,
    "max_content_results": 30,
    "max_comments": 15
};

// Run the Actor and wait for it to finish
const run = await client.actor("devcake/zhihu-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["露营"],
    "max_items": 30,
    "max_content_results": 30,
    "max_comments": 15,
}

# Run the Actor and wait for it to finish
run = client.actor("devcake/zhihu-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "露营"
  ],
  "max_items": 30,
  "max_content_results": 30,
  "max_comments": 15
}' |
apify call devcake/zhihu-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,devcake/zhihu-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/uH4Si53lhxXOx4Dbt/builds/aBcFhIzbjvzJeylKb/openapi.json
