# Zhihu 知乎 Scraper - Questions, Answers, Comments & Users (`vulnv/zhihu-scraper`) Actor

Scrape Zhihu (知乎): keyword search for answers and articles, question details, a question's answers, answer comments, user profiles and a user's answers. Export upvotes, comments, favorites, full answer text and author data to JSON/CSV/Excel. No login, cookies or proxy needed.

- **URL**: https://apify.com/vulnv/zhihu-scraper.md
- **Developed by:** [VulnV](https://apify.com/vulnv) (community)
- **Categories:** Social media, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 search results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Zhihu Scraper - Search 知乎 Questions, Answers, Comments & Users

**Scrape Zhihu (知乎) and export answers, articles, questions, comments and user profiles to JSON, CSV or Excel.** Search answers and articles by keyword, pull a question's details, collect a question's answers, get an answer's comments, look up a user's profile, or list every answer a user has written - all from one Actor. Each result is one clean, flat row with upvotes, comments, favorites, the full plain-text body and author data.

No login, no cookies, no proxies and no China IP to configure - data is fetched through a fast, managed pipeline. Just pick an operation, add your input and press **Start**.

> **Unofficial notice:** This is an independent tool and is **not** affiliated with, endorsed by, or connected to Zhihu, 知乎 or Zhihu Inc. "Zhihu" and "知乎" are trademarks of their respective owners, used here only to describe what the Actor scrapes.

### What is Zhihu?

Zhihu (知乎) is China's largest question-and-answer community, often described as the Chinese Quora, with over 100 million monthly users. Its long-form answers and Zhuanlan (专栏) articles cover technology, science, careers, finance, products and everyday life, and are a key source of expert opinion and consumer sentiment in China. This scraper gives you programmatic access to Zhihu's search, questions, answers, comments and users so you can research topics, track opinion and find experts at scale without a Zhihu account.

### Operations

Pick one **Operation** and fill in the matching field:

| Operation | Input field | What you get |
|-----------|-------------|--------------|
| **Search answers & articles by keyword** | `keywords` | Answers and articles matching each keyword, with upvotes, comments, favorites and author data. |
| **Question details** | `questionUrls` | The full question: title, description, answer count, views, followers, comments. |
| **Question answers** | `questionUrls` | The answers to each question (up to the first 200), with full text and engagement. |
| **Answer comments** | `answerUrls` | Top-level comments on each answer, with likes, replies and IP region. |
| **User profile** | `userUrls` | A user's profile: followers, answer and article counts, headline, IP region. |
| **User's answers** | `userUrls` | Every answer a user has written, newest first, paginated. |

`questionUrls` accept `zhihu.com/question/...` URLs or bare numeric question IDs. `answerUrls` accept answer URLs (`zhihu.com/question/.../answer/...`) or bare numeric answer IDs. `userUrls` accept `zhihu.com/people/...` URLs or bare user tokens (the part after `/people/`) - one per line.

### How to use it (step by step)

1. Choose an **Operation**.
2. Fill in the matching input:
   - Search -> add one or more **Search keywords** (Chinese and English both work).
   - Question details / Question answers -> paste **Question URLs / IDs**.
   - Answer comments -> paste **Answer URLs / IDs**.
   - User profile / User's answers -> paste **User URLs / tokens**.
3. Set **Maximum results per input** (default 100, or 0 for all available) for the list operations.
4. For question answers, optionally choose **Default** ranking or **Recently updated first**. For comments, optionally choose **Top** or **Newest**.
5. Press **Start**. Export the dataset as JSON, CSV, Excel, XML or via the API.

### Input

| Field | Type | Description |
|-------|------|-------------|
| `operation` | string | **Required.** `search`, `question_detail`, `question_answers`, `answer_comments`, `user_profile`, or `user_answers`. |
| `keywords` | array | Search keywords (for `search`). |
| `questionUrls` | array | Question URLs or IDs (for `question_detail`, `question_answers`). |
| `answerUrls` | array | Answer URLs or IDs (for `answer_comments`). |
| `userUrls` | array | User URLs or tokens (for `user_profile`, `user_answers`). |
| `maxItems` | integer | Max results per keyword / question / answer / user. `0` = all available. Default `100`. |
| `answerSort` | string | `default` or `updated` (question answers only). |
| `commentSort` | string | `score` or `ts` (comments only). |

### Output

Every row carries a `record_type` of `answer`, `article`, `question`, `comment` or `user`. Common answer and article fields:

| Field | Description |
|-------|-------------|
| `content_id`, `url` | Answer or article id and canonical zhihu.com URL. |
| `title`, `excerpt`, `content` | Question or article title, a short excerpt and the full plain-text body. |
| `voteup_count`, `comment_count`, `favorite_count`, `thanks_count` | Engagement metrics. |
| `question_id`, `question_title`, `question_url` | The question an answer belongs to. |
| `author_id`, `author_url_token`, `author_name`, `author_headline`, `author_follower_count`, `author_url` | Author data. |
| `created_time`, `updated_time` | Created and last-edited time (ISO-8601 UTC). |

Question (`question`) rows add `detail`, `answer_count`, `visit_count`, `follower_count` and `comment_count`. Comment (`comment`) rows add `comment_id`, `answer_id`, `like_count`, `dislike_count`, `child_comment_count`, `is_hot` and `ip_location`. User (`user`) rows add `user_id`, `url_token`, `name`, `headline`, `gender`, `ip_location`, `follower_count`, `answer_count`, `articles_count` and `profile_url`.

Example answer row:

```json
{
  "record_type": "answer",
  "operation": "question_answers",
  "input": "https://www.zhihu.com/question/2065714833606104237",
  "content_id": "2073124822138341002",
  "url": "https://www.zhihu.com/question/2065714833606104237/answer/2073124822138341002",
  "title": "菲尔兹奖得主谈「人工智能可能会杀死数学」，您如何看待这个问题？",
  "voteup_count": 191,
  "comment_count": 50,
  "favorite_count": 109,
  "thanks_count": 5,
  "question_id": "2065714833606104237",
  "author_name": "Yuhang Liu",
  "author_url_token": "yuhang-liu-34",
  "author_follower_count": 386336,
  "created_time": "2026-08-18T11:11:14+00:00"
}
```

### Common use cases

- **Topic and opinion research** - see what Zhihu's experts and users say about any product, brand, technology or trend.
- **Expert (KOL) discovery** - find the top answerers for a niche, then pull their profile and full answer history.
- **Brand and competitor monitoring** - track questions about a brand, the answers they attract and the comment sentiment.
- **Market and AI research** - build datasets of long-form Chinese Q\&A content, comments and engagement.

### Notes on reliability

- Question answers return up to the first 200 answers of a question in Zhihu's default or recently-updated order.
- Search results are Zhihu's own ranking; consecutive result pages can overlap and duplicates are removed automatically, so a keyword may return slightly fewer results than requested.
- Avatar and image URLs are served by Zhihu's CDN and can change; fetch them promptly.
- Upvote, comment and follower counts reflect what Zhihu returns at scrape time.
- If a question, answer or user is deleted, private or restricted, that input is skipped and the run continues.

### FAQ

**Do I need a Zhihu account, cookies or a proxy?** No. Just add your input and press Start.

**Can I paste full URLs instead of IDs?** Yes - every input field accepts full zhihu.com URLs or bare IDs / tokens. An answer URL pasted into Question URLs uses its question.

**Can I run several keywords or URLs at once?** Yes - add multiple lines. Cross-input duplicate answers and articles are removed automatically.

**Can I try it for free?** Yes. Users on the free Apify plan can fetch up to 10 results in total to try the Actor. Upgrade to a paid Apify plan to run it without that limit.

**What export formats are supported?** JSON, CSV, Excel, XML and HTML, plus the Apify API and integrations such as Google Sheets, Zapier, Make and webhooks.

### Pricing

This Actor is **pay per result**: you are charged for each record it returns (search results, questions, answers, comments and user profiles), plus standard Apify platform usage. Different record types have different prices; see the **Pricing** tab for current rates. You are never charged for inputs that return nothing.

### Related scrapers

- **[Xiaohongshu Scraper](https://apify.com/vulnv/xiaohongshu-scraper)** - 小红书 / RedNote notes, users and comments.
- **[Bilibili Scraper](https://apify.com/vulnv/bilibili-scraper)** - 哔哩哔哩 videos, creators and comments.
- **[Reddit Posts Search Scraper](https://apify.com/vulnv/reddit-posts-search-scraper)** - Reddit posts by keyword.

# Actor input Schema

## `operation` (type: `string`):

What to scrape. Each operation uses a different input field below:

• Search → Keywords
• Question details → Question URLs / IDs
• Question answers → Question URLs / IDs
• Answer comments → Answer URLs / IDs
• User profile → User URLs / tokens
• User's answers → User URLs / tokens

## `keywords` (type: `array`):

For the "Search" operation. Add one or more keywords - Chinese and English both work. Each keyword is searched independently and cross-keyword duplicates are removed.

## `questionUrls` (type: `array`):

For the "Question details" and "Question answers" operations. Paste zhihu.com/question/... URLs or bare numeric question IDs (e.g. 46563853) - one per line.

## `answerUrls` (type: `array`):

For the "Answer comments" operation. Paste answer URLs (zhihu.com/question/.../answer/...) or bare numeric answer IDs - one per line.

## `userUrls` (type: `array`):

For the "User profile" and "User's answers" operations. Paste zhihu.com/people/... URLs or bare user tokens (the part after /people/) - one per line.

## `maxItems` (type: `integer`):

Upper bound on results per keyword / question / answer / user (for the list operations: search, question answers, answer comments, user's answers). Cost scales linearly. Set 0 to fetch all available. Question answers are limited to the first 200 answers per question.

## `answerSort` (type: `string`):

For "Question answers" only. Zhihu's default ranking, or most recently updated first.

## `commentSort` (type: `string`):

For "Answer comments" only. Order comments by score (top) or by time (newest).

## Actor input object example

```json
{
  "operation": "search",
  "keywords": [
    "人工智能"
  ],
  "maxItems": 100,
  "answerSort": "default",
  "commentSort": "score"
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped Zhihu results in the default dataset.

## `overview` (type: `string`):

Overview table of the scraped results.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "人工智能"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("vulnv/zhihu-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "keywords": ["人工智能"] }

# Run the Actor and wait for it to finish
run = client.actor("vulnv/zhihu-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "人工智能"
  ]
}' |
apify call vulnv/zhihu-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,vulnv/zhihu-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jHb8Wg6I6wrq0IHCb/builds/q9cTyz2zqcNjVktsY/openapi.json
