Zhihu Email Scraper
Pricing
from $2.49 / 1,000 results
Zhihu Email Scraper
Zhihu Email Scraper SD - Zhihu Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Zhihu results by keyword, location and email domain - Zhihu email extractor.
Pricing
from $2.49 / 1,000 results
Rating
0.0
(0)
Developer
Leads Scraper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
14 days ago
Last modified
Categories
Share
Zhihu Email Scraper
Zhihu Email Scraper for Chinese Expert and Creator Outreach
The Zhihu Email Scraper is an Apify Actor that collects publicly indexed contact emails from Zhihu (知乎) and delivers them as a structured dataset. Zhihu is China's long-form Q&A and knowledge platform, and it is unusually rich in professionals who publish a working email address.
The Zhihu Email Scraper does not log into the platform. It builds Google queries with the site:zhihu.com operator, parses the result blocks and extracts the addresses Google has already made public.
That fits Zhihu's shape well. Consultants, founders, engineers, doctors and 专栏 (Zhuanlan) columnists routinely put a contact or 商务合作 mailbox at the bottom of an answer or in a profile bio.
Give the Zhihu Email Scraper keywords, an optional location and the email domains you want. It returns account names, profile URLs, bio snippets and normalised, deduplicated emails.
Who this Zhihu email extractor is for
B2B marketers selling into China. Recruiters sourcing Chinese technical and professional talent. Content teams looking for subject-matter experts. Agencies building KOC and expert media lists for the Chinese market.
Key Features of the Zhihu Email Scraper
| Feature | What it means in practice |
|---|---|
| Google-based contact discovery | The Zhihu Email Scraper reads publicly indexed zhihu.com results — no login, no cookies, no browser |
| Query expansion | Each keyword runs as a base query, a quoted query, an intitle: query and one variant per modifier |
| Domain-filtered extraction | Only emails on your customDomains list are kept, so you can target @qq.com or @163.com |
| Global deduplication | Every address is deduplicated across all queries and pages in the run |
| Email normalisation | Handles name [at] domain [dot] com, name (at) domain, name @ domain.com, zero-width characters and the full-width @ |
| Junk filter | Rejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals |
| Boundary-correct matching | @gmail.com will not match inside @gmail.company or @gmail.com.br |
| Soft-wrap repair | Drops a hit that is only the tail of another address in the same result block |
| Concurrency control | An asyncio worker pool runs several queries in parallel with a shared stop signal |
| Retry logic | Up to 3 attempts per page with exponential backoff and a fresh proxy session per request |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried, not treated as empty |
| Resumable state | Progress is checkpointed in the key-value store and survives PERSIST_STATE, MIGRATING and ABORTING events |
| Fallback parser | If Google's markup changes, the crawler degrades to "emails without account details" rather than "no results" |
| Structured output | Fourteen fields per row, exportable to CSV, JSON or Excel |
Zhihu answers are long. That works in your favour: a longer indexed page gives Google more snippet material, and the Zhihu Email Scraper has more text to search for a contact line.
How the Zhihu Email Scraper Works
- Read input. Keywords, location, email domains and limits are validated.
- Build queries. The Zhihu Email Scraper composes searches such as
site:zhihu.com 咨询顾问 contact "@163.com". - Fetch search results. Pages are requested asynchronously through the Apify
GOOGLE_SERPproxy usingaiohttp. - Parse structurally. The parser locates each
<h3>title, then the smallest surrounding block — it does not depend on Google's CSS class names. - Extract emails. A domain-filtered regex pulls addresses from the block text, including obfuscated and full-width forms.
- Deduplicate and store. Each unique lead is pushed to the Apify dataset immediately.
Base queries always run first, so the strongest matches land in your dataset early. Blocked or failed queries are re-queued once at the end of the run.
The Zhihu Email Scraper never opens zhihu.com. There is no browser, no JavaScript rendering, no authentication and no platform API key.
What Data Does It Extract?
Each row pairs one email with the account metadata Google printed alongside it — usually a Zhihu display name, sometimes a column or organisation title.
username and profileUrl are populated when Google exposes a zhihu.com/people/... handle. Where Google prints only a Chinese display name, you still receive accountName, fullName, the snippet and the email.
Descriptions are stripped of navigation labels and vote or comment counters before they reach the dataset, so the snippet you read is the actual bio or answer text.
Input Fields and Configuration
| Field | Type | Default | Meaning |
|---|---|---|---|
keywords | array (required) | ["founder", "consultant"] | Search terms — niche, job title or industry |
location | string | "" | Optional location phrase added to every query |
customDomains | array | ["@gmail.com","@yahoo.com"] | Only emails on these domains are kept; the @ is optional |
maxEmails | integer 1–10000 | 20 | Stop after this many unique emails |
countryCode | string | "" | Two-letter country for the search proxy (US, GB, DE...) |
expandQueries | boolean | true | Search each keyword × domain pair in several phrasings |
queryModifiers | array | ["email","contact","business","cooperation"] | Extra words combined with each keyword when expansion is on |
maxPagesPerQuery | integer 1–50 | 30 | Pagination cap per query |
maxConcurrency | integer 1–20 | 5 | Parallel queries |
Using countryCode with the Zhihu Email Scraper
countryCode is a real lever on a regional platform like Zhihu. Google returns different result sets depending on the country the search proxy appears to originate from.
HK, TW and SG frequently surface Chinese-language pages that a default US-region search buries. US is useful for finding bilingual accounts and overseas Chinese professionals.
The practical technique is to run the Zhihu Email Scraper several times with different country codes and merge the exports. Deduplication is per run, so removing overlap afterwards is trivial.
Choosing the right email domains
@gmail.com and @yahoo.com are the defaults, but Zhihu users mostly publish @qq.com, @163.com, @126.com, @foxmail.com and @sina.com addresses.
Zhihu also skews professional, so university and corporate domains are worth trying — @edu.cn, or a specific company domain when you are prospecting one organisation. The Zhihu Email Scraper accepts any domain you list, with or without the leading @.
Choosing keywords
Keywords are inserted into the Google query verbatim, so Chinese input works as well as English. 产品经理, 律师, 留学中介, 创业者, 投资人 and 数据分析 are all valid.
You can also swap the default modifiers for Chinese equivalents such as 邮箱, 联系, 合作 and 咨询 to run a fully Chinese-language expansion set.
Output Schema and Dataset Fields
| Field | Meaning |
|---|---|
network | Platform name |
keyword | The keyword that produced the lead |
query | The exact Google query used |
title | Raw result title |
accountName | Account label Google prints (handle or display name) |
fullName | Display name parsed from a profile-style title; empty for post captions |
username | URL-safe handle when Zhihu exposes one; otherwise null |
profileUrl | Canonical account URL when a handle is known; otherwise empty |
url | Direct platform link when exposed, else the profile URL |
description | Bio or answer snippet, cleaned of labels and counters |
email | Lower-cased email address |
emailDomain | The matched domain, e.g. @163.com |
possiblyTruncated | true when Google's snippet ellipsis touched the email — verify before sending |
foundAt | ISO 8601 UTC timestamp |
query and keyword are stored on every row, which makes the output auditable. You can see which phrasing found which expert and refine the next Zhihu Email Scraper run from that evidence.
How to Use the Zhihu Email Scraper
- Open the Zhihu Email Scraper on Apify and click Try for free.
- Enter two or three specific keywords — a narrow professional niche outperforms a broad category.
- Replace
customDomainswith the Chinese mail providers listed above. - Set
countryCodetoHK,TW,SGor leave it empty and compare the results. - Set
maxEmailsto the list size you actually need. - Run it, then export the dataset as CSV, JSON or Excel.
Keep maxPagesPerQuery and maxConcurrency at their defaults for a first pass. Raise concurrency only when your Apify plan has memory headroom.
The Zhihu Email Scraper is also callable from the Apify API and SDK clients, so recurring expert-discovery runs are easy to schedule.
Use Cases for the Zhihu Email Scraper
| Use case | How the Zhihu Email Scraper helps |
|---|---|
| B2B prospecting in China | Find consultants, founders and specialists who publish a work email |
| Expert and KOC sourcing | Build lists of credible subject-matter voices for content collaborations |
| Technical recruiting | Source engineers, data scientists and product managers by discipline |
| Academic and research outreach | Reach researchers and educators active in a topic area |
| Partnership development | Identify Zhihu columnists open to 商务合作 |
| Market research | See how a professional category presents itself on Chinese social media |
| PR and thought leadership | Contact writers who already cover your category |
| CRM enrichment | Add public Zhihu contact emails to records you already hold |
Zhihu is the closest Chinese analogue to Quora, so the Quora Email Scraper is the natural companion when you need the same expert niche in English.
For a fuller picture of the Chinese ecosystem, run the Zhihu Email Scraper alongside the Weibo Email Scraper for microblogging KOLs and the WeChat Email Scraper for official-account publishers.
Why Use This Zhihu Email Scraper
Generic web crawlers point at arbitrary domains and hope for a hit. This is a purpose-built Zhihu scraper: the domain, username pattern, zhihu.com/people/ profile template and query modifiers are all configured for the platform.
Parsing is structural, not cosmetic. Because the Zhihu Email Scraper finds the <h3> and walks outward to the smallest surrounding block, a Google redesign degrades results gracefully instead of breaking the run.
Chinese-text handling is deliberate. Full-width at signs, zero-width characters and [at] obfuscation are normalised before matching — the exact places generic extractors lose leads on Chinese content.
Runs are resumable and checkpointed, and each row flags whether the snippet may have clipped the address.
Limitations of the Zhihu Email Scraper
Read this before buying. Every item here is a real constraint, not a bug.
- Only publicly indexed emails. The Zhihu Email Scraper finds addresses already visible in Google's index. Private data and login-gated content are out of reach.
- Google caps a single query at roughly 300 results. Query expansion exists specifically to work around that ceiling.
- Google indexes Zhihu less completely than Western platforms. Chinese sites are crawled more thinly, most indexed text is Chinese, and coverage shifts over time. Expect lower volume than a comparable Quora or LinkedIn run.
usernameandprofileUrlmay be empty. They are only filled when Google exposes a handle. Some rows will carryaccountNameandfullNameonly. That is a Google limitation.possiblyTruncated: truemeans verify. Google's snippet ellipsis may have clipped the address.- Requires the Apify
GOOGLE_SERPproxy. The Zhihu Email Scraper cannot run without Apify proxy credentials. - Free Apify plans are capped at 100 emails per run. Paid plans are uncapped.
- No volume guarantee. Results vary with keywords, domains, country code and location.
The Zhihu Email Scraper is not affiliated with, endorsed by or connected to Zhihu. Use the output in line with PIPL, GDPR and your own outreach policy.
Related Actors
| Actor | What it collects |
|---|---|
| Zhihu Email and Phone Number Scraper | Emails and phone numbers from Zhihu |
| Zhihu Phone Number Scraper | Public phone numbers from Zhihu |
| Behance Email Scraper | Public contact emails from Behance |
| Bigo Live Email Scraper | Public contact emails from Bigo Live |
| Bluesky Email Scraper | Public contact emails from Bluesky |
| Bumble Email Scraper | Public contact emails from Bumble |
| Clubhouse Email Scraper | Public contact emails from Clubhouse |
| Dailymotion Email Scraper | Public contact emails from Dailymotion |
| DeviantArt Email Scraper | Public contact emails from DeviantArt |
| Discord Email Scraper | Public contact emails from Discord |
| Dribbble Email Scraper | Public contact emails from Dribbble |
| Facebook Email Scraper | Public contact emails from Facebook |
| Goodreads Email Scraper | Public contact emails from Goodreads |
| Hinge Email Scraper | Public contact emails from Hinge |
| Instagram Email Scraper | Public contact emails from Instagram |
| KakaoTalk Email Scraper | Public contact emails from KakaoTalk |
| Kick Email Scraper | Public contact emails from Kick |
| Lemon8 Email Scraper | Public contact emails from Lemon8 |
| Likee Email Scraper | Public contact emails from Likee |
| LINE Email Scraper | Public contact emails from LINE |
| LinkedIn Email Scraper | Public contact emails from LinkedIn |
| Mastodon Email Scraper | Public contact emails from Mastodon |
| Medium Email Scraper | Public contact emails from Medium |
| Mixcloud Email Scraper | Public contact emails from Mixcloud |
| Patreon Email Scraper | Public contact emails from Patreon |
| Pinterest Email Scraper | Public contact emails from Pinterest |
| Quora Email Scraper | Public contact emails from Quora |
| Reddit Email Scraper | Public contact emails from Reddit |
| Rumble Email Scraper | Public contact emails from Rumble |
| Snapchat Email Scraper | Public contact emails from Snapchat |
| SoundCloud Email Scraper | Public contact emails from SoundCloud |
| Substack Email Scraper | Public contact emails from Substack |
| Telegram Email Scraper | Public contact emails from Telegram |
| Threads Email Scraper | Public contact emails from Threads |
| TikTok Email Scraper | Public contact emails from TikTok |
| Tinder Email Scraper | Public contact emails from Tinder |
| Tumblr Email Scraper | Public contact emails from Tumblr |
| Twitch Email Scraper | Public contact emails from Twitch |
| Vimeo Email Scraper | Public contact emails from Vimeo |
| VK Email Scraper | Public contact emails from VK |
| WeChat Email Scraper | Public contact emails from WeChat |
| Weibo Email Scraper | Public contact emails from Weibo |
Example Run
A realistic Zhihu Email Scraper input for sourcing independent consultants:
{"keywords": ["独立顾问", "品牌咨询", "consultant"],"location": "北京","customDomains": ["@qq.com", "@163.com", "@foxmail.com", "@gmail.com"],"maxEmails": 200,"countryCode": "HK","expandQueries": true,"queryModifiers": ["email", "contact", "business", "cooperation"],"maxPagesPerQuery": 30,"maxConcurrency": 5}
One dataset item from that run:
{"network": "Zhihu","keyword": "品牌咨询","query": "site:zhihu.com 品牌咨询 contact \"@163.com\" \"北京\"","title": "李文博 - 知乎","accountName": "li-wen-bo-consult","fullName": "李文博","username": "li-wen-bo-consult","profileUrl": "https://www.zhihu.com/people/li-wen-bo-consult","url": "https://www.zhihu.com/people/li-wen-bo-consult","description": "品牌战略顾问,北京。商务合作与咨询请邮件 liwenbo.brand@163.com","email": "liwenbo.brand@163.com","emailDomain": "@163.com","possiblyTruncated": false,"foundAt": "2025-03-14T09:41:22Z"}
Every row the Zhihu Email Scraper writes carries all fourteen fields, so the JSON shape is stable and safe to map into a database schema.
Leads are pushed to the dataset as soon as they are found, so you can begin reviewing results while the Zhihu Email Scraper is still running.
Zhihu Email Scraper FAQ
What is the Zhihu Email Scraper?
An Apify Actor that extracts publicly indexed contact emails from Zhihu search results and returns them as a structured dataset.
Does the Zhihu Email Scraper log into Zhihu?
No. It never opens zhihu.com, never authenticates and never uses Zhihu's API. All data comes from Google search results.
Can it scrape private profiles or hidden emails?
No. If an address is not publicly visible in Google's index, the Zhihu Email Scraper cannot find it.
Which email domains work best for Zhihu?
@qq.com, @163.com, @126.com, @foxmail.com and @sina.com are the highest-yield choices. Corporate and @edu.cn domains work well for professional niches.
Do Chinese-language keywords work?
Yes. Keywords are placed into the Google query as written, so 产品经理 and 商务合作 behave exactly like English terms.
Does it handle the full-width @ used in Chinese text?
Yes. Full-width at signs, zero-width characters and [at]/(at) obfuscation are normalised before matching.
How is Zhihu different from Quora for outreach?
Zhihu skews toward Chinese professionals and long-form answers, and contact lines are usually written in Chinese. Use the Zhihu Email Scraper for the Chinese market and the Quora Email Scraper for English-language experts.
What does possiblyTruncated mean?
Google's snippet ellipsis may have clipped the email. Verify those rows before adding them to a send list.
Why are some username and profileUrl values empty?
Google sometimes prints only a display name with no zhihu.com/people/ handle. The Zhihu Email Scraper fills those fields when a handle is exposed and leaves them empty otherwise.
Does the Zhihu Email Scraper need a proxy?
Yes. It requires the Apify GOOGLE_SERP proxy and cannot run without Apify proxy credentials.
How many emails can I collect per run?
Free Apify plans are capped at 100 emails per run. Paid plans are uncapped, up to the maxEmails value you set (maximum 10000).
Can I run the Zhihu Email Scraper on a schedule?
Yes. Use Apify Schedules or the API. Runs are resumable and state is checkpointed in the key-value store.
What export formats are supported?
The Zhihu Email Scraper writes to a standard Apify dataset, which exports to JSON, CSV, Excel, XML and HTML.
Leave a review
If the Zhihu Email Scraper saved you time, please leave a star rating and a short review on the Actor page.
Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.
If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.
Support
Questions, feature requests or need a custom build? Email neurodata.apify@gmail.com and we will get back to you.