# Haodf Scraper: 好大夫在线 Doctor & Hospital Reviews (`getascraper/haodf-scraper`) Actor

Scrape Haodf (好大夫在线) doctor profiles, aggregate reputation stats, and real per-condition patient reviews (effect/attitude/skill ratings, anonymized patient location, dates). No login or API key needed. Includes a new-reviews monitor mode for tracking a shortlist of doctors over time.

- **URL**: https://apify.com/getascraper/haodf-scraper.md
- **Developed by:** [GetAScraper](https://apify.com/getascraper) (community)
- **Categories:** Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 75.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.88 / 1,000 doctor records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Haodf Scraper: 好大夫在线 Doctor & Hospital Reviews

<table width="100%">
<tr>
<td style="padding:24px 28px;background:#F0FDFA;border:1px solid #99F6E4;border-top:4px solid #0F766E;border-radius:12px">
<span style="font-size:23px;font-weight:800;color:#1C1917;line-height:1.3">Real patient reviews from 好大夫在线, structured and ready to use</span><br>
<span style="font-size:15px;color:#57534E;line-height:1.6">Pull doctor profiles, per-condition patient ratings, and hospital stats from China's largest doctor review platform. No login, no account, no API key, just clean structured data.</span>
</td>
</tr>
</table>

<table width="100%">
<tr>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #99F6E4;border-radius:10px 0 0 10px;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#115E59">🩺 Real per-condition ratings</span><br>
<span style="font-size:12px;color:#57534E">Effect, attitude, and skill scores broken out by the exact condition each patient was treated for, not one blended average.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #99F6E4;border-left:none;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#115E59">🔎 Two ways to target</span><br>
<span style="font-size:12px;color:#57534E">Auto-discover doctors by medical department, or point the Actor at specific doctor URLs or IDs.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #99F6E4;border-left:none;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#115E59">🔔 Monitor mode</span><br>
<span style="font-size:12px;color:#57534E">Track a shortlist of doctors across repeat runs and get only the reviews that are genuinely new since last check.</span>
</td>
<td style="padding:14px 12px;width:25%;background:#FFFFFF;border:1px solid #99F6E4;border-left:none;border-radius:0 10px 10px 0;vertical-align:top">
<span style="font-size:15px;font-weight:800;color:#115E59">🏥 Optional hospital data</span><br>
<span style="font-size:12px;color:#57534E">Add hospital grade, doctor and faculty counts, and service-patient volume without slowing down a base run.</span>
</td>
</tr>
</table>

好大夫在线 (Haodf) is China's largest doctor and hospital review platform, covering hundreds of thousands of doctors across every province and specialty. This Actor turns its public doctor profiles and patient reviews into a clean, structured dataset: no login, no account, and no API key required to run it.

### 🙋 Why use this Actor

- **I am a hospital marketing manager** benchmarking my own doctors' patient recommend scores and review volume against competing hospitals in the same department, so I know where we're actually losing reputation ground.
- **I am a pharma or medical device field rep** identifying the highest-volume, highest-rated specialists treating a given condition, so I can prioritize which doctors are worth an outreach visit.
- **I am a medical tourism agency coordinator** matching international patients to specialists with a track record of real, patient-reviewed outcomes, instead of relying on a hospital's own marketing copy.
- **I am an academic researcher** pulling structured patient review text and ratings at scale for sentiment or health-outcomes research, the same kind of public review corpus that has already been used in published studies analyzing patient sentiment in China.

### 🚀 How to use it

<table width="100%">
<tr>
<td style="padding:16px 14px;width:33%;background:#F0FDFA;border:1px solid #99F6E4;border-radius:10px 0 0 10px;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#0F766E;letter-spacing:1px">STEP 1</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Choose your doctors</span><br>
<span style="font-size:12px;color:#57534E">Pick a department to auto-discover doctors, or paste specific doctor URLs or IDs.</span>
</td>
<td style="padding:16px 14px;width:33%;background:#F0FDFA;border:1px solid #99F6E4;border-left:none;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#0F766E;letter-spacing:1px">STEP 2</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Run the Actor</span><br>
<span style="font-size:12px;color:#57534E">It pulls each doctor's profile, reputation stats, and their most-reviewed conditions.</span>
</td>
<td style="padding:16px 14px;width:33%;background:#F0FDFA;border:1px solid #99F6E4;border-left:none;border-radius:0 10px 10px 0;vertical-align:top">
<span style="font-size:12px;font-weight:800;color:#0F766E;letter-spacing:1px">STEP 3</span><br>
<span style="font-size:14px;font-weight:700;color:#1C1917">Get three ready-to-use views</span><br>
<span style="font-size:12px;color:#57534E">Doctor profiles, individual patient reviews, and department-level breakdowns.</span>
</td>
</tr>
</table>

Want to track the same doctors over time instead of a one-off pull? Turn on **Only new reviews since last check** and rerun the Actor later with the same **State name**. Only reviews that weren't there last time come back.

### 📥 Input

| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `doctorUrls` | array of URLs | No | Direct Haodf doctor profile links or their bare numeric IDs, one per line. The cheapest, most reliable way to target specific doctors. Leave empty to auto-discover doctors from `department` instead. |
| `department` | enum | No | Which Haodf department directory to auto-discover doctors from, only used when `doctorUrls` is empty. Defaults to Respiratory Medicine (呼吸内科). Any valid Haodf department slug works, not only the suggested list. |
| `proxyConfiguration` | object | No | Apify Proxy settings. Defaults to the cheaper datacenter proxy, which was confirmed to reach Haodf's doctor, review, and hospital pages without issue. |
| `includeHospitalDetails` | boolean | No | Fetch each doctor's hospital page once (deduplicated per hospital) to add hospital grade, total doctor and faculty counts, and service-patient volume. Defaults to false to keep runs fast and cheap. |
| `maxDoctors` | integer | No | Maximum number of doctors to scrape in this run, whether sourced from `doctorUrls` or department auto-discovery. Defaults to 15. |
| `maxDiseaseTagsPerDoctor` | integer | No | Caps how many of a doctor's most-reviewed conditions get fetched, each one is a separate set of aggregate stats and up to 10 reviews. Defaults to 5. |
| `onlyNewReviews` | boolean | No | Turns the run into a monitor: only reviews not seen before for this doctor and condition under `stateName` are pushed to the Reviews dataset. Defaults to false. |
| `stateName` | string | No | Name for this tracked doctor set's saved monitor history, so multiple independent trackers (one per client or department, for example) don't overwrite each other. Only relevant when `onlyNewReviews` is on. Defaults to `default`. |
| `resetState` | boolean | No | Clears the saved seen-review history for `stateName` before the run starts, so every review currently on the page counts as new again. Defaults to false. |

### 📊 Data table

#### Doctors (default view)

| Field | Type | Description |
| --- | --- | --- |
| `doctorId` | string | The doctor's numeric Haodf ID. |
| `doctorUrl` | string | Link to the doctor's public Haodf profile. |
| `name` | string | Doctor's name. |
| `title` | string | Professional title, for example 主任医师 (Chief Physician) or 副主任医师 (Associate Chief Physician). |
| `hospitalId` | string | The affiliated hospital's Haodf ID. |
| `hospitalName` | string | Affiliated hospital name. |
| `hospitalUrl` | string | Link to the hospital's public Haodf profile. |
| `departmentName` | string | Department name within the hospital. |
| `departmentUrl` | string | Link to the department's Haodf page. |
| `specialty` | string | The doctor's stated specialty focus. |
| `recommendScore` | number | Patient recommend score out of 5, as shown on the doctor's own profile. |
| `totalPatientsHelped` | number | Cumulative patients-helped count reported on the profile. |
| `totalArticles` | number | Number of educational articles the doctor has published. |
| `totalThanksAndVotes` | number | Combined count of patient thank-you notes and votes. |
| `awards` | array of strings | Doctor of the Year (年度好大夫) awards, one entry per year won. |
| `diseaseTagCount` | number | Total number of condition tags listed on the doctor's profile. |
| `diseaseTagsFetched` | array of strings | Which of those condition tags this run actually fetched reviews for. |

**Example output**

```json
{
  "doctorId": "11908",
  "doctorUrl": "https://www.haodf.com/doctor/11908.html",
  "name": "王建国",
  "title": "主任医师",
  "hospitalId": "5678",
  "hospitalName": "北京协和医院",
  "hospitalUrl": "https://www.haodf.com/hospital/5678.html",
  "departmentName": "呼吸内科",
  "departmentUrl": "https://www.haodf.com/keshi/xxxxx.html",
  "specialty": "支气管哮喘、慢性阻塞性肺疾病诊疗",
  "recommendScore": 4.5,
  "totalPatientsHelped": 9459,
  "totalArticles": 44,
  "totalThanksAndVotes": 1213,
  "awards": ["2023年度好大夫", "2024年度好大夫"],
  "diseaseTagCount": 19,
  "diseaseTagsFetched": ["哮喘", "慢性阻塞性肺疾病", "肺结节"]
}
```

#### Reviews view

| Field | Type | Description |
| --- | --- | --- |
| `doctorId` | string | The reviewed doctor's Haodf ID. |
| `doctorName` | string | The reviewed doctor's name. |
| `hospitalName` | string | The doctor's affiliated hospital. |
| `diseaseKey` | string | Internal slug for the condition this review was left under. |
| `diseaseName` | string | Condition name in Chinese, for example 哮喘 (Asthma). |
| `reviewId` | string | Stable review ID, safe to use as a unique key across runs. |
| `content` | string | The patient's full written review. |
| `effect` | string | Patient rating of treatment effect. |
| `attitude` | string | Patient rating of the doctor's attitude. |
| `skill` | string | Patient rating of the doctor's skill. |
| `remedy` | string | Treatment or remedy context the patient mentioned. |
| `patientProvince` | string | Patient's province, as pre-anonymized by Haodf itself. |
| `patientCity` | string | Patient's city, as pre-anonymized by Haodf itself. |
| `reviewDate` | string | When the review was posted. |
| `isNew` | boolean | True when this review was not seen in a prior run under the same monitor state. |

**Example output**

```json
{
  "doctorId": "11908",
  "doctorName": "王建国",
  "hospitalName": "北京协和医院",
  "diseaseKey": "xiaochuan",
  "diseaseName": "哮喘",
  "reviewId": "48213092",
  "content": "医生非常耐心，详细解释了病情和用药方案，复诊后症状明显改善。",
  "effect": "显效",
  "attitude": "很好",
  "skill": "很好",
  "remedy": "吸入用药调整",
  "patientProvince": "北京",
  "patientCity": "朝阳",
  "reviewDate": "2026-08-02",
  "isNew": true
}
```

#### Department breakdown view

| Field | Type | Description |
| --- | --- | --- |
| `dimension` | string | What this row summarizes, for example `department` or `hospital`. |
| `label` | string | The department or hospital name for this row. |
| `doctorCount` | number | Number of doctors scraped under this dimension in the run. |
| `reviewCount` | number | Number of reviews fetched under this dimension in the run. |
| `avgRecommendScore` | number | Average patient recommend score across the doctors in this row. |
| `hospitalGrade` | string | Official hospital grade, for example 三甲 (Tier 3A), only present when hospital enrichment is on. |
| `totalDoctorsInHospital` | number | Hospital-wide doctor count, only present when hospital enrichment is on. |
| `totalFacultiesInHospital` | number | Hospital-wide faculty count, only present when hospital enrichment is on. |
| `servicePatientCount` | string | Hospital-wide cumulative service-patient count, as reported by the hospital's own page. |

**Example output**

```json
{
  "dimension": "department",
  "label": "呼吸内科",
  "doctorCount": 15,
  "reviewCount": 63,
  "avgRecommendScore": 4.4,
  "hospitalGrade": "三甲",
  "totalDoctorsInHospital": 312,
  "totalFacultiesInHospital": 42,
  "servicePatientCount": "1,200,000+"
}
```

### 💰 Pricing

Pricing is pay per result, billed on the doctor and review records actually saved to your dataset. Empty runs cost nothing. There are no fixed monthly subscriptions or hidden maintenance fees.

### ⭐ Enjoying Haodf Scraper?

<table width="100%">
<tr>
<td style="padding:20px 24px 14px;background:#F0FDFA;border:1px solid #99F6E4;border-left:5px solid #0F766E;border-radius:10px 10px 0 0">
<span style="font-size:20px;letter-spacing:4px">⭐ ⭐ ⭐ ⭐ ⭐</span><br>
<span style="font-size:17px;font-weight:800;color:#1C1917">If this saved you hours of manually checking doctor pages one by one, let us know.</span><br>
<span style="font-size:14px;color:#57534E">A 5-star rating takes 10 seconds and helps other hospital marketing teams and field researchers find this Actor. Your feedback also tells us what to build next.</span>
</td>
</tr>
<tr>
<td style="padding:0;background:#0F766E;border:1px solid #99F6E4;border-top:none;border-radius:0 0 10px 10px;text-align:center">
<a href="https://apify.com/getascraper/haodf-scraper/reviews" style="display:block;padding:13px 16px;color:#FFFFFF;text-decoration:none;font-weight:800;font-size:15px;letter-spacing:0.3px">★&nbsp;&nbsp;Rate this Actor on Apify</a>
</td>
</tr>
</table>

### 💡 Tips

- Start with a low `maxDoctors` and `maxDiseaseTagsPerDoctor` to see the shape of the data before scaling up to a full department.
- Turn on `includeHospitalDetails` only when you actually need hospital-level context. It adds one extra request per unique hospital, not per doctor.
- For ongoing tracking, give each tracked doctor set its own `stateName` so a hospital-benchmarking tracker and a KOL-monitoring tracker never overwrite each other's history.

### ❓ FAQ

##### 好大夫在线怎么爬取医生评价数据?

You can pull doctor profiles and patient reviews from 好大夫在线 with this Actor. Point it at a department or specific doctor links, run it, and get structured doctor, review, and hospital data back with no login required.

##### 如何查询医生的患者推荐度?

Each doctor record includes `recommendScore`, the patient recommend score shown on the doctor's own Haodf profile, along with total patients helped and total patient thanks and votes for further context.

##### Does this Actor need a Haodf account or login?

No. Every page this Actor reads is public. It never logs in, uses a cookie, or requires an API key from Haodf or from you.

##### Does it extract patient names or private contact details?

No. Haodf pre-anonymizes patient reviews itself, showing only province and city, never a name or contact detail. This Actor passes that anonymization through unchanged and never attempts to identify a patient.

##### How fresh is the data?

Every run reads live pages at request time. There is no cached or pre-scraped snapshot, so results reflect what is on Haodf right now.

##### Can I track new reviews instead of re-pulling everything each time?

Yes. Turn on `onlyNewReviews` and rerun the Actor later against the same doctors and `stateName`. Only reviews that are new since the last run are returned.

### 🔗 Other actors

- [CNKI Scraper: 中国知网 Citations, Rankings & Academic Search](https://apify.com/getascraper/cnki-scraper) ↗ - pulls citation counts, rankings, and academic search results from China's largest academic database.
- [Zhaopin Scraper 智联招聘: China Jobs & Salary API](https://apify.com/getascraper/zhaopin-jobs-scraper) ↗ - scrapes job listings and salary data from one of China's largest recruitment platforms.
- [Lagou Tech Jobs Scraper: China IT Recruitment Data](https://apify.com/getascraper/lagou-tech-jobs-scraper) ↗ - collects tech job postings and compensation data from China's leading IT recruitment site.
- [China Recall Scraper: 中国缺陷产品召回](https://apify.com/getascraper/china-samr-product-recall-scraper) ↗ - monitors official Chinese government product recall notices by brand and category.
- [Employee Reviews Scraper for TeamBlind](https://apify.com/getascraper/teamblind-reviews-scraper) ↗ - pulls anonymous employee reviews and company ratings for reputation research.

# Actor input Schema

## `doctorUrls` (type: `array`):

Direct Haodf doctor profile links or their bare numeric IDs, one per line (both https://www.haodf.com/doctor/11908.html and 11908 work). This is the cheapest, most reliable way to target specific doctors. Leave empty to auto-discover doctors from the Department field below instead.

## `department` (type: `string`):

Only used when Doctor URLs above is empty. Picks which Haodf department directory page to auto-discover doctors from. You can also type in any other department slug from Haodf's own department picker, not just the ones suggested here.

## `proxyConfiguration` (type: `object`):

Haodf's doctor, review, and hospital pages were confirmed to return clean, unblocked responses across real-Chrome, Googlebot, and empty User-Agent tests, so Apify's cheaper datacenter proxy is used by default. Residential proxy is not needed for this target.

## `includeHospitalDetails` (type: `boolean`):

Fetch each doctor's hospital profile page once (deduplicated across doctors sharing a hospital) to add hospital grade, total doctor/faculty counts, and hospital-wide service-patient counts to the Department Breakdown view. Adds one extra request per unique hospital, so it defaults off to keep runs fast and cheap.

## `maxDoctors` (type: `integer`):

Maximum number of doctors to scrape in this run, whether they come from Doctor URLs or from department auto-discovery. Keep this low for a quick test run; raise it once you know the run time and cost you're comfortable with.

## `maxDiseaseTagsPerDoctor` (type: `integer`):

Most active doctors on Haodf are reviewed under several different conditions (disease tags), each with its own review page. This caps how many of a doctor's most-reviewed disease tags get fetched, since each tag is one additional request yielding its own aggregate stats and up to 10 reviews.

## `onlyNewReviews` (type: `boolean`):

Turns this run into a monitor: only reviews whose ID has not been seen before for this doctor and disease tag under State Name below are pushed to the Reviews dataset. Run this Actor again later on a schedule with the same State Name to get alerted only to genuinely new patient reviews.

## `stateName` (type: `string`):

Identifies this tracked doctor set's saved monitor state, so you can run several independent trackers (e.g. one per client or department) without them overwriting each other's seen-review history. Only relevant when Only new reviews since last check is enabled.

## `resetState` (type: `boolean`):

Clears the saved seen-review history for State Name before this run starts, so every review currently on the page is treated as new again. Use this once if you want to restart tracking from scratch.

## Actor input object example

```json
{
  "doctorUrls": [],
  "department": "huxineike",
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "includeHospitalDetails": false,
  "maxDoctors": 15,
  "maxDiseaseTagsPerDoctor": 5,
  "onlyNewReviews": false,
  "stateName": "default",
  "resetState": false
}
```

# Actor output Schema

## `doctors` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "doctorUrls": [],
    "department": "huxineike",
    "proxyConfiguration": {
        "useApifyProxy": true
    },
    "includeHospitalDetails": false,
    "maxDoctors": 15,
    "maxDiseaseTagsPerDoctor": 5,
    "onlyNewReviews": false,
    "stateName": "default",
    "resetState": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("getascraper/haodf-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "doctorUrls": [],
    "department": "huxineike",
    "proxyConfiguration": { "useApifyProxy": True },
    "includeHospitalDetails": False,
    "maxDoctors": 15,
    "maxDiseaseTagsPerDoctor": 5,
    "onlyNewReviews": False,
    "stateName": "default",
    "resetState": False,
}

# Run the Actor and wait for it to finish
run = client.actor("getascraper/haodf-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "doctorUrls": [],
  "department": "huxineike",
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "includeHospitalDetails": false,
  "maxDoctors": 15,
  "maxDiseaseTagsPerDoctor": 5,
  "onlyNewReviews": false,
  "stateName": "default",
  "resetState": false
}' |
apify call getascraper/haodf-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,getascraper/haodf-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cupyY0NRBcI3lLNRF/builds/nbPjjegfgAyqGg9z4/openapi.json
