# IT之家 Scraper — China Tech News & Stats (`hipersoft/ithome-scraper`) Actor

Scrape IT之家 (ithome.com) tech news across channels: title, summary, views, comments, keywords, image and URL. Android, Apple, Windows, digital, games, AI and more.

- **URL**: https://apify.com/hipersoft/ithome-scraper.md
- **Developed by:** [hiper soft](https://apify.com/hipersoft) (community)
- **Categories:** News, Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.001 / article scraped

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## IT Home Scraper — China Tech News, Views, Comments & Metadata (ithome.com)

**Scrape IT Home** (IT之家, ithome.com) — one of China's biggest consumer-tech news sites — into clean JSON, CSV, Excel or XML. This IT Home scraper extracts **article titles, summaries, view counts, comment counts, keyword tags, cover images, posted dates and URLs** across the Top News, Android, Digital and Games channels — all from one Actor, no code and no login required. A fast, cheap, no-code way to turn ITHome's front pages into a structured tech-news dataset.

Perfect for **China tech-news monitoring, product-launch tracking, trend research, price-drop alerts and media intelligence** — and one of the cheapest ways to do it, since you pay only per article you get.

![IT Home Scraper input — channels and options in the Apify Console](https://api.apify.com/v2/key-value-stores/SUfvnaFLd9z9eBCtV/records/ithome-scraper-input.png)

### What does the IT Home Scraper do?

The IT Home Scraper crawls and extracts public news data from **IT之家 (ithome.com)** by channel. Pick any mix of the **Top News, Android, Digital and Games** channels, set how many articles you want, and it returns structured data points — title, summary, views, comments, keywords, image, posted date and canonical URL — ready for analytics, dashboards, datasets, LLM pipelines or trend research. It's a no-code crawler and extractor for Chinese consumer-tech coverage, with no API key to manage.

### What data can you scrape from IT Home?

| Group | Fields |
|---|---|
| 📰 **Article** | id, title, description/summary, canonical **URL**, cover **image** |
| 📊 **Engagement** | **view count** (阅读), **comment count** per article |
| 🏷️ **Tags** | keyword tags attached to the article |
| 🔀 **Channel** | source channel — Top News (资讯), Android (安卓), Digital (数码), Games (游戏) |
| 🕒 **Timestamps** | posted date, collected-at timestamp |

### Use cases

- **Tech-news monitoring** — track new coverage across IT之家's Top News, Android, Digital and Games channels in near real time.
- **Product-launch & rumor tracking** — catch phone, chip and gadget announcements as soon as they hit the front page.
- **Trend & topic research** — mine keyword tags, views and comments to see what Chinese tech readers care about.
- **Build a tech-news dataset** — thousands of articles with engagement stats for analytics, BI or model training.
- **LLM & RAG pipelines** — feed clean summaries and metadata into knowledge bases and summarization workflows.
- **Competitive & market intelligence** — watch how brands and products are covered and how much engagement they pull.
- **Content curation & newsletters** — pull the most-read stories per channel to power a digest.
- **Price & deal tracking** — surface pricing and pre-order stories the moment they publish.

### How to scrape IT Home data

1. Click **Try for free / Start** to open the IT Home Scraper.
2. Choose the **channels** you want — Top News, Android, Digital and/or Games.
3. Set **Max articles** to cap how many results you collect across all selected channels.
4. Click **Run**.
5. Download the results as **JSON, CSV, Excel or XML**, or pull them from the API.

### Input

Pick one or more channels and a limit — that's it. No code, no login, no setup.

```json
{
  "channels": ["news", "android", "digi", "game"],
  "maxItems": 200
}
```

| Field | Type | Description |
|---|---|---|
| `channels` | array | Which IT之家 channels to scrape. Options: `news` (Top News 资讯), `android` (Android 安卓), `digi` (Digital 数码), `game` (Games 游戏). Pass any combination. |
| `maxItems` | integer | Maximum number of articles to collect across all selected channels (default 200). |

#### Choose your channels

Select any mix of the four channels — `news`, `android`, `digi`, `game`. Each maps to an IT之家 section, and articles are tagged with the `channel` they came from so you can filter or split the dataset later.

#### Set how many articles you get

Use `maxItems` to control run size and cost. Since you pay per article, a smaller limit is a cheap way to sample a channel before scaling up to a full dataset.

### Output

![IT Home Scraper output — a tech-news dataset with views, comments and metadata](https://api.apify.com/v2/key-value-stores/SUfvnaFLd9z9eBCtV/records/ithome-scraper-output.png)

Each article is one dataset item:

```json
{
  "id": 983883,
  "title": "25.99 万元！小米澎程 N70 Max 预售价格公布",
  "description": "……",
  "channel": "news",
  "keywords": ["闪讯"],
  "views": 2081,
  "comments": 0,
  "image": "https://img.ithome.com/newsuploadfiles/thumbnail/....jpg",
  "postedAt": "2026-07-30T18:51:05.443Z",
  "url": "https://www.ithome.com/0/983/883.htm",
  "collectedAt": "2026-07-30T12:00:00.000Z"
}
```

#### Output schema

| Field | Type | Description |
|---|---|---|
| `id` | integer | Unique IT之家 article ID. |
| `title` | string | Article headline. |
| `description` | string | Article summary / lead text. |
| `channel` | string | Source channel — `news`, `android`, `digi` or `game`. |
| `keywords` | array | Keyword tags attached to the article. |
| `views` | integer | View count (阅读) for the article. |
| `comments` | integer | Number of comments on the article. |
| `image` | string (URL) | Cover / thumbnail image URL. |
| `postedAt` | string (ISO date) | When the article was published. |
| `url` | string (URL) | Canonical article URL on ithome.com. |
| `collectedAt` | string (ISO date) | When this record was collected. |

### Need more Chinese tech & community data?

Building a wider China tech dataset? Pair the IT Home Scraper with our other scrapers:

- [Bilibili Scraper](https://apify.com/hipersoft/bilibili-scraper) — videos, channels and stats from China's biggest video community.
- [SSPAI Scraper](https://apify.com/hipersoft/sspai-scraper) — apps, gear and productivity articles from 少数派.
- [Juejin Scraper](https://apify.com/hipersoft/juejin-scraper) — developer articles, tags and engagement from 掘金.
- [Guokr Scraper](https://apify.com/hipersoft/guokr-scraper) — science and tech explainers from 果壳.

### FAQ

**Do I need an API key or login to scrape IT Home?**
No. There's no setup, no key and no login — just pick your channels, set a limit and run.

**How many articles can I scrape per run?**
As many as you like. Use `maxItems` to cap results across all selected channels, and scale up to thousands when you need a full dataset.

**Which channels can I scrape?**
Top News (资讯), Android (安卓), Digital (数码) and Games (游戏). Select any combination in the `channels` input.

**How does billing work?**
You pay per article you collect, so cost scales with your `maxItems` and channel selection — one of the cheapest ways to build a China tech-news dataset.

**What export formats are supported?**
JSON, CSV, Excel and XML, plus webhooks and the Apify API.

**How fresh is the data?**
Each run pulls the latest articles currently published on the selected IT之家 channels, with a `postedAt` date and a `collectedAt` timestamp on every item.

**Can I filter by channel?**
Yes — choose exactly the channels you want in the input, and every item is tagged with its source `channel` so you can split or filter the dataset afterward.

**Can I automate or integrate the IT Home Scraper?**
Yes. The IT Home Scraper can be connected with almost any cloud service or web app thanks to [integrations on the Apify platform](https://apify.com/integrations). It works with [Make](https://apify.com/integrations/make), [Zapier](https://apify.com/integrations/zapier), [Slack](https://docs.apify.com/platform/integrations/slack), [Airbyte](https://docs.apify.com/platform/integrations/airbyte), [GitHub](https://docs.apify.com/platform/integrations/github), [Google Drive](https://docs.apify.com/platform/integrations/drive) and [many more](https://apify.com/integrations), plus the [Apify API](https://docs.apify.com/api/v2), JavaScript/Python clients and MCP. Or use [webhooks](https://docs.apify.com/platform/integrations/webhooks) to trigger an action whenever a run finishes — get a notification, or kick off another process such as loading your data downstream.

**Is scraping IT Home legal?**
The Actor collects only public data. You are responsible for how you use it and for complying with IT之家's terms and applicable laws.

### Related Actors

- [Bilibili Scraper](https://apify.com/hipersoft/bilibili-scraper) — videos, channels and view stats.
- [SSPAI Scraper](https://apify.com/hipersoft/sspai-scraper) — apps, gear and productivity articles.
- [Juejin Scraper](https://apify.com/hipersoft/juejin-scraper) — developer articles and engagement.
- [Guokr Scraper](https://apify.com/hipersoft/guokr-scraper) — science and tech explainers.

### Notes

Original clean-room implementation. Returns only public data; you are responsible for how you use the data and for complying with IT之家's terms. This is an independent tool and is not affiliated with or endorsed by IT Home (IT之家 / ithome.com).

# Actor input Schema

## `channels` (type: `array`):

Which IT之家 news channels to scrape.

## `maxItems` (type: `integer`):

Maximum articles to scrape across all selected channels.

## Actor input object example

```json
{
  "channels": [
    "news"
  ],
  "maxItems": 200
}
```

# Actor output Schema

## `results` (type: `string`):

The scraped results as dataset items.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("hipersoft/ithome-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("hipersoft/ithome-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call hipersoft/ithome-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hipersoft/ithome-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/d5KHHdWiQlkQJ3R7L/builds/zD0PsFlNMejOgh4fl/openapi.json
