# Changelog of Baidu News Scraper (`searchapi/baidu-news-scraper`) Actor

- **URL**: https://apify.com/searchapi/baidu-news-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/searchapi/baidu-news-scraper.md

## Changelog

### \[3.1.1] - 2026-09-06

#### Fixed

- Probe all bounded public-DNS answers independently before selecting a browser hostname mapping; this avoids a silent, unreachable Baidu endpoint causing repeated navigation timeouts.
- Classify Baidu interactive verification as `TARGET_VERIFICATION` in sanitized logs and retain safe HTTP/content-type/final-host diagnostics in each query summary.
- Read `Content-Type` case-insensitively while still validating status and media type before any page extraction.

### \[3.1.0] - 2026-09-06

#### Added

- Bounded public-DNS mappings for only the official `news.baidu.com`, `www.baidu.com`, and `wappass.baidu.com` hosts when the local resolver cannot resolve them.
- Direct persistent Playwright execution so browser host mappings, shared cookies, retries, pacing, and optional authorized proxies work together.
- `requestDelayMs`, source/title/snippet filters, deterministic `sortBy` options, and a validated `OUTPUT` key-value run summary.

#### Fixed

- Baidu News no longer reports a redirect-host DNS failure before it can classify the target response.
- Dataset audit tooling now handles runs that have no dataset directory, including fail-closed verification runs.
- Pinned Browserslist to a patched release; Crawlee's upstream moderate `stream-json` advisory remains without a published fix.

### \[3.0.2] - 2026-08-29

#### Added

- Single and fair batch query modes with total/per-query limits and bounded query concurrency.
- Stable SHA-256 article IDs, query/global/result positions, extraction provenance, and a valid linked 56-field dataset schema.
- Consistent Chromium fingerprints, persistent sessions, optional China Residential/custom proxy support, and sanitized structural debug logs.
- All-or-nothing dataset writes, recursive empty-field removal, dataset audit tooling, and coverage for batch fairness and pagination-link validation.

#### Changed

- Migrated from the generic `www.baidu.com/s?tn=news` start route to Baidu's public `news.baidu.com/ns` News interface.
- Pagination now follows the website's own link after validating the exact Baidu host/path, query, News template, and increasing offset.
- Aligned Apify SDK 3.7.2, Crawlee 3.18.1, Playwright 1.61.1, and the Playwright Chrome Docker image.
- Proxy-disabled input no longer initializes Apify Proxy; `GOOGLE_SERP` is rejected as Google-only.
- Sanitized highlighted HTML and added a bounded structural snippet fallback.

#### Fixed

- Prevented user query URLs and raw error payloads from reaching logs.
- Prevented partial datasets when a later page or batch query is challenged.
- Corrected `publishedAtRaw`/`relativeTime` normalization and removed fabricated false ad/sponsored flags.
- Corrected live 10-item pagination and restored 100% snippet/source/date coverage in representative runs.

### \[2.0.0] - 2026-06-12

#### Added

- **47 canonical output fields** grouped by category, produced via the shared `normalizeNewsRecord` in `_canonical/normalize.js`:
  - **Identity (4)**: `type`, `resultType`, `position`, `page`
  - **Title / text (9)**: `title`, `titleHtml`, `snippet`, `snippetHtml`, `description`, `descriptionHtml`, `emphasis`, `link`, `url`
  - **Source (5)**: `source`, `sourceUrl`, `sourceDomain`, `domain`, `displayUrl`
  - **Media (3)**: `thumbnailUrl`, `imageUrl`, `images` (gallery array)
  - **Authorship (3)**: `author`, `authorUrl`, `tags`
  - **Dates (3)**: `publishedAt`, `publishedAtRaw`, `relativeTime`
  - **Engagement (2)**: `commentCount`, `shareCount`
  - **Flags (4)**: `isBreaking`, `isPaywall`, `isSponsored`, `isAd`
  - **Engine identity (2)**: `engineId`, `engineData`
  - **Search context (8)**: `searchQuery`, `query`, `searchUrl`, `market`, `country`, `language`, `scrapedAt`, `searchMetadata`
  - **Branding (1)**: `favicon`
  - **Categorization (1)**: `category`
- **Shared normalizer**: uses `normalizeNewsRecord` from `_canonical/normalize.js`
- **Chinese date parser**: `parseChineseDate()` handles "X分钟前", "X小时前", "X天前", "刚刚", "昨天", "前天", "X年X月X日", "前天 HH:MM"
- **`isBreaking` detection**: scans for "最新", "置顶", "突发" markers in the result card
- **Multi-image gallery extraction**: collects all `<img>` URLs from the result card into `images[]`
- **Baidu redirect resolution**: HEAD request follows `https://www.baidu.com/link?url=...` to the real article URL
- **2 views**: `overview` (key fields) and `fullDetails` (all fields)
- **New input fields**: `startUrls`, `maxPages`, `resolveRedirects`, `market`, `category`

#### Changed

- **Breaking** — actor version bumped to 2.0.0
- **Breaking** — schema expanded from ~10 fields to 47 fields
- Cleaner `main.js` — extraction moved to `routes/handlers.js`
- `package.json` no longer runs `crawlee install-playwright-browsers` postinstall
- Pagination respects `maxPages` cap

### \[1.0.0] - 2025-07-01

#### Added

- Initial release of the actor
- Basic news extraction with redirect resolution
