# Changelog of Baidu Search Scraper (`searchapi/baidu-search-scraper`) Actor

- **URL**: https://apify.com/searchapi/baidu-search-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/searchapi/baidu-search-scraper.md

## Changelog

### \[3.2.2] - 2026-09-06

- Classify Baidu CAPTCHA/verification failures as `TARGET_VERIFICATION` in sanitized logs and per-query `OUTPUT` diagnostics.
- Preserve safe HTTP status, response content type, final page host, and `searchResultCountText` observations in the run summary without storing challenge payloads.

### \[3.2.1] - 2026-09-06

- Replaced the Crawlee navigation layer with a bounded direct Playwright runner so validated Baidu host mappings are effective in this environment while shared cookies, controlled concurrency, retries, optional authorized proxies, and pagination remain supported.
- Added DNS-over-HTTPS validation before the system-DNS fallback and classified Baidu verification interstitials as target challenges without fabricating records.
- Hardened `searchResultCountText` extraction with known-selector and bounded text-pattern fallbacks, preserving the first valid count across pages.
- Corrected fair multi-query output so `queryPosition` identifies the source query rather than the per-query result ordinal.

### \[3.2.0] - 2026-09-06

- Added bounded public-DNS browser mappings for official `www.baidu.com` and `wappass.baidu.com` when the local resolver cannot resolve them.
- Avoided initializing Apify Proxy when the proxy editor is explicitly disabled.
- Added DNS mode to the safe run summary and pinned Browserslist to a patched release.

### \[3.1.0] - 2026-09-05

- Added typed result-type, domain, keyword, rating, and sort filters.
- Added safe HTTP/content-type provenance, displayed result-count text, extraction method, and richer dataset fields.
- Added media-route cleanup, consistent browser fingerprints, developer controls, output/KVS schemas, and validated run summaries.
- Buffered query results so multi-query relevance output is fair and global ordering/deduplication happens before dataset writes.

### \[3.0.0] - 2026-08-29

#### Added

- Stable SHA-256 IDs, global/query/page positions, global deduplication, and recursive empty omission.
- Deduplicated multi-query mode with fair per-query quotas and bounded browser concurrency.
- Explicit `pn` pagination, configurable retries/timeouts, and validated Baidu-only legacy start URLs.
- External typed 44-field dataset schema and safe structural status logging.
- Chinese Residential proxy input and explicit rejection of the Google-only `GOOGLE_SERP` group.

#### Changed

- Aligned Playwright 1.61.1 with the Chrome 1.61.1 cloud image and exact Apify/Crawlee dependencies.
- Replaced random Chrome version rotation and webdriver-disabling flags with one internally consistent profile.
- Reduced blind page waits and skipped media/font resources without removing result images from the DOM.
- Pagination no longer depends on a drifting Next link; pages use Baidu's verified ten-result offset.
- Results are buffered until the requested workflow succeeds, so a later challenge cannot leave a misleading partial dataset.

#### Removed

- Full request/result URL and exception-message logging.
- Fabricated Google favicon URLs for Baidu results.
- Automatic or Google-specific proxy assumptions.

### \[2.0.1] - 2026-06-13

#### Fixed

- **Crawlee validation error**: `userAgent` was being passed at the top level of `PlaywrightCrawler` options, which is no longer valid in Crawlee 3.x. Moved into `launchContext.userAgent` so the schema validation passes and the actor actually starts.
- **Extra headers**: Added `preNavigationHooks` to send `Accept-Language: zh-CN,zh;q=0.9,en;q=0.8` and Chinese `User-Agent` to look more like a real Chinese browser. Reduces captcha likelihood.

### \[2.0.0] - 2026-06-12

#### Added

- New dataset fields: `type`, `resultType`, `page`, `url`, `description` (snippet alias), `domain`, `favicon`, `thumbnail`, `titleHtml`, `snippetHtml`, `emphasis`, `siteLinks`, `breadcrumbs`, `richResultType`, `isAd`, `isSponsored`, `isOfficial`, `author`, `publishedAt`, `publishedAtRaw`, `relativeTime`, `rating`, `ratingCount`, `price`, `priceNumeric`, `currency`, `searchQuery`, `query`, `searchUrl`, `market`, `country`, `language`, `searchMetadata`
- **Shared normalizer**: uses `normalizeOrganicRecord` from `_canonical/normalize.js` for cross-engine consistency
- **Smart date parsing**: handles Chinese relative dates ("X小时前", "昨天", "X年X月X日") via `parseChineseDate`
- **Ad detection**: `data-tuiguuang` attribute, `.ec_*` classes, "推广" / "广告" text indicators
- **Official site detection**: "官方" label, `[aria-name="label"]` attribute
- **Rich result types**: `organic`, `baike`, `video`, `image`, `official`, `rich`
- **Related searches**: extracted from `#rs` and pushed as separate records with `resultType=relatedSearch`
- **Sub-links**: Baidu "sublinks" block captured into `siteLinks[]`
- **Title/snippet HTML**: preserves `<em>` emphasis tags from Baidu
- **Emphasis extraction**: collects highlighted terms (query matches) into `emphasis[]`
- **Chinese title handling**: `parseChineseDate` for "X小时前", "X月X日", "昨天", etc.
- **Multi-page support**: `maxPages` parameter (default 5)
- **New input field**: `startUrls` (array) — supply your own Baidu search URLs
- **New input field**: `includeAds` (boolean, default false)
- **New input field**: `includeRelated` (boolean, default true) — push "相关搜索" as records
- **New input field**: `market` (e.g. `zh-CN`, `en-US`)
- **New input field**: `maxPages` (default 5)
- **2 views**: `overview` (key fields) and `fullDetails` (all 41 fields)
- **Favicon auto-resolution** from result domain

#### Changed

- **Breaking** — actor version bumped to 2.0.0
- **Breaking** — schema expanded from 7 fields to 41 fields
- **Breaking** — input fields expanded: `query`, `startUrls`, `maxItems`, `maxPages`, `market`, `includeAds`, `includeRelated`, `proxyConfiguration`
- Improved pagination detection (handles `#page .n`, `a.n`, `.page-n a`)
- Removed unused `errors/` directory and inline `await` after `Actor.init()`
- Cleaner `main.js` — extraction moved to `routes/handlers.js`
- `package.json` no longer runs `crawlee install-playwright-browsers` (browsers pre-installed in base image)

#### Fixed

- Resolves Baidu redirect links (`https://www.baidu.com/link?url=...`) to actual destination URLs via HEAD request
- Strips extra whitespace and empty values from `source` field
- Handles missing optional DOM elements gracefully

### \[1.0.0] - 2025-07-01

#### Added

- Initial release of the actor
- Full README documentation with usage examples and field descriptions
- Input schema with validation for all parameters
- Dataset output schema defining all output fields
- Structured source layout: `src/main.js`, `src/routes/handlers.js`, `src/routes.js`, `src/errors/`

#### Changed

- Dockerfile updated: all `COPY` commands use `--chown=myuser` to prevent EACCES build errors
- Removed `postinstall: npx crawlee install-playwright-browsers` (browsers pre-installed in base image)
- `playwright` kept in `dependencies` (not devDependencies) for correct runtime import resolution

#### Fixed

- Build reliability: resolved `EACCES: permission denied, open '/home/myuser/package-lock.json'` on Apify Cloud (exit code 243)
