# Changelog of Baidu Images Scraper (`searchapi/baidu-images-scraper`) Actor

- **URL**: https://apify.com/searchapi/baidu-images-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/searchapi/baidu-images-scraper.md

## Changelog

### \[3.2.1] - 2026-09-05

- Fixed Baidu Images runs on hosts whose system DNS cannot resolve `image.baidu.com` by using a bounded public-DNS browser mapping only when normal resolution fails.
- Replaced Crawlee's local CONNECT proxy path with a direct Playwright context so the DNS fallback is effective while shared cookies, retries, concurrency, optional authorized proxies, API extraction, and DOM fallback remain available.
- Fixed Apify proxy session naming, added filter/developer QA inputs, and verified real multi-page, multi-query, proxy, filter, no-results, and invalid-input runs.

### \[3.2.0] - 2026-09-05

- Added validated width, height, file-format, source-domain, and output-sort filters.
- Added developer controls for page delay, concurrency, retries, timeouts, media blocking, and diagnostics.
- Added strict output, key-value-store, and run-validation contracts for API metadata, pagination, fallback extraction, and multi-query summaries.
- Added safe malformed-response handling, fair multi-query interleaving, and a rendered DOM fallback when the structured response is unavailable.

### \[3.1.0] - 2026-08-29

- Updated structured extraction for the current nested `data.images` response and verified `gsm` pagination; 70-result three-page tests now return 70 distinct records.
- Moved API requests into the loaded browser session, with status/content-type validation and DOM fallback.
- Added single and fair multi-query modes, per-query limits, query concurrency, stable SHA-256 IDs, global positions, and explicit no-results completion.
- Added current Baidu watermark, AI-edit, commodity, source-type, clarity, copyright, and image-format fields.
- Removed fabricated Google favicons and transient/raw request diagnostics from output and logs.
- Aligned Apify, Crawlee, Playwright, and the Chrome Docker image; production dependencies now audit with zero vulnerabilities.
- Added Residential/custom proxy support with fail-closed verification handling; `GOOGLE_SERP` is explicitly rejected as incompatible.

### \[3.0.0] - 2026-07-20

- Added validated API-first pagination with bounded retry/backoff and DOM fallback.
- Removed unsupported category/color and result-URL inputs.
- Added stable ID plus image-URL deduplication and strict input validation.
- Expanded and corrected the dataset schema to 46 normalized fields.
- Added numeric dimensions, source validation, richer Baidu metadata, tests, and nonzero terminal failures.

### \[2.0.0] - 2026-06-12

#### Added

- **51 canonical output fields** grouped by category, produced via the shared `normalizeImageRecord` in `_canonical/normalize.js`:
  - **Identity (4)**: `type`, `resultType`, `position`, `page`
  - **Title (2)**: `title`, `altText`
  - **URLs (8)**: `imageUrl`, `url`, `thumbnail`, `thumbnailUrl`, `previewImage`, `originalUrl`, `hostPageUrl`, `sourceUrl`
  - **Domain (5)**: `hostPageDomain`, `sourceName`, `sourceDomain`, `domain`, `favicon`
  - **Dimensions (4)**: `width`, `height`, `aspectRatio`, `sizePixels`
  - **File metadata (4)**: `fileFormat`, `mimeType`, `fileSize`, `fileSizeFormatted`
  - **Visual (3)**: `isAnimated`, `color`, `colorHex`
  - **Authorship (4)**: `creator`, `creatorUrl`, `creatorAvatar`, `copyright`
  - **License / source (2)**: `license`, `source`
  - **Page (2)**: `pageTitle`, `pageUrl`
  - **Categorisation (2)**: `caption`, `tags`
  - **Related (1)**: `similarImages[]`
  - **Engine identity (2)**: `engineId`, `engineData`
  - **Search context (8)**: `searchQuery`, `query`, `searchUrl`, `market`, `country`, `language`, `scrapedAt`, `searchMetadata`
- **Shared normalizer**: uses `normalizeImageRecord` from `_canonical/normalize.js`
- **File size parsing**: `parseFileSize()` converts "2.5 MB", "800 KB", "1.2 GB" to bytes
- **File size formatting**: `formatFileSize()` converts bytes to human-readable "2.5 MB", "800 KB"
- **MIME type detection**: `formatToMime()` converts "jpeg", "png", "webp", "gif" to "image/jpeg", "image/png", etc.
- **Aspect ratio computation**: `computeAspectRatio()` returns width / height rounded to 4 decimals
- **Size pixels computation**: `computeSizePixels()` returns width × height
- **`isAnimated` detection**: based on file extension (gif, webp, apng)
- **Color extraction**: parses inline `style="background-color: #..."` on the image element
- **Creator / license extraction**: looks for `[class*="author"]`, `[class*="creator"]`, `[class*="license"]` elements
- **Page title extraction**: looks for `[class*="title"]`, `[class*="caption"]`, `[class*="alt"]` elements
- **Multi-image grid scrolling**: loops up to 40 times, scrolling the grid until `maxItems` is reached or no new images appear (6 stall rounds)
- **2 views**: `overview` (key fields) and `fullDetails` (all 51 fields)
- **New input fields**: `startUrls`, `maxPages`, `category`, `color`, `market`

#### Changed

- **Breaking** — actor version bumped to 2.0.0
- **Breaking** — schema expanded from ~10 fields to 51 fields
- Cleaner `main.js` — extraction moved to `routes/handlers.js`
- `package.json` no longer runs `crawlee install-playwright-browsers` postinstall
- Pagination respects `maxPages` cap

### \[1.0.0] - 2025-07-01

#### Added

- Initial release of the actor
- Basic image extraction
