# Changelog of Bing Images Scraper (`searchapi/bing-images-scraper`) Actor

- **URL**: https://apify.com/searchapi/bing-images-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/searchapi/bing-images-scraper.md

## Changelog

### \[3.0.1] - 2026-08-29

#### Added

- Added `single` and `batch` input modes with deduplicated query lists, fair per-query quotas, and bounded parallel query processing.
- Added stable query-aware `id` values and `queryPosition` to the documented dataset contract.
- Added a configurable results wait and explicit query-count progress logs.

#### Changed

- Pinned Apify, Crawlee, and Playwright versions and aligned cloud installs with `npm ci`.
- Added patched transitive dependency overrides and the required Actor generation metadata.

### \[3.0.0] - 2026-07-20

#### Changed

- Replaced the URL-driven and hidden browser modes with a strict search-and-filter input contract.
- Added SafeSearch, market, scroll, retry, timeout, proxy, and debug controls.
- Upgraded to Apify 3.7.2, Crawlee 3.17.0, Playwright 1.61.1, and the matching Apify Docker image.
- Removed `playwright-extra` and the Puppeteer stealth plugin; Crawlee now provides session-bound consistent fingerprints.
- Normalized the dataset schema to the 43 fields actually emitted by current Bing result cards.

#### Fixed

- Fixed `proxyConfiguration` being silently discarded by input validation.
- Added response status/content-type validation and explicit challenge handling.
- Fixed pagination/scroll limits so only valid source-backed records count toward `maxItems`.
- Added available `meta.desc` descriptions, corrected JPEG MIME types, and removed fabricated favicon/description fallbacks.
- Filtered cards without valid image, thumbnail, title, or host-page URLs.
- Added global stable deduplication and nonzero failure for selector or extraction changes.

#### Verified

- Default input: 15 records with 43/43 schema parity, no duplicates, missing required fields, invalid URLs, or block records.
- Minimal input: exactly 1 record.
- Filtered input: exactly 5 large, strict-SafeSearch records.
- Multiple-scroll input: exactly 60 unique records.
- `npm test`, dataset validation, input/dataset schema validation, and dependency audit pass.

### \[2.0.1] - 2026-06-13

#### Fixed

- **Critical: actor crashed with "Execution context was destroyed"** — caused by overly-broad consent banner selectors (`button[class*="accept"]` and `button[id*="accept"]`) that matched Bing's own UI controls and triggered unintended navigations. Replaced with a strict allow-list of known GDPR banner IDs/classes (`#bnp_btn_accept`, `.bnp_btn_accept`, `#consent-banner button`, `#onetrust-accept-btn-handler`, `.ot-pc-refuse-all-handler`).
- **Critical: `page.evaluate` silently returned 0 results** — the extraction code called `detectFileFormat`, `isAnimatedByUrl`, and `isAnimatedByFormat` (imported from `../extractors/index.js`) **inside `page.evaluate`**, where Node-side imports are not available. The `try/catch` in the loop silently swallowed the `ReferenceError`, so every card was skipped. The helpers are now inlined directly in the browser-side code.
- **Critical: extracted 0 records even though `.iusc[m]` cards were present** — same root cause as above; once the helpers were inlined, 15/15 records were extracted with all key fields populated (title, imageUrl, thumbnailUrl, hostPageUrl, sourceName, width, height, fileFormat, imageId, contentId).
- **"See more" button clicker was clicking Bing nav buttons** — the selector `.cbtn` matched a generic Bing UI class with `count: 1` in our DOM, causing a navigation that destroyed the page context mid-extraction. Replaced with a narrow allow-list (`input[value="See more images"]`, `a.see_more`, `.btn_seemore`, `.b_seemore`, `button.b_sbBtn[aria-label*="more"]`).
- Wrapped the main extraction `page.evaluate` in a `try/catch` so a transient context-destroyed error no longer aborts the run — the actor simply waits 2s and tries the next step.
- Bumped `waitForFunction` threshold from `> 0` to `>= 20` so we don't start extracting before Bing has actually rendered the grid.
- `preNavigationHooks` now uses `waitUntil: 'domcontentloaded'` (was `'load'`) so we don't block on slow sub-resources.

#### Verified

- 15/15 records extracted for `query: "artificial intelligence", maxItems: 15, market: en-US`
- 100% fill rate on: `position`, `title`, `imageUrl`, `thumbnailUrl`, `hostPageUrl`, `sourceUrl`, `sourceName`, `hostPageDomain`, `width`, `height`, `fileFormat`, `isAnimated`, `imageId`, `contentId`, `searchQuery`, `searchUrl`, `resultType`
- `description` is null across all 15 — Bing's `m` attribute does not include `desc2` for this query (source limitation, not a bug)

### \[2.0.0] - 2026-06-12

#### Added

- New dataset fields: `type`, `resultType`, `page`, `searchQuery`, `searchUrl`, `country`, `language`, `scrapedAt`, `altText`, `description`, `hostPageUrl`, `sourceName`, `contentId`, `thumbnailId`, `cacheKey`, `md5`, `aspectRatio`, `fileFormat`, `fileSize`, `isAnimated`, `isAdult`, `color`, `creator`, `creatorUrl`, `license`, `favicon`, `additionalInfo`, `searchMetadata`
- `width` and `height` are now properly extracted from the detail-view URL parameters (`expw`/`exph`) with `.img_info` text as fallback
- `fileFormat` is detected from the image URL extension (jpg, png, webp, gif, etc.)
- `isAnimated` is set when the file format is animated (gif, webp, apng, avif) or the URL contains animated markers
- `hostPageUrl` is a richer URL field (renamed from `sourceUrl` for clarity, but `sourceUrl` is kept as a backwards-compatible alias)
- `searchMetadata` object containing engine, market, country, language, page, query, resultType, scrapedAt
- New input fields: `mode` (`search` | `urls`), `startUrls`, `headless`, `market`
- New view "Full Details" showing every field
- New input mode `urls` to scrape arbitrary Bing Images URLs via `startUrls`
- Uses shared `_canonical/normalize.js` for cross-actor consistency

#### Changed

- **Breaking** — `position` is now a global counter across pages (not per-page)
- `query` is kept as a backward-compat alias of `searchQuery`
- `sourceUrl` is kept as a backward-compat alias of `hostPageUrl`
- `width`/`height` are now integers (px); previously they were not always extracted

#### Fixed

- Better deduplication using the Bing `mid` (image ID) as the primary key

### \[1.0.0] - 2025-07-01

#### Added

- Initial release of the actor
- Full README documentation with usage examples and field descriptions
- Input schema with validation for all parameters
- Dataset output schema defining all output fields
- Structured source layout: `src/main.js`, `src/routes/handlers.js`, `src/routes.js`, `src/errors/`

#### Changed

- Dockerfile updated: all `COPY` commands use `--chown=myuser` to prevent EACCES build errors
- Removed `postinstall: npx crawlee install-playwright-browsers` (browsers pre-installed in base image)
- `playwright` kept in `dependencies` (not devDependencies) for correct runtime import resolution

#### Fixed

- Build reliability: resolved `EACCES: permission denied, open '/home/myuser/package-lock.json'` on Apify Cloud (exit code 243)
