# Changelog of DuckDuckGo Search Scraper (`searchapi/duckduckgo-search-scraper`) Actor

- **URL**: https://apify.com/searchapi/duckduckgo-search-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/searchapi/duckduckgo-search-scraper.md

## Changelog

### 3.1.0 - 2026-09-22

- Removed the public debug input and debug-only log-level switch.
- Added required rich dataset fields, a full schema contract, and an audit for populated fields, ordering, timestamps, duplicate IDs, sensitive keys, and schema parity.
- Added a live sample from the current public DuckDuckGo organic search endpoint.
- Verified a fresh run for `artificial intelligence` returned 20 records across two public HTML pages; headed Chrome showed the matching current organic result cards and Search Assist modules.

### 3.0.3 - 2026-08-29

- Replaced the broken Firefox-only crawler with fast public HTML search extraction.
- Added fair multi-query processing, bounded pagination, region/Safe Search/date filters, global and per-query limits, concurrency, timeouts, and bounded retries.
- Added stable query-specific IDs, validated redirect decoding, recursive empty omission, explicit provenance, and a truthful linked 36-field dataset schema.
- Added coherent transport profiles, ordinary cookie continuity, request pacing, authorized proxy support, and explicit rejection of incompatible `GOOGLE_SERP` input.
- Added fail-closed CAPTCHA, content-type, malformed-layout, and unsafe-pagination handling.
- Fixed Apify proxy session IDs to use only the platform's permitted characters.
- Excluded sponsored cards and their opaque DuckDuckGo/Bing tracking URLs from the organic dataset.
- Removed Playwright and dead browser modules; upgraded deterministic dependencies and resolved production audit findings.

All notable changes to the DuckDuckGo Search Scraper are documented here.

### \[2.0.1] - 2026-06-13

#### Verified

- End-to-end run against `lite.duckduckgo.com` for query "artificial intelligence" returned **10/10 records** with all fields populated:
  - All records have real titles, links, snippets (100–400 chars), and decoded display URLs
  - 5/10 records have parsed `date` (mixed ISO datetimes and date-only formats from the Microsoft .NET 7-decimal datetime emitted by the lite API)
  - `domain` correctly extracted from the decoded link (e.g. `en.wikipedia.org`, `britannica.com`)
  - No captcha / no proxy required for lite.duckduckgo.com
- Parser logic in `parsers.js` (4-row table layout, position detection, DDG `/l/?uddg=` redirect decoding, ISO date normalisation) still matches the current DOM.

### \[2.0.0] - 2026-06-12

#### Added

- **Schema expanded** from 32 to **46 fields**.
- **New fields**: `resultType`, `titleHtml`, `url`, `displayUrl`, `host`, `description`, `imageUrl`, `reviewCount`, `dateRaw`, `publishedAtRaw`, `category`, `tags`, `isAd`, `isSponsored`, `query`, `market`.
- **New view**: `fullDetails` — all 46 fields in a single table.

### \[2.0.0] - 2026-06-12 (DDG refactor)

#### Added

- **Identity & metadata**: `position`, `page`, `type`, `resultType`.
- **URLs**: `link`, `displayedUrl`, `domain`, `favicon`.
- **Snippet**: `snippet`, `snippetHtml`, `emphasis`, `sitelinks`, `breadcrumbs`.
- **Media**: `thumbnail`.
- **Ratings & price**: `rating`, `ratingCount`, `price`, `priceNumeric`, `currency`.
- **Dates**: `date`, `publishedAt`.
- **Rich result**: `richResultType`, `additionalInfo`, `cacheLink`.
- **Run context**: `searchUrl`, `searchMetadata`, `searchQuery`, `country`, `language`, `scrapedAt`.
- **Shared validator/normalizer** wired to `_canonical/`.

### \[1.0.0] - 2025-05-20

#### Added

- Initial production release of the DuckDuckGo Search Scraper.
- `PlaywrightCrawler`-based architecture using Firefox (`apify/actor-node-playwright-firefox`).
- Modular project structure.
- `validate-datasets.js` for local dataset output validation.
- `Dockerfile` using `apify/actor-node-playwright-firefox` image.
- `README.md` with full Apify usage guide.
