# Changelog of 1688 Image Search — $0.005/photo, $0.01/run on success (`crawleast/1688-image-search-scraper`) Actor

- **URL**: https://apify.com/crawleast/1688-image-search-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/crawleast/1688-image-search-scraper.md

## Changelog

### 0.3.30 — README only: Related Actors (2026-09-28)

No pipeline, billing, input or output changes; every other file is identical to the 0.3.29 build.

- **Related Actors now lead with the 1688 Product Scraper and the TikTok Shop Scraper** — keyword research and TikTok winner discovery that feed photo sourcing. The Pet Supplies Scraper link stays; the Pipeline line is updated to match.
- **The 1688 → AliExpress AI Product Matcher link is removed** because that Actor is being withdrawn from the Store. (0.3.29, 2026-09-21, was the README-only build that had added it.)

### 0.3.28 — Docs & metadata only (2026-09-16)

No pipeline, billing or output changes; `src/` is unchanged apart from one comment.

- **Buyer-facing copy no longer names competitors.** The value table and the reverse-image-search comparison section now describe capability and billing-model differences instead of naming rival Actors, and the run-success row points readers at each Store page's own live stats rather than quoting a third party's number. One source comment lost its competitor reference.
- **`meta.categories` realigned to `["ECOMMERCE","AGENTS"]`** to match the live Store listing. The local file declared only `ECOMMERCE`, so a future push would have silently dropped the AI-agents category placement.

### 0.3.27 — Task #270: detail-budget "regression" investigated (NO regression) + enrichment observability (2026-09-04)

0.3.26 investigated for alleged detail-budget regression — CONFIRMED NO regression; detail slice correctly tracks opt-in enrichDetails (default false); pin restored to 0.3.26 on 2026-09-04.

The reported fingerprint (`budget … detail 0s`, `detailResults.requested=0`) reproduced ONLY on runs whose input had `enrichDetails=false` — the schema default for the opt-in paid enrichment ($0.003/delivered detail). A decisive A/B (same input, `enrichDetails=true`, one run pinned 0.3.26 vs one 0.3.24) returned byte-identical results (`requested=3, delivered=3, detail 330s, chargedUsd=0.014`). `planBudgetSecs` (`detailSlice = enrichDetails ? … : 0`) is unchanged across 0.3.24→0.3.26; the only 0.3.26 `main.js` change was a try/catch around the image-search phase (P1-1 resilience), which neither returns early nor gates the detail phase. Card-only runs were charged the card price only (`$0.005`, `detail-item=0`) — no overcharge, no silent under-delivery.

This release adds durable guards so the opt-in/opt-out state can never again be mistaken for a regression:

- `src/main.js`: new `SUMMARY.enrichment` field — `"enriched"` vs `"card-only (enrichDetails=false)"` — computed by the exported pure helper `enrichmentLabel({enrichDetails, mode})` (offerIds mode forces `enriched`), mirroring the `enrichEnabled` gate. + unit test.
- `cloud-tests/run-case.mjs`: executor gains a **sequence mode** (ordered run1→run2 with per-step waits) and a minimal **assertion framework** (`summary.<path> === value`, `requested === min(matchCount, maxResultsPerImage)`, `budgetDetailNonZero`, `logContains`), exit non-zero on failure. Plain-input cases remain backward-compatible.
- `cloud-tests/cases/warm_replay_enrich.json`: new permanent regression case — run1 (enrich small order) primes the flywheel write-back, run2 asserts warm-replay gate ACCEPTED + `requested === min(matches, maxResultsPerImage)` + `enrichment="enriched"`.
- `cloud-tests/smoke/staging-enrich.json`: converted to a structured case asserting the budget `detail` slice is non-zero and `enrichment="enriched"`.
- Billing, input schema and output shape unchanged.

### 0.3.20 — Task #236 Flywheel 2.0: browser-channel warm-jar consumption (2026-08-26)

The shared warm-token flywheel now consumes cached cookie jars ONLY through
a real browser channel. Task #234 proved (36 cloud runs, 7/7 clean-jar
replays) that warm jars replay across runs/IPs/sessions exclusively in a
browser — HTTP/impit transports are 100% RGV587-intercepted — and that the
historical failures came from duplicate/stale `_m_h5_tk` entries in the jar.

- New `src/jar-hygiene.js`: `dedupeJar` (most-specific-domain-wins) +
  `_m_h5_tk` re-parse from the deduped jar (+ unit tests).
- New `src/warm-replay.js`: run-level acceptance gate — deduped jar seeded
  into a fresh stealth context, ONE signed in-page mtop canary decides;
  accepted → the identity is reused for the whole run, rejected → silent
  self-mint fallback (pre-flywheel semantics).
- `src/main.js`: one acceptance gate per run; flywheel write-back now
  publishes the FULL deduped cookie jar (was: token fields only).
- `src/crawler.js`: pipeline cache-hit consumes the pre-accepted replay
  browser (or validates the jar itself); all mint write-backs dedupe +
  re-parse before persisting.
- Cleanup: `src/experiments.js`, its test file and the `__experiment`
  input gate are removed (experiment concluded).
- Billing, input schema and output shape unchanged.

### 0.3.11 — Task #234 replay-binding experiment harness (NOT a product release) (2026-08-26)

Experimental build for the hard-method replay experiment matrix.
Production behavior is UNCHANGED: the harness activates ONLY when the
run input carries an explicit `__experiment` object (production inputs
never do); an env kill switch `EXP_ENABLED=0` disables even that.
Without `__experiment` in the input, execution is byte-for-byte
identical to 0.3.9. (Original plan was an env-level ON switch, but the
platform API refuses env-var injection for this token — activation is
therefore input-driven; documented in the Task #234 report.)

- New `src/experiments.js`: E0 positive control, E1 in-run replay,
  E2A/E2B cross-run full-jar replay (isolated KV key `exp_replay_v1`),
  E4 fresh-IP control, E5 clean-marker isolation, ip-echo primitive,
  optional warm-up sequence and HTTP-channel canary.
- `src/main.js`: env-gated dispatch hook before input resolution.
- Production warm-cache key `mtop_token_v1`, pinned builds and
  schedules are untouched.

### 0.3.8 — README trust elements (Task #229) (2026-08-26)

Documentation-only release; no code, input, output or billing changes.

- **Trust at a glance table**: seven verifiable commitments (pay only for
  delivered results, no hidden fees, honest degradation disclosed in
  SUMMARY, subscriber discounts, schema-declared MCP/agentic-payment
  readiness, fast Issue response, live Store-page reliability stats).
- **One-click examples**: direct links to the three public quick-start
  tasks (quick start, full details & landed cost, offer-ID lookup).
- **Review call-to-action**: compliance-safe, non-incentivized review
  request in the Reviews section.
- Verified complete: the existing "Copy to your AI assistant" section
  (run-sync endpoint, `?timeout=900`, 408 polling recovery).

### 0.3.7 — Shared warm-token cache (Task #228) (2026-08-26)

Token-cache plumbing only; no input, output or pipeline-behaviour changes.

- **2c — shared-cache REST read**: when the injected KV env vars
  (`MTOP_KV_API_TOKEN` + `MTOP_KV_STORE_ID`) are present, `get()` reads the
  minter-published warm token from the shared KV store via the Apify REST
  API BEFORE the SDK path. A fresh (<25 min) entry is used as-is
  (`shared-cache: hit`); any failure silently degrades to the existing
  SDK/bootstrap path (`shared-cache: miss`). Works for external users too,
  whose SDK cannot see the owner's named store.
- **2d — flywheel write-back**: after a self-bootstrapped token is cached
  locally, it is also PUT back to the shared store (non-fatal on failure).
- **2e — pool_health unchanged**: circuit-breaker entries stay on the
  run-local/default store; they are never read from or written to the
  shared store.
- **Secrets hygiene**: the KV token value never reaches logs, datasets or
  error messages (`shared-cache: hit/miss` only; `scrubSecret` helper).
- New tests: `tests/shared-cache-228.test.js` (hit / miss / degrade /
  flywheel / pool_health isolation / token-never-leaks).

### 0.3.6 — Pricing v2: run-base-fee deferred billing (2026-08-26)

Billing-model change (platform pricing update + settlement code); no input,
output or pipeline-behaviour changes.

- **New `run-base-fee` event ($0.01)**: charged once at run completion, and
  only when the run delivered at least one match. Runs that deliver no
  matches are not charged. Settled after ALL flushes (`settleRunBaseFee`),
  guarded by try/catch — a charge failure (e.g. the event still in its
  14-day pending window) logs a warning and never fails the run.
- **Pricing update** (effective 2026-09-09, 14-day notice for existing
  paying users): `image-search` tiers flattened to $0.005 (FREE/BRONZE) /
  $0.0048 (SILVER) / $0.0045 (GOLD+) — discount depth capped at 10%;
  `detail-item` unchanged; $0.04 minimum per run retained.
- **README**: pricing section rewritten (all-inclusive messaging, base-fee
  row, 1/10/100-image worked examples), cheapest-style claims removed.

### 0.3.5 — Version-line restoration + default build pinning (2026-08-25)

No functional changes — content is identical to the 0.3.4 line plus the
docs-only pricing update that was shipped as 0.2.12/0.2.13.

- **Version-line restoration**: the 0.2.12/0.2.13 pushes were traced to
  this repository (CLI origin, timestamps match the 0.2.0 snapshot and
  pricing-docs pushes); they were legitimate but retired the `latest`
  tag onto the 0.2 series. Version files restored to the 0.3 line
  (package.json 0.3.5, actor.json 0.3) so `latest` points back at the
  current code line.
- **Default build pinning**: Actor `defaultRunOptions.build` pinned to
  the exact 0.3.5 build number, so Console default runs and Store
  "Try" no longer follow a bare `latest` tag. The daily canary
  schedule intentionally keeps `latest` as a hijack early-warning.

### Pricing update — unified tiered subscription discounts (2026-08-24)

Billing + docs-only change (no code, input or output changes):

- **Unified tiered PPE pricing** (aligned with 1688 Pet Supplies Scraper),
  discounts deepened from −5/−15/−25 to −10/−20/−30:
  - `image-search`: FREE $0.005 (unchanged) / BRONZE $0.0045 / SILVER $0.004 / GOLD+ $0.0035
  - `detail-item`: FREE $0.003 (unchanged) / BRONZE $0.0027 / SILVER $0.0024 / GOLD+ $0.0021
- Pure price reduction → effective immediately under Apify policy (no
  14-day waiting period). Appended as a new `pricingInfos` record;
  price history preserved (append-only). `minimalMaxTotalChargeUsd`
  stays $0.04.
- Unchanged: no start fee, failed/undelivered work never charged.
- README discount line updated to −10/−20/−30; worked example now
  states it is computed at FREE-tier prices.

### 0.2.0 — Pre-publication snapshot (2026-08-24)

功能修复批（外部测试报告 IS-1/2/3/4）与产品化批合流为统一版本快照；staging 验证中，待用户手动门禁后上架。

### 0.3.3 — External test-report fix batch (IS-1/2/3/4, 2026-08-24)

Fixes against the independent 20-case test report (build 0.3.2). Billing
semantics untouched: pay-per-success events unchanged, pricingInfos
append-only.

**IS-1 \[High] titleEn Chinese residue (34/34 dictionary-only):** root
cause was `translateTitle(..., { useApi: false })` in enrichProduct plus a
dead LibreTranslate endpoint. Fix: keyless machine-translation chain
(Google gtx → MyMemory) with 5–6s timeouts and a per-run 3-strike circuit
breaker per endpoint, wired with `useApi: true`. API outputs still
containing CJK are rejected and fall back to the domain dictionary;
`titleCn` stays authoritative. README discloses machine-translation
best-effort semantics. New tests/w2-translation.test.js (7 cases,
fetch-mocked).

**IS-2 \[Med] cold-start canary NETWORK got only 1 of 2 configured
attempts:** canary retry loop now treats NETWORK like RISK_CONTROL
(retry once on a fresh identity). Intermediate retryable attempts are
marked diagnostic (no dataset row); only the FINAL outcome is delivered,
preserving the one-item-per-input-image contract.

**IS-3 \[Low] card-only runs wrote 2 rows per image (null vs 0
detailDelivered):** the final settlement flush re-pushed every imageResult
after the per-image flush. Card-only runs now skip the final re-flush and
settle `detailDelivered: 0` at per-image flush time — exactly one row per
input image. Enrich-mode snapshot semantics unchanged.

**IS-4 \[Low] post-cap in-flight details delivered free:** intentional
graceful-stop design; now documented in README Limits ("Budget cap stops
gracefully, never mid-delivery").

**Cloud verification (build 0.3.3):** card-only run yIyhIgzReUSCrD4lT —
1 dataset row for 1 input image (IS-3 closed), detailDelivered=0,
charged {image-search:1} = $0.005, ledger matches. Enrich retry run
SvE6e7QuLD2gqAm02 — 3/3 details delivered, charged {image-search:1,
detail-item:3} = $0.014, translationMethod=api 3/3 with zero Chinese
residue in titleEn (IS-1 closed). First enrich attempt B1iehpMi3l34Kxyh7
hit transient RGV587 risk control on the detail layer (0/3, correctly
UNCHARGED) and self-healed on retry — same pattern as the 0.3.1 staging
evidence.

### 0.3.2 — Ops-audit closeout batch (Tina P1×5 + P2, 2026-08-24)

**P1-1 (self-provable competitor claims):** README ✅/❌ table — dev00 raw
run counts removed (not reproducible), replaced with "~80–85% run success
rate per Apify Store stats (as of Aug 2026)"; devcake per-row price kept
with an "as of Aug 2026" date anchor.

**P1-2 (G-09 zero-result billing evidence):** staging build 0.3.2, both
runs SUCCEEDED with ZERO actor charges:

- 404 image run Raac0FXlP4EWJlgNd (31s): dataset status
  IMAGE_FETCH_FAILED, chargeable=false; platform chargedEventCounts
  {image-search:0, detail-item:0}; SUMMARY chargedEvents={}, chargedUsd=0.
- NO_INPUT empty-input run SlxbcY0LSiETJgzEB (31s): OUTPUT errorCode
  NO_INPUT ("nothing to do, nothing charged"); chargedEventCounts
  {image-search:0, detail-item:0}; SUMMARY chargedEvents={}, chargedUsd=0.
  Both runs bill only platform compute (~$0.002). No apify-actor-start
  synthesis event exists in pricingInfos — nothing to disable.

**P1-3 (minimal alias mapping restored, contract G-04):**
input-validation R4 now MAPS `imageUrl`→imageUrls (scalar wrapped in
array) and `images`→imageUrls; `imageFile` routed by value shape
(data:URI / pure base64 → imagesBase64, otherwise → imageUrls); `query`
still DROPPED with a warning (this Actor searches by image, not keywords).
NO_INPUT OUTPUT message now reports "Detected unsupported fields: …".
Plan deviation 留痕: W2 had dropped ALL four aliases; the ops audit ruled
the minimal mapping is the G-04 original intent. Tests rewritten + new
P1-4 cap tests (49/49 green). README migration note updated accordingly.

**P1-4 (schema/code/README boundary parity):** code clamps tightened to
maxImages ≤20 (was 400) and maxResultsPerImage ≤20 (was 60);
DEFAULT_MAX_IMAGES 20 → 10, matching input schema (default 10, cap 20)
and the README input table. README "Up to 20 images × 20 matches" stays
true.

**P1-5 (snapshot semantics promoted):** dataset schema description +
Overview view description now state "runs write progressive snapshots;
consumers take the last imageResult block"; README Sample output carries
the same note up front.

**P2-1:** PLATINUM/DIAMOND tiers verified via API — both events carry
explicit tier keys at the −25% price (0.00375 / 0.00225); README wording
("GOLD/PLATINUM/DIAMOND −25%") stays.
**P2-4:** Copy-to-assistant section adds `?timeout=900` / client-timeout
guidance for large enriched orders (gateway 408 recovery via run polling).
**P2-5:** dedup/benchmark numbers unified on the recorded source of truth
(20-image benchmark run MrWX8o2ZcMFXC2ZQ0: 371 unique targets,
371/371 delivered & billed; compensation retry 50 offered / 50 delivered).

### 0.3.0 — W3 productization batch (2026-08-24)

**Schemas (all four `apify validate-schema` classes green):**

- Input schema v2: 8 buyer fields (imageUrls / imagesBase64 / offerIds /
  maxImages \[default 10, cap 20] / maxResultsPerImage \[default=cap 20] /
  enrichDetails / targetMarket \[US,EU,UK,AU,global] / maxTotalChargeUsd
  \[default 2, min 0.04]); NO required fields; diagnostic switches
  (useMintedToken, maxRunTimeSecs) removed from the Store form (runtime
  still honours them); wide aliases stay runtime-mapped, documented.
- Dataset schema (draft-07, exact `$schema`): one row per query image
  (`imageResult`) or one summary row (`offerIdsResult`); search-card fields
  - enriched-detail fields incl. two supplier shapes (card vs detail-page),
    riskFlags as {flag,message} objects; views.overview buyer table. NO
    always-null fields: reviewCount/rating verified absent from W2 data and
    omitted entirely; repurchaseRate kept nullable (can land when a genuine
    product-level source exists — 846/846 null in W2, disclosed).
- Reverse validation gate (tests/w3-schema.test.js, AJV allErrors):
  233 real dataset items / 4514 listing cards / 846 enriched details from
  ALL W2 smoke captures pass; corrupted items are rejected (batch-reject
  semantics).
- KV schema: SUMMARY + OUTPUT collections, loose (contentTypes only).
- Actor output schema: results → dataset items link, runSummary → KV
  SUMMARY record link.

**Pricing:** tiered subscription discounts added (append-only pricingInfos,
old entry preserved): image-search FREE $0.005 / BRONZE $0.00475 /
SILVER $0.00425 / GOLD+ $0.00375; detail-item FREE $0.003 / BRONZE
$0.00285 / SILVER $0.00255 / GOLD+ $0.00225; minimalMaxTotalChargeUsd
0.04; still no apify-actor-start event.

**Listing:** README twelve-section marketplace rewrite (hook with hard
numbers, ✅/❌ comparison naming dev00's 85.3% success rate and
devcake-style per-row billing, real sample output, worked-example pricing,
honest limits incl. snapshot semantics / dedup billing / compensation
retry, FAQ, "Copy to your AI assistant", Related cross-link, review CTA).
Store description 265 chars; SEO description 150 chars; SEO name 43 chars.

**Prefill / example run:** golden image
`https://cbu01.alicdn.com/img/ibank/15418323391_1847764777.jpg` validated
on cloud build 0.2.11 (run A9JTov8FNAu9WyPu5: SUCCEEDED in 11s, 1 image ×
5 cards, billed 1 image-search) and set as input prefill + Actor-level
exampleRunInput (drives Apify's daily default-input health check).

**Staging:** build 0.3.0 pushed; 48h staging window opened with a default
input run + a small enriched order; no Store publication until W4 gate.

### 0.2.11 — SM-04 compensation retry (2026-08-24, leader-approved, guardrails)

- ONE serial compensation round after the parallel pool pass: retries only the
  pool pass's failed offerIds (no parallel risk-control stacking), gated on
  deadline slack = conservative estimate (180s/group) + 60s safety margin.
  When the gate closes, the run stays honest-partial and says so
  (`compensationRetry.reason = 'deadline_slack_insufficient'`).
- Billing semantics unchanged: the compensation round charges only what it
  actually delivers; charged = delivered ledger parity is preserved.
- New SUMMARY field: `compensationRetry { triggered, reason, offered,
  delivered, failed }`.
- Pure gate exported as `planCompensationRetry` + unit-tested (42/42 green).
- README W3 backlog: dataset snapshot semantics + cross-image dedup calibre
  (billed detail-item counts unique offerIds; card-level enriched count may
  exceed it when one offer appears under several query images).

### 0.2.10 — SM-04 retest batch (2026-08-24, leader decision A)

**Frozen-plan revision (留痕):** the big-order deadline budget is re-planned.

- `DETAIL_SLICE_BIG_MS` raised 780s → 2200s; new big-order shape is
  bootstrap 48s + image slice 90–300s + detail slice ≤2200s + finish 48s,
  kept inside the new 2400s actor timeout. Small orders (≤100 detail targets)
  stay on the 330s standard slice — unchanged.
- `.actor/actor.json` defaultRunOptions timeoutSecs raised 1200 → 2400
  (min/memory unchanged; the cap is a bounded protection, small runs still
  finish in ~45–550s and are billed on actual usage).
- Detail identity groups now run on a BOUNDED PARALLEL POOL of ≤3
  (`DETAIL_PARALLEL_IDENTITIES`, mirroring the image-phase precedent) instead
  of strictly serial — previous serial throughput (~55 offers/100s) made a
  378-offer order impossible inside any platform timeout. Snapshot flush is
  serialised through a promise-chain lock; final settlement recomputes
  failed = requested − delivered.
- `effectiveDetailSliceMs` (0.2.9) keeps the slice honest when the platform
  timeout leaves less room than the plan.
- `cloud-tests/run-case.mjs`: run timeout now env-driven (`RUN_TIMEOUT_SECS`,
  default 1200) with poll wall = timeout + 5 min.

### 0.2.9 — SM-04 fix batch (2026-08-23)

- enrich-mode flush leak fixed: raw image cards no longer settle to the
  dataset before enrichment (persistState used to splice the queue empty;
  details were charged but never delivered). Progressive flushes now re-queue
  a full snapshot each time; consumers take the LAST imageResult block
  (append-only dataset). `summary.imageFlushCount` records snapshot count.
- Platform-timeout-aware detail slice: the detail slice shrinks to the real
  remaining window (ACTOR_TIMEOUT_AT) and logs a WARNING when clipped.

### 0.1.0 — Scaffold (2026-08-23)

- Repository initialized: `package.json` (deps mirrored from Actor 1), `.actor/actor.json`
  (timeoutSecs 1200, defaultRunOptions 1024 MB, categories ECOMMERCE), `.gitignore`.
- 13 modules reused verbatim from `china-pet-supplies-scraper/src/`:
  token-cache, utils (createDeadline timeout/cost protection inherited), mtop-search,
  crawler (full copy, untrimmed — keyword-search segment isolated in W2), browser-detail,
  browser-bootstrap, translator, dropship-score, landed-cost, compliance, sanitize,
  data-safety, i18n-mappings.
- `translator.js`: `KEYWORD_MAP` import replaced with an empty-map placeholder
  (`keyword-mapper.js` is intentionally not carried over).
- `tests/scaffold-imports.test.js`: asserts all reused modules import and key exports
  exist (createDeadline, createResultFlusher, sanitizeProduct, …).
- `cloud-tests/` harness adapted from Actor 1 (BUILD_NUMBER pin logic kept, token via env).
- Known scaffold gaps (by design, finalized in W3): storages/output schema set
  (invalid placeholder schema files would break `apify validate-schema`), marketplace README.
