# Changelog of Panorama Firm (PL) Scraper — Polish business leads (`alwaysprimedev/panoramafirm-scraper`) Actor

- **URL**: https://apify.com/alwaysprimedev/panoramafirm-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/alwaysprimedev/panoramafirm-scraper.md

## Changelog

All notable changes to this actor will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this actor adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

### \[Unreleased]

#### Changed

- **Output schema:** `categories`, `phones`, `emails`, `socialLinks` are now comma-separated strings (or `null` when empty) instead of arrays of strings. Most listings have a single value for each — the array form rendered as "1 item" / "0 items" placeholders in the Apify dataset preview, hiding the actual content. Multi-value entries are preserved as `"a, b, c"`. To regain individual items, callers can `split(", ")`.
- `reviews` now returns `null` instead of `[]` when no reviews are present, for consistent "missing data" semantics across fields.

#### Fixed

- Keyword search now uses `?what=...&where=...` query parameters instead of path-style `/szukaj/keyword,city`. The path form was silently rewritten by the site to a location-only canonical (e.g. `searchTerms: ["nike"]` returned random Pomerania businesses instead of Nike-named ones).
- New `searchLocation` input field separates the city/region from the keyword. Putting both into one term (e.g. `"nike Warszawa"`) yields 6× fewer results than splitting them.

#### Build

- Switched `.actor/Dockerfile` to multi-stage build to avoid `tsconfig.tsbuildinfo` cache silently producing an empty `dist/` on Apify cloud.
- Added `.dockerignore` excluding `node_modules`, `dist`, `tsconfig.tsbuildinfo`, `tests`, `tools`, `tmp`, `.env`.

#### Pending before first Store publish

- Replace `.actor/icon.svg` placeholder with a finalized 256×256 PNG.
- Add screenshots (input form, dataset preview) to README and `.actor/screenshots/`.

### \[0.1.0] — 2026-05-01

#### Added

- Initial scraper for panoramafirm.pl built on Crawlee (`CheerioCrawler`).
- Extracts: name, categories, description, full address (split into street / postalCode / city / voivodeship), phones (normalized to E.164), emails, website, NIP, rating, review count, individual reviews (author, rating, body, ISO date), opening hours, geo coordinates, social links.
- Two start modes: `searchTerms` (free-text queries) and `startUrls` (search, category or detail URLs).
- Pagination follows `<link rel="next">`.
- NIP-based deduplication (`deduplicateByNip`, default `true`) — same business under multiple categories collapses to a single record.
- 404 / 410 detail pages are logged and skipped without failing the run.
- Anti-bot challenge page (`Potwierdź swoją tożsamość`) is detected; the offending session is retired and the request is retried via a fresh session.
- Crawlee session pool with cookie persistence and per-session error scoring.
- Empty / inactive listings are written with `null` fields rather than skipped, so callers can decide.
