# Changelog of Propriétés Privées · Full profiles infos (emails, phones) (`corent1robert/proprietes-privees-negociateur`) Actor

- **URL**: https://apify.com/corent1robert/proprietes-privees-negociateur/changelog.md
- **Full Actor documentation**: https://apify.com/corent1robert/proprietes-privees-negociateur.md

## Changelog

All notable changes to this project will be documented in this file.

### \[1.12.6] - 2026-08-31

#### Fixed

- Store health check / Try Apify: LinkedIn enrichment no longer crashes the run (`log.warning` was called unbound → empty dataset → `UNDER_MAINTENANCE`). Enrichment is skipped on the 5-advisor prefill and wrapped in try/catch so a search/API error still saves advisors.
- Prefill `maxResults` lowered 50 → **5** so the default input finishes well under Apify's 5-minute Store test.

### \[1.11] - 2026-08-20

#### Fixed

- Apify runs no longer finish with **0 advisors**: **residential FR proxy** (datacenter hangs on this host from the platform), a real browser User-Agent (the old `ProprietesPrivees-Scraper/1.0` UA was blocked), and a hard fail when every profile fetch fails.

### \[1.10] - 2026-03-10

#### Changed

- **Phase 2 speedup** — Default concurrency 20→40, pool connections = concurrency×2 (was bottleneck: 20 conn for 40 req/batch). ETA full scrape ~40–50 min (was ~2 h).

#### Added

- **Generic email filter** — commerces@, contact@, info@, commercial@, etc. excluded from output (only actionable professional emails kept).
- **LinkedIn URL cleanup** — Fragments (#...) stripped for cleaner URLs.

#### Changed

- Email extraction (HTML + API) now filters generic addresses consistently.
- README FAQ updated for empty email clarification.

### \[1.9] - 2026-03-10

#### Added

- **Listing average price** — `listingValueAverage` = total/count (€ per listing).
- **Facebook URL** — `facebookUrl` for advisor's personal Facebook profile (excludes proprietes.privees brand page).
- **Business URL** — `urlBusiness` for "Entreprises et Commerces" section (business.proprietes-privees.com/conseillers/…).
- **Presentation excerpt** — `presentationExcerpt` first ~200 chars of advisor's zone presentation text.

### \[1.8] - 2026-03-10

#### Added

- **Phase 2 progress** — Batch logs now show % complete and ETA (e.g. `Batch 5/157 — 100/3129 (3%), ETA ~18min`).
- **Listings data** — `listingsCount` and `listingValueTotal` extracted from profile page (like Sextant).
- **Phone split** — `phonePro` (01–05) and `phonePerso` (06–07) when detectable.
- **Shared phone filter** — 0932508467 (PP hotline) excluded from `telephones`.
- **Error page detection** — Bad Gateway / 502 / 503 pages treated as failed, not pushed to dataset.

### \[1.7] - 2026-03-10

#### Added

- **LinkedIn enrichment** — Same system as sextant (OpenAI + Autom.dev). Runs when `OPENAI_API_KEY` and `AUTOM_DEV_API_KEY` are set as environment variables. No UI — keys in Actor settings. Output: `linkedinUrl`, `linkedInCorrelationScore`.

### \[1.6] - 2026-03-10

#### Added

- **Unit tests** — `npm test` runs parse tests + benchmark.
- **Parse speedup** — Extract `<main>` before Cheerio on detail pages for ~97% faster parsing (Phase 2).

#### Changed

- actor.json `readme` references `./README.md` (like sextant).

### \[1.5] - 2026-03-09

#### Changed

- **Phase 2 reliability** — Retries increased to 3 attempts (was 2) with exponential backoff (800ms → 1.2s → 1.8s). HTTP 429 and 5xx trigger retries.
- **Default concurrency** — Phase 2 concurrency reduced from 30 to 20 to limit rate limiting.
- Backoff also applied to HTTP 429/5xx responses.

### \[1.4] - 2026-03-09

#### Added

- **Benchmark script** — `npm run benchmark` / `npm run benchmark:quick` to test listConcurrency × concurrency.
- BENCHMARK Phase1/Phase2 timing logs for performance analysis.

#### Changed

- `listConcurrency` max raised from 20 to 50 for benchmarking.
- Fix `log.warning` fallback when running locally (console.warn).

### \[1.3] - 2026-03-09

#### Changed

- **Simplified input** — Removed maxPages, maxResults, concurrency, fetchTimeout from UI. List mode scrapes everything by default.
- Interface: only Mode + startUrls (URLs mode).

### \[1.2] - 2026-03-09

#### Changed

- **Phase 1 now parallel** — 15 listing pages fetched simultaneously (was sequential). ~5–10x faster for full scrape.
- Strip `<script>` tags before Cheerio parsing for faster DOM load.
- New internal `listConcurrency` (15) for Phase 1 batches.

### \[1.0.1] - 2026-03-09

#### Added

- **URLs mode** — Enrich specific advisor profile URLs (like KW, Optimhome, Sextant)
- `startUrls` input with requestListSources editor
- `maxResults` parameter (alias for maxNegociateurs, backward compatible)

#### Changed

- **mode** select: List | URLs (Step 1 UX like version-1.0)
- Input schema: sectionCaption, sectionDescription, additionalProperties
- README: Two Modes section, Use Cases table

### \[1.0.0] - 2026-03-09

#### Added

- Initial release
- Scrape Propriétés Privées advisors directory
- Phase 1: Collect URLs from paginated listing
- Phase 2: Extract detail pages (name, title, zone, phone, email, URLs, photo)
- Configurable maxPages, maxNegociateurs, concurrency, fetchTimeout
- English logs, README, and documentation
