# Changelog of LinkedIn Jobs Scraper — Hiring Leads with Emails | $2.99/1k (`glasswing/linkedin-leads`) Actor

- **URL**: https://apify.com/glasswing/linkedin-leads/changelog.md
- **Full Actor documentation**: https://apify.com/glasswing/linkedin-leads.md

## Changelog

### 2026-09-28

- **Fixed: company websites that redirect are reported where they land.** LinkedIn often
  lists an old or alias domain; the row now carries the site the company actually serves
  (`osfhealthcare.com` -> `https://www.osfhealthcare.org/`). A redirect to a non-business
  host (a login wall, a social network) keeps the listed site.

### 2026-09-27

Fixes from a 40-record audit of live output.

- **Fixed: wrong company websites for well-known brands.** A guessed domain (e.g.
  `whatnot.io`, `audible.io`) could be used even when it was really an unrelated site that
  merely used the brand's name as an ordinary English word ("whatnot", "audible" are both
  dictionary words). A guessed domain is now only kept once the fetched page's own
  `<title>`/`og:site_name` actually names the brand — never from a raw body-text match —
  and a domain-broker/parked "for sale" page is rejected outright even when its title
  happens to contain the brand (because the brand name is also the domain name). `.com`
  is always tried before `.io`/`.co`/`.net`. When nothing verifies, the website is left
  empty instead of filled with a wrong guess.
- **Fixed: domain guesses built from the embellished LinkedIn display name.** "Honeywell
  Technologies" produced `honeywelltechnologies.com` instead of the real `honeywell.com`.
  Guesses are now built from the LinkedIn company page's URL **slug** (`honeywell`) when
  one is available, never the display name, and the name/slug-corroboration check no
  longer accepts an arbitrary business-descriptor suffix (only a known legal suffix like
  "inc"/"llc" is still allowed onto a shorter form).
- **Fixed: a second candidate domain could leak contacts into the wrong company.** Once a
  candidate site was accepted, the lookup previously kept trying further TLD guesses and
  could add emails/phones scraped from an unrelated second domain into the same lead.
  The lookup now stops at the first verified (or known) site.
- **Fixed: `summary` glued together ATS boilerplate with no separators**, e.g. `"Date
  Posted:2026-08-25Country:United States of AmericaLocation:US-TX-HOUSTON-011 ~ 11
  Greenway Plz..."`. Recognized label/value boilerplate (Date Posted, Country, Location,
  Job ID, Category, Employment Type, etc.) now gets `": "` and `" · "` separators inserted
  so the field reads as `"Date Posted: 2026-08-25 · Country: United States of America ·
  Location: ..."`. This only ever inserts characters — a description that already read
  fine is unchanged.
- **Fixed: `logo` was always `null`.** LinkedIn lazy-loads this image — the real URL is in
  the `data-delayed-url` attribute, not `src`, which was empty for every record. The
  extractor now reads `data-delayed-url` first (falling back to `src`) and never uses the
  generic `data-ghost-url` placeholder asset as a logo.
- **README:** documented that generic email/phone fill is highest for small and mid-sized
  companies (large enterprises rarely publish either publicly), and documented the new
  website-verification behavior.
- Added `tests/domain-resolution.test.mjs`, `tests/summary-cleanup.test.mjs`, and
  `tests/logo-extraction.test.mjs` (fixture/mocked, no live network), wired into
  `npm test` alongside the existing `fixtures/bench.mjs`.
