# Changelog of LinkedIn Decision-Maker Finder | No Cookies, No Account (`tqm/linkedin-decisionmaker-finder`) Actor

- **URL**: https://apify.com/tqm/linkedin-decisionmaker-finder/changelog.md
- **Full Actor documentation**: https://apify.com/tqm/linkedin-decisionmaker-finder.md

## Changelog

### 2026-09-26: the zero-row outage, two faults, not one

Every run was returning zero rows. There turned out to be **two independent faults**, and they have
to be separated because only one of them is about the proxy.

#### 1. Brave refuses `site:linkedin.com/in`

Not rate limiting, not IP reputation: the engine blocks that one query shape. Measured from an Apify
datacenter IP, an Apify RESIDENTIAL IP and an unproxied home connection, all three identical:

| Query | Result |
|---|---|
| `site:linkedin.com/in "Head of Engineering at Trainline"` | **HTTP 429** |
| `site:linkedin.com/in` (bare) | **HTTP 429** |
| `site:linkedin.com` (no `/in`) | HTTP 200 |
| `site:bbc.co.uk trains` | HTTP 200 |
| `"Head of Engineering at Trainline"` (no operator) | HTTP 200 |

The 429 body is Brave's "your request has been flagged as being suspicious" captcha page. The
operator works and other domains work, so no proxy tier can buy its way out of this one.

**The operator is gone.** The quoted headline phrase was already doing nearly all the constraining,
and the parser only ever reads a `linkedin.com/in` href out of a result block, so a non-LinkedIn
result is skipped rather than mis-parsed. It costs some yield, honestly recorded:

| | rows | high confidence |
|---|---:|---:|
| Trainline | 16 | 10 |
| Stripe | 12 | 8 |

against **18 and 18** company-matched rows on the old form (PRECISION.md, 2026-09-08). The
alternative was not 18 rows, it was zero.

#### 2. Brave now rate-limits Apify's shared datacenter pool

Independent of the operator. With the operator already removed, an **interleaved** A/B of the same
14 queries, alternating tier inside one Actor run:

| Tier | Result |
|---|---|
| automatic (datacenter) | 14 of 14 **HTTP 429**, 0 result blocks |
| **RESIDENTIAL** | 14 of 14 **HTTP 200**, 244 result blocks |

Interleaved on purpose: run one tier after the other and a pool the first half burnt is
indistinguishable from a tier difference.

**Moved to `RESIDENTIAL`**, and unlike the last time that was considered, it is affordable. Billed
transfer over three 14-query runs is 0.36 to 0.53 MB, which at $8/GB is $0.0029 to $0.0041 of proxy
per company, roughly $0.21 to $0.29 per 1,000 queries against the README's own $4 per 1,000 ceiling.
All-in a run costs about $0.0058 against $0.0448 of net revenue, about an **87% margin**.

The exposure worth knowing is the **zero-row run**: it still spends all 14 residential queries and
earns only the start fee, so a mistyped company name now costs about $0.004 rather than about
nothing. That is the price of the Actor answering at all.

#### Workflows moved to a self-hosted runner

`ubuntu-latest` cannot start on this account: jobs fail in ~3 seconds with no log and the annotation
*"recent account payments have failed or your spending limit needs to be increased"*. The last
successful hosted run in this estate was 2026-09-10, so for two weeks every build and every Store
release was silently impossible, including the fix for a bug that fails a buyer's first click.

All three workflows now run on `[self-hosted, hermes-tqm]`. **Hermes, not `tqm-scrapers-01`**: on the
scraper VM every runner executes as `tech`, that user has passwordless sudo, and the Neon production
URL sits on disk at `/opt/tqm-leadgen/.env`, so an Actor-repo workflow could read production database
credentials. Safe only while this repo is private with no outside contributors.

### 2026-09-10 — `anyTitleMatched`: the field to actually filter on

`titleMatched` is scoped to the **single query that surfaced the person**, because the Actor runs one
search per title. So someone found by the `VP Engineering` search whose headline reads
*"Engineering Manager at Acme"* reported `titleMatched: false` — even when `Engineering Manager` was
also on the buyer's list. The Actor was **understating its own precision**.

Measured on a live 25-row run (Stripe; `CTO`, `VP Engineering`, `Head of Engineering`,
`Engineering Manager`): `titleMatched` reports **2 of 25**; the headline carries one of the four
requested titles on **7 of 25**. (Both figures are post-fix — the first pass read 3 and 9, and three
of those rows were the substring false positives described below.)

- **Added `anyTitleMatched`** — does the headline carry ANY requested title. Documented in the Store
  listing as the field to filter on.

#### 🔴 And a false positive that was already live

Adding the field exposed one in the existing matcher. `normalise()` strips **every** separator, so a
three-letter title matches inside an unrelated word: `normalise('CTO')` is `cto`, a substring of
*dire**cto**r*. On the same 25-row run it wrongly matched *"Representative Director @ Stripe, Japan"*
and *"Director Of Engineering at Stripe"* — and the first of those was **already wrong in the live
`titleMatched` field**, not something this change introduced.

Title matching now uses `titleAppears()`, which requires **word boundaries** while tolerating
punctuation and spacing inside the title itself, so *"Head of Engineering"* still matches
*"Head of Engineering, Crypto @ Stripe"*. It is deliberately strict about filler words —
*"VP Engineering"* does not match *"VP of Engineering"*; put both in `titles` if you want both. An
over-permissive matcher is the thing this function exists to prevent.

`companyAppears()` keeps the glued comparison **on purpose** — *"BearingPoint"* must still match
*"Bearing Point"*, and a company name is long enough for that to be safe in a way a 3-letter title
is not.

> ⚠️ **A normaliser built for one comparison is not automatically right for the next one.** The
> glued form exists so company names survive spacing differences. Reusing it on a 3-letter title
> turned it into a substring search.

- **`titleMatched` keeps its meaning** — "does the headline carry *that one* query's title" — and is
  only made correct. It is a live field on a paid listing, and quietly redefining what it reports
  would break anyone already filtering on it. Its documentation was accurate all along; it was the
  *useful* question that was missing, not the honest one.

### Unreleased

#### 2026-09-08 (later still) — `set-start-fee.yml`: change the listing's start fee without the Console

The `tqm` Store account has no local credentials by design, so a pricing change previously meant
doing it by hand in the Console. `APIFY_TOKEN_TQM` already lives in this repo for the mirror, and
`PUT /v2/acts/{id}` accepts `pricingInfos`, so the change can be made reviewably instead.

**The job can only ever LOWER the price**, because Apify treats the two directions as different
operations:

| direction | consequence |
|---|---|
| lowering | immediate, reversible, unlimited |
| **raising** | 14 days notice, **one** significant change per month, **cannot be cancelled once scheduled** |

A fat-fingered decimal in the raising direction is therefore not a recoverable mistake — it books the
listing's only pricing change for the month and then charges buyers. So the job refuses to raise,
mutates exactly one field, and verifies against a **fresh read** (not against what it sent) that the
new price landed *and* that the per-result tiered pricing survived. It aborts before writing if the
actor is not `PAY_PER_EVENT`, or if the primary per-result event is missing its tiers — writing that
structure back would damage the listing.

#### 2026-09-08 (later still) — `includeUnmatched` now defaults to FALSE, and the cost model was wrong

Measured across 5 real companies at `maxResults: 25`, plus a nonsense-company control:

| company | loose (was default) | strict (now default) |
|---|---|---|
| Stripe | 25 rows / **18** matched | 25 / **25** |
| Trainline | 25 / **18** | 25 / **25** |
| Monzo | 25 / **19** | 25 / **25** |
| Pluralsight | 25 / **16** | 25 / **25** |
| Cronofy (small) | 25 / **11** | 18 / **18** |
| `Zzqxwv Nonexistent Holdings` | 25 / **0** | **FAILED, 0 rows** |
| **total** | **125 / 82 (66%)** | **118 / 118 (100%)** |

**Strict does not cost recall the way it looks like it should.** The search loop runs until it fills
`maxResults`, so on 4 of 5 companies it returned the *same* row count with every row confirmed — it
simply searched harder. Only tiny Cronofy fell short, and there the buyer gets 18 usable rows instead
of 11 while paying for 7 fewer.

**The billing argument is the decisive one.** On per-result pricing, a loose run against a company
name that does not match LinkedIn headline conventions bills for every unmatched row: the control
returned 25 rows, 0 matched, and would have charged for all 25. Strict returns nothing and fails with
a diagnosis naming the flag, so a mistyped company costs no result charges at all. Research mode is
one boolean away — opt IN to noise, never be opted in by default.

`.actor/INPUT_SCHEMA.json` and the Store README are updated to match; the README stated the old
default as fact in three places, which LISTING-AUDIT-01 counts as a listing defect.

#### Cost: MONEY-03's premise does not survive measurement

| engine | cost/run | dominant component |
|---|---:|---|
| Google (old) | **$0.0489** | `PROXY_SERPS` **$0.042** — 14 queries × $0.003 |
| Brave (now) | **$0.0005–0.0020** | compute only; **proxy $0.00** |

MONEY-03 concluded *"cost tracks RUNTIME, not rows — a browser plus residential proxy."* It tracked
**queries**: a flat $0.042 of SERP proxy per run, charged for all 14 whatever came back. That is also
why *"the zero-row run was the most expensive"* — the outage still paid for every query and returned
nothing. Runtime was never the driver.

Measured on Brave, cost is **flat at ~$0.0006 whether the run returns 5 rows or 100**:

| `maxResults` | 5 | 25 | 50 | 100 |
|---|---:|---:|---:|---:|
| cost | $0.00075 | $0.00053 | $0.00064 | $0.00053 |

So the **$0.05 start fee is now 25–80× the run cost** and taxes exactly the small trial runs that
build ranking — a 5-row trial costs $0.0675 instead of $0.0175, **3.9×**. Agreed to drop it to the
$0.00001 minimum, matching MONEY-02 for the other six; lowering a price is immediate on Apify, with
no notice period. **That change is Console-side and is not made by this commit.**

#### 2026-09-08 (later) — FIXED: the search engine is now Brave, and the Actor returns people again

The outage below is resolved. The engine was swapped Google → **Brave**, which publishes plain
`https://www.linkedin.com/in/<slug>` hrefs.

**Nothing about the product promise changes.** The listing says it *"queries the public search index"*
and *"never touches LinkedIn — no cookie, no session, no account to be banned"*. Both remain exactly
true; the engine was always an implementation detail.

| | Before | After |
|---|---|---|
| Host | `www.google.com/search` | `search.brave.com/search` |
| Proxy | `GOOGLE_SERP` | **automatic (datacenter)** — `GOOGLE_SERP` is Google-only and errors against Brave |
| Result pairing | href + `<h3>` inside one `<a>` | href + title inside one `data-type="web"` block |

**The pairing guarantee is preserved, by different means.** The Google parser matched href-then-`<h3>`
inside a single anchor, because an earlier forward-searching version paired 26 of 36 rows (72%) with
another person's profile URL. Brave does not wrap the title in the result anchor, so pairing is
enforced by **segmenting on `data-type="web"`** — one segment per organic result — and taking the
first href and the title from within that segment. A href and a title can only meet if they are in
the same result, which is the same guarantee the anchor gave.

Verified offline against a live Brave response before deploying: 6/6 results parsed, and every
person's name matches their own profile slug — no drift.

**Also fixed: LinkedIn's own pages were becoming people.** Brave returns LinkedIn login/interstitial
pages among the organic results for a `site:linkedin.com/in` query, and such a block still carries a
real profile href. Measured on Trainline: the title `"LinkedIn: Log In or Sign Up"` came back paired
with `https://uk.linkedin.com/in/oraziocotroneo` — **a real person's URL under a name that is not a
name.** That is not a junk row, it is a plausible identity for the wrong human, the same class as
LI-URL-01's 759 synthesised URLs. `looksLikeChrome()` now rejects them on the title, since the href
beside them is perfectly valid and cannot be used to tell them apart.

Three smaller things that came with it:

- Brave emits the profile href **twice** per result (title link and thumbnail link); only the first
  is taken, or every person would be deduped against themselves.
- Brave appends the site name to result titles — `"Alan Curiel - Stripe | LinkedIn"`, sometimes
  `"… | Professional Profile | LinkedIn"`. `splitTitle()` now strips it. Left in, every headline ends
  `"| LinkedIn"`, which is noise in the output and dilutes `companyMatched` / `titleMatched`.

⚠️ **Expect `titleMatched` and `seniorityMatched` to fall.** Brave's titles carry a shorter headline
than Google's did — `"Vivian Ren - Stripe"` where Google gave `"Kapil Agarwal - Software Engineer at
Stripe"`. `companyMatched` is unaffected. The four flags are the product and they stay honest, but
the *rates* in PRECISION.md were measured on Google and no longer describe this actor.

#### 2026-09-08 — OUTAGE: Google stopped publishing result URLs, and this Actor exits green anyway

**The Actor currently returns zero people for every input**, including the README's own Trainline
example. Two canary runs, 14 queries each, 0 rows, both `SUCCEEDED`.

**Cause.** Google's result anchors no longer contain the destination URL. They now point at an
opaque redirect:

```
href="/goto?url=CAESXAHrOzAV1ZPOYqrNfTgYv5-tnKlY6qFhRf59lfrnj3tH3fk4Bij1N1DC..."
```

The `RESULT` regex requires a literal `https://xx.linkedin.com/in/<slug>` inside the href, so it
matches nothing. Measured 2026-09-08 against the live `GOOGLE_SERP` proxy — the search itself is
**fine**: `site:linkedin.com/in "CTO at Stripe"` returns HTTP 200 and 10 organic results. Only the
URL is gone. The page carries the person's name, headline and follower count, but the profile URL
appears nowhere in the HTML.

Ruled out, all measured rather than reasoned about:

| Attempted recovery | Result |
|---|---|
| Resolve `/goto?url=<token>` directly | **HTTP 400** — needs session context |
| Legacy user agents (Googlebot, curl, Lynx, FF78) hoping for old `/url?q=` markup | 0 URLs on all four |
| Bing (its `u=a1<base64>` redirect used to be decodable) | 10 results, format also changed, 0 URLs |
| DuckDuckGo HTML endpoint | HTTP 202, 0 URLs |
| Brave Search | **works** — 5 plain `https://www.linkedin.com/in/...` URLs |

Swapping the search source is a product decision, not a hotfix — it changes what the listing's
"No Cookies, No Account" claim rests on — so it is **not** done here.

#### Fixed here: a zero-row run no longer exits green

On pay-per-result a buyer who gets nothing has paid the start fee for a green tick and an empty
table, and cannot tell "this company has no decision-makers" from "this Actor is broken". Three
counters (`serpOk`, `resultBlocks`, `anchors`) now separate the cases that need opposite responses:

| Condition | Diagnosis |
|---|---|
| `serpOk === 0` | the search never answered — proxy/network fault |
| `resultBlocks === 0` | it answered with no results — company name or titles are wrong |
| `resultBlocks > 0`, `anchors === 0` | **results exist and cannot be read — Actor fault, do not retry** |
| `anchors > 0`, all dropped | every link filtered — try `includeUnmatched` |

The run now calls `Actor.fail()` with that sentence instead of `log.info('done')`. The third row is
the current outage, and it is exactly the case that looked identical to the second until today.

### 1.0.0 — 2026-08-18

First release. Built for TQM's ENRICH-WF `person` stage after Apollo's free plan was measured
returning `last_name_obfuscated` ("Wa\*\*\*r") on 50/50 people across 5 companies — a masked surname is
unusable input for any email-resolution provider, which all need first + last + domain.

- Queries the search index via Apify's `GOOGLE_SERP` proxy rather than scraping LinkedIn. Direct
  company-page scraping was measured first and yields ~1 person; `/people/` is client-rendered and
  yields 0; authenticated scraping risks the account.
- Phrase query `"{title} at {company}"`, not two loose quoted terms — measured 50–63% company-match
  vs 10% for the loose form.
- `companyMatched` is computed from the **headline only**. An earlier build matched 1,200 characters
  of surrounding result HTML and scored 2/25, one of which was a person at a *different* company
  with a similar name.
- `titleMatched` reported separately, because Google honours the phrase loosely.
- Dedupe on profile slug, not URL — LinkedIn serves one person from every country subdomain.
- Rows without a surname are dropped.

### 2026-09-04 — the dataset schema documented nothing

Found by a pre-launch audit comparing every listing claim against the code across the portfolio.

`.actor/dataset_schema.json` was `"fields": {}` — **zero documented output fields** — while its one
`views.overview` referenced five fields (`fullName`, `headline`, `matchedTitle`, `companyMatched`,
`profileUrl`) that were never declared. Every sibling actor documents 18–25 fields; this one, the
highest-priced in the portfolio and the one with the largest niche, documented none.

All **12** fields are now titled, described and exampled, and three views are declared —
`overview`, `outreach` (mail-merge shape) and `evidence` (the four match flags side by side). Every
field referenced by a view is now declared; the schema build asserts it.

The descriptions carry the *reasoning*, not just the field name, because the four booleans are the
product:

- `companyMatched` is judged on the **headline only**, and says so — an earlier version searched the
  surrounding result HTML and matched the search engine's own page furniture.
- `titleMatched` is reported separately from the query because the engine honours a phrase loosely.
- `seniorityMatched` is deliberately broader than the title queries.
- `looksPastRole` is the "right company, person has left" flag.
- `lastName` notes that first-name-only rows are dropped rather than returned.

Also: `titles` had `default: []` in the input schema while the real default is seven hard-coded
titles, so a buyer could not see what they were opting into. Added as a `prefill`.

No behaviour change; schema and listing metadata only.
