Give it a company name, get the decision-makers: full name, headline and public LinkedIn profile URL. It never touches LinkedIn - no cookie, no session, no account of yours at risk. Every row shows its own match evidence, so you set your own precision bar instead of trusting ours.
2026-09-26: the zero-row outage, two faults, not one
Every run was returning zero rows. There turned out to be two independent faults, and they have
to be separated because only one of them is about the proxy.
1. Brave refuses site:linkedin.com/in
Not rate limiting, not IP reputation: the engine blocks that one query shape. Measured from an Apify
datacenter IP, an Apify RESIDENTIAL IP and an unproxied home connection, all three identical:
Query
Result
site:linkedin.com/in "Head of Engineering at Trainline"
HTTP 429
site:linkedin.com/in (bare)
HTTP 429
site:linkedin.com (no /in)
HTTP 200
site:bbc.co.uk trains
HTTP 200
"Head of Engineering at Trainline" (no operator)
HTTP 200
The 429 body is Brave's "your request has been flagged as being suspicious" captcha page. The
operator works and other domains work, so no proxy tier can buy its way out of this one.
The operator is gone. The quoted headline phrase was already doing nearly all the constraining,
and the parser only ever reads a linkedin.com/in href out of a result block, so a non-LinkedIn
result is skipped rather than mis-parsed. It costs some yield, honestly recorded:
rows
high confidence
Trainline
16
10
Stripe
12
8
against 18 and 18 company-matched rows on the old form (PRECISION.md, 2026-09-08). The
alternative was not 18 rows, it was zero.
2. Brave now rate-limits Apify's shared datacenter pool
Independent of the operator. With the operator already removed, an interleaved A/B of the same
14 queries, alternating tier inside one Actor run:
Tier
Result
automatic (datacenter)
14 of 14 HTTP 429, 0 result blocks
RESIDENTIAL
14 of 14 HTTP 200, 244 result blocks
Interleaved on purpose: run one tier after the other and a pool the first half burnt is
indistinguishable from a tier difference.
Moved to RESIDENTIAL, and unlike the last time that was considered, it is affordable. Billed
transfer over three 14-query runs is 0.36 to 0.53 MB, which at $8/GB is $0.0029 to $0.0041 of proxy
per company, roughly $0.21 to $0.29 per 1,000 queries against the README's own $4 per 1,000 ceiling.
All-in a run costs about $0.0058 against $0.0448 of net revenue, about an 87% margin.
The exposure worth knowing is the zero-row run: it still spends all 14 residential queries and
earns only the start fee, so a mistyped company name now costs about $0.004 rather than about
nothing. That is the price of the Actor answering at all.
Workflows moved to a self-hosted runner
ubuntu-latest cannot start on this account: jobs fail in ~3 seconds with no log and the annotation
"recent account payments have failed or your spending limit needs to be increased". The last
successful hosted run in this estate was 2026-09-10, so for two weeks every build and every Store
release was silently impossible, including the fix for a bug that fails a buyer's first click.
All three workflows now run on [self-hosted, hermes-tqm]. Hermes, not tqm-scrapers-01: on the
scraper VM every runner executes as tech, that user has passwordless sudo, and the Neon production
URL sits on disk at /opt/tqm-leadgen/.env, so an Actor-repo workflow could read production database
credentials. Safe only while this repo is private with no outside contributors.
2026-09-10 — anyTitleMatched: the field to actually filter on
titleMatched is scoped to the single query that surfaced the person, because the Actor runs one
search per title. So someone found by the VP Engineering search whose headline reads
"Engineering Manager at Acme" reported titleMatched: false — even when Engineering Manager was
also on the buyer's list. The Actor was understating its own precision.
Measured on a live 25-row run (Stripe; CTO, VP Engineering, Head of Engineering,
Engineering Manager): titleMatched reports 2 of 25; the headline carries one of the four
requested titles on 7 of 25. (Both figures are post-fix — the first pass read 3 and 9, and three
of those rows were the substring false positives described below.)
Added anyTitleMatched — does the headline carry ANY requested title. Documented in the Store
listing as the field to filter on.
🔴 And a false positive that was already live
Adding the field exposed one in the existing matcher. normalise() strips every separator, so a
three-letter title matches inside an unrelated word: normalise('CTO') is cto, a substring of
director. On the same 25-row run it wrongly matched "Representative Director @ Stripe, Japan"
and "Director Of Engineering at Stripe" — and the first of those was already wrong in the live
titleMatched field, not something this change introduced.
Title matching now uses titleAppears(), which requires word boundaries while tolerating
punctuation and spacing inside the title itself, so "Head of Engineering" still matches
"Head of Engineering, Crypto @ Stripe". It is deliberately strict about filler words —
"VP Engineering" does not match "VP of Engineering"; put both in titles if you want both. An
over-permissive matcher is the thing this function exists to prevent.
companyAppears() keeps the glued comparison on purpose — "BearingPoint" must still match
"Bearing Point", and a company name is long enough for that to be safe in a way a 3-letter title
is not.
⚠️ A normaliser built for one comparison is not automatically right for the next one. The
glued form exists so company names survive spacing differences. Reusing it on a 3-letter title
turned it into a substring search.
titleMatched keeps its meaning — "does the headline carry that one query's title" — and is
only made correct. It is a live field on a paid listing, and quietly redefining what it reports
would break anyone already filtering on it. Its documentation was accurate all along; it was the
useful question that was missing, not the honest one.
Unreleased
2026-09-08 (later still) — set-start-fee.yml: change the listing's start fee without the Console
The tqm Store account has no local credentials by design, so a pricing change previously meant
doing it by hand in the Console. APIFY_TOKEN_TQM already lives in this repo for the mirror, and
PUT /v2/acts/{id} accepts pricingInfos, so the change can be made reviewably instead.
The job can only ever LOWER the price, because Apify treats the two directions as different
operations:
direction
consequence
lowering
immediate, reversible, unlimited
raising
14 days notice, one significant change per month, cannot be cancelled once scheduled
A fat-fingered decimal in the raising direction is therefore not a recoverable mistake — it books the
listing's only pricing change for the month and then charges buyers. So the job refuses to raise,
mutates exactly one field, and verifies against a fresh read (not against what it sent) that the
new price landed and that the per-result tiered pricing survived. It aborts before writing if the
actor is not PAY_PER_EVENT, or if the primary per-result event is missing its tiers — writing that
structure back would damage the listing.
2026-09-08 (later still) — includeUnmatched now defaults to FALSE, and the cost model was wrong
Measured across 5 real companies at maxResults: 25, plus a nonsense-company control:
company
loose (was default)
strict (now default)
Stripe
25 rows / 18 matched
25 / 25
Trainline
25 / 18
25 / 25
Monzo
25 / 19
25 / 25
Pluralsight
25 / 16
25 / 25
Cronofy (small)
25 / 11
18 / 18
Zzqxwv Nonexistent Holdings
25 / 0
FAILED, 0 rows
total
125 / 82 (66%)
118 / 118 (100%)
Strict does not cost recall the way it looks like it should. The search loop runs until it fills
maxResults, so on 4 of 5 companies it returned the same row count with every row confirmed — it
simply searched harder. Only tiny Cronofy fell short, and there the buyer gets 18 usable rows instead
of 11 while paying for 7 fewer.
The billing argument is the decisive one. On per-result pricing, a loose run against a company
name that does not match LinkedIn headline conventions bills for every unmatched row: the control
returned 25 rows, 0 matched, and would have charged for all 25. Strict returns nothing and fails with
a diagnosis naming the flag, so a mistyped company costs no result charges at all. Research mode is
one boolean away — opt IN to noise, never be opted in by default.
.actor/INPUT_SCHEMA.json and the Store README are updated to match; the README stated the old
default as fact in three places, which LISTING-AUDIT-01 counts as a listing defect.
Cost: MONEY-03's premise does not survive measurement
engine
cost/run
dominant component
Google (old)
$0.0489
PROXY_SERPS$0.042 — 14 queries × $0.003
Brave (now)
$0.0005–0.0020
compute only; proxy $0.00
MONEY-03 concluded "cost tracks RUNTIME, not rows — a browser plus residential proxy." It tracked
queries: a flat $0.042 of SERP proxy per run, charged for all 14 whatever came back. That is also
why "the zero-row run was the most expensive" — the outage still paid for every query and returned
nothing. Runtime was never the driver.
Measured on Brave, cost is flat at ~$0.0006 whether the run returns 5 rows or 100:
maxResults
5
25
50
100
cost
$0.00075
$0.00053
$0.00064
$0.00053
So the $0.05 start fee is now 25–80× the run cost and taxes exactly the small trial runs that
build ranking — a 5-row trial costs $0.0675 instead of $0.0175, 3.9×. Agreed to drop it to the
$0.00001 minimum, matching MONEY-02 for the other six; lowering a price is immediate on Apify, with
no notice period. That change is Console-side and is not made by this commit.
2026-09-08 (later) — FIXED: the search engine is now Brave, and the Actor returns people again
The outage below is resolved. The engine was swapped Google → Brave, which publishes plain
https://www.linkedin.com/in/<slug> hrefs.
Nothing about the product promise changes. The listing says it "queries the public search index"
and "never touches LinkedIn — no cookie, no session, no account to be banned". Both remain exactly
true; the engine was always an implementation detail.
Before
After
Host
www.google.com/search
search.brave.com/search
Proxy
GOOGLE_SERP
automatic (datacenter) — GOOGLE_SERP is Google-only and errors against Brave
Result pairing
href + <h3> inside one <a>
href + title inside one data-type="web" block
The pairing guarantee is preserved, by different means. The Google parser matched href-then-<h3>
inside a single anchor, because an earlier forward-searching version paired 26 of 36 rows (72%) with
another person's profile URL. Brave does not wrap the title in the result anchor, so pairing is
enforced by segmenting on data-type="web" — one segment per organic result — and taking the
first href and the title from within that segment. A href and a title can only meet if they are in
the same result, which is the same guarantee the anchor gave.
Verified offline against a live Brave response before deploying: 6/6 results parsed, and every
person's name matches their own profile slug — no drift.
Also fixed: LinkedIn's own pages were becoming people. Brave returns LinkedIn login/interstitial
pages among the organic results for a site:linkedin.com/in query, and such a block still carries a
real profile href. Measured on Trainline: the title "LinkedIn: Log In or Sign Up" came back paired
with https://uk.linkedin.com/in/oraziocotroneo — a real person's URL under a name that is not a
name. That is not a junk row, it is a plausible identity for the wrong human, the same class as
LI-URL-01's 759 synthesised URLs. looksLikeChrome() now rejects them on the title, since the href
beside them is perfectly valid and cannot be used to tell them apart.
Three smaller things that came with it:
Brave emits the profile href twice per result (title link and thumbnail link); only the first
is taken, or every person would be deduped against themselves.
Brave appends the site name to result titles — "Alan Curiel - Stripe | LinkedIn", sometimes
"… | Professional Profile | LinkedIn". splitTitle() now strips it. Left in, every headline ends
"| LinkedIn", which is noise in the output and dilutes companyMatched / titleMatched.
⚠️ Expect titleMatched and seniorityMatched to fall. Brave's titles carry a shorter headline
than Google's did — "Vivian Ren - Stripe" where Google gave
"Kapil Agarwal - Software Engineer at Stripe"
. companyMatched is unaffected. The four flags are the product and they stay honest, but
the rates in PRECISION.md were measured on Google and no longer describe this actor.
2026-09-08 — OUTAGE: Google stopped publishing result URLs, and this Actor exits green anyway
The Actor currently returns zero people for every input, including the README's own Trainline
example. Two canary runs, 14 queries each, 0 rows, both SUCCEEDED.
Cause. Google's result anchors no longer contain the destination URL. They now point at an
opaque redirect:
The RESULT regex requires a literal https://xx.linkedin.com/in/<slug> inside the href, so it
matches nothing. Measured 2026-09-08 against the live GOOGLE_SERP proxy — the search itself is
fine: site:linkedin.com/in "CTO at Stripe" returns HTTP 200 and 10 organic results. Only the
URL is gone. The page carries the person's name, headline and follower count, but the profile URL
appears nowhere in the HTML.
Ruled out, all measured rather than reasoned about:
Attempted recovery
Result
Resolve /goto?url=<token> directly
HTTP 400 — needs session context
Legacy user agents (Googlebot, curl, Lynx, FF78) hoping for old /url?q= markup
0 URLs on all four
Bing (its u=a1<base64> redirect used to be decodable)
10 results, format also changed, 0 URLs
DuckDuckGo HTML endpoint
HTTP 202, 0 URLs
Brave Search
works — 5 plain https://www.linkedin.com/in/... URLs
Swapping the search source is a product decision, not a hotfix — it changes what the listing's
"No Cookies, No Account" claim rests on — so it is not done here.
Fixed here: a zero-row run no longer exits green
On pay-per-result a buyer who gets nothing has paid the start fee for a green tick and an empty
table, and cannot tell "this company has no decision-makers" from "this Actor is broken". Three
counters (serpOk, resultBlocks, anchors) now separate the cases that need opposite responses:
Condition
Diagnosis
serpOk === 0
the search never answered — proxy/network fault
resultBlocks === 0
it answered with no results — company name or titles are wrong
resultBlocks > 0, anchors === 0
results exist and cannot be read — Actor fault, do not retry
anchors > 0, all dropped
every link filtered — try includeUnmatched
The run now calls Actor.fail() with that sentence instead of log.info('done'). The third row is
the current outage, and it is exactly the case that looked identical to the second until today.
1.0.0 — 2026-08-18
First release. Built for TQM's ENRICH-WF person stage after Apollo's free plan was measured
returning last_name_obfuscated ("Wa***r") on 50/50 people across 5 companies — a masked surname is
unusable input for any email-resolution provider, which all need first + last + domain.
Queries the search index via Apify's GOOGLE_SERP proxy rather than scraping LinkedIn. Direct
company-page scraping was measured first and yields ~1 person; /people/ is client-rendered and
yields 0; authenticated scraping risks the account.
Phrase query "{title} at {company}", not two loose quoted terms — measured 50–63% company-match
vs 10% for the loose form.
companyMatched is computed from the headline only. An earlier build matched 1,200 characters
of surrounding result HTML and scored 2/25, one of which was a person at a different company
with a similar name.
titleMatched reported separately, because Google honours the phrase loosely.
Dedupe on profile slug, not URL — LinkedIn serves one person from every country subdomain.
Rows without a surname are dropped.
2026-09-04 — the dataset schema documented nothing
Found by a pre-launch audit comparing every listing claim against the code across the portfolio.
.actor/dataset_schema.json was "fields": {} — zero documented output fields — while its one
views.overview referenced five fields (fullName, headline, matchedTitle, companyMatched,
profileUrl) that were never declared. Every sibling actor documents 18–25 fields; this one, the
highest-priced in the portfolio and the one with the largest niche, documented none.
All 12 fields are now titled, described and exampled, and three views are declared —
overview, outreach (mail-merge shape) and evidence (the four match flags side by side). Every
field referenced by a view is now declared; the schema build asserts it.
The descriptions carry the reasoning, not just the field name, because the four booleans are the
product:
companyMatched is judged on the headline only, and says so — an earlier version searched the
surrounding result HTML and matched the search engine's own page furniture.
titleMatched is reported separately from the query because the engine honours a phrase loosely.
seniorityMatched is deliberately broader than the title queries.
looksPastRole is the "right company, person has left" flag.
lastName notes that first-name-only rows are dropped rather than returned.
Also: titles had default: [] in the input schema while the real default is seven hard-coded
titles, so a buyer could not see what they were opting into. Added as a prefill.
No behaviour change; schema and listing metadata only.