LinkedIn Profile Phone Number Scraper: Seniority & Title Filter avatar

LinkedIn Profile Phone Number Scraper: Seniority & Title Filter

Pricing

$19.99/month + usage

Go to Apify Store
LinkedIn Profile Phone Number Scraper: Seniority & Title Filter

LinkedIn Profile Phone Number Scraper: Seniority & Title Filter

LinkedIn Profile Phone Number Scraper extracts publicly listed phone numbers from LinkedIn profiles and linked pages. Build targeted contact lists by role, industry, or company. Ideal for sales teams running outbound campaigns.

Pricing

$19.99/month + usage

Rating

0.0

(0)

Developer

Scraper Engine

Scraper Engine

Maintained by Community

Actor stats

1

Bookmarked

35

Total users

2

Monthly active users

13 days ago

Last modified

Share

LinkedIn Phone Number Scraper — Seniority, Job Title and Function

LinkedIn Phone Number Scraper By Seniority & Title Filter finds publicly indexed linkedin.com pages that show a phone number, and returns each one as a structured JSON row: phone_number normalised to a +-and-digits form, the profile url, the result title and description it came from, plus a jobTitle, a seniorityLevel and a jobFunction the Actor works out from that text. Optionally keep only the bands you want. Give it keywords and a country, press Start, and rows land in the dataset live.

⚠️ This Actor reads Google's index, not LinkedIn profiles, and seniority is not a LinkedIn field. It never sends a request to LinkedIn. It runs site:linkedin.com "<dial code>" "<keyword>" queries through Apify's GOOGLE_SERP proxy, pulls phone numbers out of the result blocks, and then classifies each result by running regular expressions over the result's title and snippet text. seniorityLevel and jobFunction are the Actor's own labels, produced by keyword matching — not values published by LinkedIn, and not a LinkedIn search filter. The rule set is printed in full below so you can see exactly what maps to what.

What is LinkedIn Phone Number Scraper By Seniority & Title Filter?

LinkedIn Phone Number Scraper By Seniority & Title Filter is an Apify Actor that turns a keyword and a country into a list of publicly indexed linkedin.com pages that show a phone number in visible text, each one tagged with a classified seniority band and job function. For every keyword it pages through Google search results, keeps only links containing linkedin.com, scans each result block for a phone-shaped string that matches your country's dial code, classifies the title and snippet, and writes one flat row per number kept.

No LinkedIn account, login, session cookie or li_at is required, because LinkedIn is never contacted — the only credential involved is your Apify token. The Actor does need Apify's GOOGLE_SERP proxy, which it selects for you.

It is built for sales and lead-generation teams who want a call list narrowed to decision-makers, recruiters sourcing by function, and developers piping classified contact records into a CRM or an AI pipeline.

What LinkedIn profile data is publicly available to scrape?

A Google result row for a linkedin.com page carries a title, a URL and a snippet — nothing more. Everything else about the person behind that URL sits on the LinkedIn page itself, which this Actor never opens.

Data CategoryOn the Google result row (returned here)Needs a LinkedIn page fetch or login
Phone number written in visible indexed text✅ Returned as phone_number
The linkedin.com profile or page URL✅ Returned as url
Google's result title, usually name plus headline✅ Returned as title
Snippet text surrounding the number✅ Returned as description
Job title, seniority band, job function✅ Derived by the Actor from the two text fields above, not read from LinkedInLinkedIn's own structured headline and experience
Employment history, company, dates, locationProfile page — use an enrichment Actor
Connections, skills, endorsements, educationLogin required
Email addresses as their own fieldNot parsed — only phone numbers are extracted
Profiles Google has not indexed with a numberOut of reach — coverage is Google's index

LinkedIn Phone Number Scraper By Seniority & Title Filter only returns publicly visible data — what any visitor sees on a Google results page. Nothing behind a login wall.

What data can I extract with LinkedIn Phone Number Scraper By Seniority & Title Filter?

Every row carries eleven keys: four read from the Google result, four echoed from your run configuration, and three produced by the Actor's classifier. No key is ever omitted or set to null.

Field NameDescription
platformPlatform label the search ran against. With the schema's only allowed value this is always Linkedin.com
keywordThe exact keyword from your keywords list that produced this result
titleText of the Google result's <h3> heading, whitespace-normalised — typically Name - Headline
descriptionThe Google snippet the number was extracted from, or "" when no snippet element matched
urlThe linkedin.com URL exactly as it appears on the result link
phone_numberThe number normalised to + followed by digits only, e.g. +442079460958
countryCountry name with the dial code stripped off, e.g. United Kingdom
dial_codeDial code used to build the query and validate the number, e.g. +44
jobTitleActor-derived. The segment of title (or, failing that, of description) that looks most like a role. "" when nothing role-like is found
seniorityLevelActor-derived. One of c_level, vp, director, manager, senior, junior, unknown
jobFunctionActor-derived. One of founder, sales, marketing, engineering, hr, finance, operations, other

Result identity and source fields — what was actually scraped

title, description, url and phone_number are the only four fields taken from the page. platform, keyword, country and dial_code are echoed from your input onto every row so a merged multi-run dataset stays sortable — keyword in particular is the partition key you will use most, since one run covers many keywords and rows are not otherwise separated.

description is the highest-information text field in the row: it is the snippet Google showed, so it usually contains the number in its original written form along with the surrounding words. It falls back to "" rather than null when the snippet element does not match.

country and dial_code come from splitting your single country input — the schema stores it as one combined string like United Kingdom (+44), and the parser splits it on the trailing parenthesised code. The dial code then does real work in two places: it becomes a quoted term in the Google query, and a candidate number is discarded unless the normalised result starts with it. Normalisation is conservative — text is NFKC-normalised, non-breaking spaces converted and zero-width characters (U+200B, U+200C, U+200D, U+FEFF) stripped, a leading 00 becomes +, and a national-format number has one trunk 0 dropped before the dial code is prepended.

One thing to know about phone_number: the Actor scans the whole result block's visible text and takes the first phone-shaped string that survives the dial-code check. That is usually the person's number, but on a company page or a snippet that quotes a switchboard it can be an organisation's line rather than the individual's.

Classification fields — what the Actor derived

jobTitle, seniorityLevel and jobFunction are computed locally from title and description after the row is parsed. They are not scraped, not supplied by LinkedIn, and not a LinkedIn search parameter. seniorityLevel is always one of the seven listed values and jobFunction always one of the eight — a result that matches nothing gets unknown and other respectively rather than being left blank. The full rule set is in the next section.

🤖 Add-on: Need additional LinkedIn data?

A number, a URL and a band is a lead, not a profile. LinkedIn Profile Company Enrichment Scraper takes the url values from this output and resolves them into full profile and company records, which is what you need to verify a classification before acting on it. LinkedIn Phone Number Scraper: Lead Scoring is the sibling to reach for when you want a numeric score and a title-keyword match instead of fixed bands, and LinkedIn Jobs Scraper With Salary Range Filters covers hiring signals for the same companies.

How seniority and job function are classified

Both labels are a heuristic over title strings. There is no machine learning, no LinkedIn field and no external lookup — the classifier lowercases title, tests it against an ordered list of regular expressions, and returns the first band that matches. If the title matches nothing, it repeats the whole pass over description as a fallback. Order matters, because the first match wins.

Seniority bands, in the order they are tested:

BandMatches (word-boundary patterns, case-insensitive)
c_levelceo, cfo, cto, coo, cmo, cio, chro — each also matching separated forms such as c.e.o — plus a generic three-letter c_o pattern that catches cpo, cro, cdo and anything shaped the same way; chief, founder, co-founder, owner, partner, managing director, and president when not preceded by vice or v
vpvice president, vp, v.p., svp, evp, avp, head of
directordirector, dir.
managermanager, mgr, team lead, lead, supervisor, principal
seniorsenior, snr, sr, staff
juniorjunior, jr, intern, trainee, entry-level, associate, assistant, graduate, apprentice
unknownNothing above matched in either the title or the snippet

Job functions, in the order they are tested: founder (founder, co-founder, owner, entrepreneur, proprietor), sales (sales, account executive, account manager, business development, biz dev, bdr, sdr, revenue, partnership(s)), marketing (marketing, brand, growth, seo, sem, content, social media, communication(s), public relations, pr, demand gen, cmo), engineering (engineer, engineering, developer, software, programmer, devops, data scien*, machine learning, architect, cto, full-stack, backend, frontend), hr (hr, human resource(s), recruit*, talent, people ops/operations/team, l&d, learning and development, chro), finance (finance, financial, accountant, accounting, controller, treasur*, audit*, cfo, investment, bookkeep*), operations (operations, ops, supply chain, logistics, procurement, project manager, program manager, coo), and other when none matched.

Read that table as what it is: a matcher over free text, not a judgement about a person's actual rank. Some consequences are worth stating outright, because they will show up in your data.

  • Head of Growth classifies as vp, Growth Lead classifies as manager, and Director of Growth classifies as director — three titles a human would probably treat as one seniority, split across three bands by the words used.
  • Ordering overrides intuition in both directions. Senior Vice President is vp, not senior, because vp is tested first — that one is deliberate. Associate Director is director, because director is tested before junior. Managing Director is c_level, not director. Partner Manager is c_level, because partner sits in the C-level list.
  • The snippet fallback can classify the wrong person. When the title carries no role word, the classifier scans the Google snippet — which may be company boilerplate, a testimonial, or text about somebody else entirely. A snippet containing the word chief will produce c_level.
  • jobTitle is a best-effort text slice, not a parsed field. The title is split on -, |, ·, , , , at and @, and the first segment containing any role-ish substring is returned. That substring test is not word-bounded, so a name like Amanda Leadbetter contains lead and can be returned as the job title. The same splitting turns Vice-President into a President fragment for jobTitle, while seniorityLevel still correctly reads vp from the unsplit string.
  • No accuracy figure is published for this classifier, and none is claimed. Treat the bands as a first-pass sort, and keep title and description beside them so a human or an LLM can check the call.

Why not build this yourself?

LinkedIn publishes no public API that returns arbitrary third-party profiles. Its developer programs are partner-gated and scoped to content you own or are authorised for, and none of them expose profile search or contact details — so there is no official API to compare against here. The alternative to this Actor is writing the Google-SERP plumbing and the classifier yourself, and that is where the time goes.

The proxy is HTTP-only. Apify's GOOGLE_SERP proxy intercepts the plain request, runs the search itself and returns SERP HTML. It is the mechanism, not an accessory, and it constrains how the request has to be formed.

Block detection is a trap, not a substring match. The obvious implementation greps the response body for captcha or /sorry/. The source comment in this Actor records why that fails: a healthy 326 KB result page carrying ten real results contained captcha twice and /sorry/index five times as page chrome and inline JavaScript. Matching those throws away good pages and returns zero rows. This Actor keys on the genuine interstitial's own body copy — the unusual traffic phrasings — plus a non-200 status, and leaves genuinely empty pages to the empty-page streak logic downstream.

Then there is the long tail. A rotating user-agent and accept-language pool, jittered delays, a three-attempt retry budget with a fresh proxy URL per attempt, ten-at-a-time pagination, a phone regex whose output has to survive zero-width padding, 00 prefixes and trunk zeros, a country string that has to be split into a name and a dial code, and an ordered classifier where a single pattern in the wrong position silently reassigns thousands of leads to the wrong band. All of it is maintained here.

How to use LinkedIn Phone Number Scraper By Seniority & Title Filter

The Actor runs on Apify. Start it from the Apify Console or call it through the Apify API — nothing else to sign up for.

  1. Open LinkedIn Phone Number Scraper By Seniority & Title Filter on Apify and click Try for free
  2. Add one or more entries to Keywords / Usernames / URLs (keywords) — a role, industry or niche term works best, because the term has to appear in Google's indexed text
  3. Pick a Country (country) from the dropdown — this sets both the quoted dial code in the query and the filter applied to every number found
  4. Optionally pick Seniority Levels (filter) (seniorityLevels) and Job Functions (filter) (jobFunctions). Leave either empty to keep everything on that dimension
  5. Set Max Phone Numbers (maxPhoneNumbers) if 20 per keyword is not what you want, and leave Proxy Configuration on Apify Proxy
  6. Click Start, then download the dataset as JSON, CSV or Excel, or read it through the Apify API

keywords and country are the two required inputs, and they fail differently. An empty or missing keywords list logs No keywords provided. Please fill the 'keywords' array with at least one value. and the run finishes immediately — no rows, nothing charged. A missing country does not stop the run: the code falls back to United Kingdom (+44), so an API caller who omits the field quietly gets UK numbers. The Console blocks that case before the run starts; the fallback only bites callers building the input JSON themselves.

How to scale to bulk lead extraction

keywords is a list, so bulk is the normal mode — add as many terms as you like and each is searched in turn within one run. Keywords are processed sequentially, not in parallel, and rows are pushed as they are found rather than at the end.

maxPhoneNumbers applies per keyword, not per run, and it counts rows after the seniority and function filters: ten keywords at maxPhoneNumbers: 200 is a 2,000-kept-row target. Each keyword pages through Google ten results at a time until its own limit is met or it hits three consecutive pages that yield nothing parseable.

One run covers one country, because country is a single string. To cover several regions, run the Actor once per country with the same keywords list — country and dial_code are stamped on every row, so the datasets merge cleanly afterwards.

What can you do with LinkedIn seniority-filtered phone lead data?

  • 📞 A sales development rep building a decision-maker call list sets seniorityLevels to ["c_level", "vp", "director"] and jobFunctions to ["sales", "marketing"], then works the rows in keyword order, opening url to verify each person before the number is dialled.
  • 🎯 An account-based marketer sizing a segment counts rows per seniorityLevel for the same keywords across two country values, using jobFunction to see which departments actually publish contact details in each market.
  • 🧹 A RevOps engineer enriching a CRM matches incoming phone_number values against existing records, routes each one by dial_code, and writes seniorityLevel and jobFunction onto the record as routing attributes with url kept as the provenance link.
  • 🔎 A recruiter sourcing a function rather than a title runs broad keywords with jobFunctions set to ["engineering"] and reads jobTitle and description to judge whether a hit is an individual practitioner or an agency before shortlisting.
  • 🤖 An AI engineer building a lead-qualification agent indexes description, title and jobTitle into a vector store with seniorityLevel, jobFunction and keyword as metadata filters, then has the model re-check the Actor's band against the raw text — exactly the sanity pass a regex classifier needs.

Every one of these is callable from an agent framework over the Apify API, since the Actor is a standard HTTP-triggered run.

How does the Actor handle rate limits and blocking?

Every request goes out through Apify's GOOGLE_SERP proxy group, which is what performs the search — you never create a proxy account or rotate an IP yourself. The group is forced on in code regardless of what you put in proxyConfiguration, and if it cannot be initialised the run raises an error rather than falling back to a direct connection.

Each page request gets up to three attempts. Before each one the Actor sleeps a random 1–2 seconds and rebuilds its headers, drawing a user-agent from a pool of four current desktop browser strings and an accept-language from three. After a failed or blocked attempt it waits a random 3–6 seconds and asks the proxy for a fresh URL, so a blocked exit IP is not immediately reused. HTTP timeouts are 60 seconds total, with 30 seconds each for connect and socket read.

A response counts as blocked on any non-200 status, or when the body contains one of three literal unusual traffic interstitial phrases. Loose tokens like captcha and /sorry/ are deliberately not matched, for the reason given above.

There is no CAPTCHA solving in this Actor, and none is claimed. When all three attempts fail, that page counts as one empty page for the keyword; after three consecutive empty or failed pages the Actor stops that keyword, keeps the rows already collected, and moves to the next one. A run never dies because one keyword got blocked.

⬇️ Input

Two inputs are required: keywords and country. Everything else has a default.

ParameterRequiredTypeDescriptionExample Value
keywordsYesarrayA list of keywords, Linkedin usernames, or profile URLs to search for. Each entry is searched separately. Prefilled with ["marketing"].["marketing", "operations"]
countryYesstringCountry to scrape related phone numbers for, as one combined "Country Name (+dial code)" string. The dial code is used to normalize and filter phone numbers. 194 values in the dropdown. Default "United Kingdom (+44)"."United Kingdom (+44)"
platformNostringTarget platform for the site: search. Enum has one value: Linkedin. Default "Linkedin"."Linkedin"
maxPhoneNumbersNointegerMaximum number of matching phone numbers to keep per keyword, counted AFTER the seniority/function filter. The scraper stops once this limit is reached. Minimum 1, maximum 10000. Default 20.50
engineNostringScraping engine. Enum has one value: legacy, which uses the Apify GOOGLE_SERP proxy to query Google search. Default "legacy"."legacy"
seniorityLevelsNoarrayKeep only leads whose classified seniority matches one of these. Empty keeps all seniorities. Allowed values c_level, vp, director, manager, senior, junior, unknown. Default [].["c_level", "vp", "director"]
jobFunctionsNoarrayKeep only leads whose classified job function matches one of these. Empty keeps all functions. Allowed values sales, marketing, engineering, hr, finance, operations, founder, other. Default [].["sales", "marketing"]
proxyConfigurationNoobjectProxy settings. The Actor always searches Google through the GOOGLE_SERP proxy group. Default {"useApifyProxy": true}.{"useApifyProxy": true}

Six honest notes on how these behave:

  • engine is read and then ignored. The value is passed into the proxy setup function and never referenced inside it. The enum has one value, so nothing changes either way — but do not expect a second engine to exist behind it.
  • platform is effectively fixed. The schema exposes one value, and two things pin the Actor to LinkedIn regardless: the search domain is built as linkedin.com from that value, and any result whose URL does not contain linkedin.com is dropped before a row is built. The value does change one visible thing — the platform field in the output, which comes back as Linkedin.com rather than Linkedin.
  • The two filters are ANDed, and both are optional. Set both and a row must satisfy both to survive. Set neither and every row with a matching phone number is kept, classified but unfiltered.
  • An unrecognised filter value is dropped, and an entirely unrecognised list disables that filter. The Console's select editor prevents this, but an API caller passing jobFunctions: ["biz dev"] gets a logged warning and a run that keeps all functions — more rows than intended, and every one of them charged. The seniority filter additionally accepts undocumented aliases beyond the schema enum: clevel, c-level, executive, exec, vice president, vicepresident, mgr, sr, jr and entry all normalise to canonical bands. jobFunctions has no alias table — only the eight canonical values, case- and separator-insensitive, so "Sales" works and "Human Resources" does not.
  • maxPhoneNumbers is bounded by the schema, not by the code. The Console enforces 1–10000. The Actor reads the value with int(), so a non-numeric value fails the run outright, and 0 or a negative number gives a run that completes with zero rows instead of an error.
  • Twelve country options do not parse their dial code. The parser only recognises an all-digit parenthesised code, so the twelve entries whose code contains a hyphen — Antigua And Barbuda (+1-268), Bahamas (+1-242), Barbados (+1-246), Dominica (+1-767), Dominican Republic (+1-809), Grenada (+1-473), Jamaica (+1-876), Saint Kitts And Nevis (+1-869), Saint Lucia (+1-758), Saint Vincent And The Grenadines (+1-784), Trinidad And Tobago (+1-868) and Vatican City (+39-06) — come back with dial_code: "" and the code still embedded in country. The run logs a warning and continues, but with no dial-code term in the query and no dial-code filter on the numbers, so results are not country-scoped. Pick United States (+1) or Canada (+1) for North American numbers instead.

keywords is documented as accepting usernames and profile URLs as well as keywords. Mechanically they are all the same thing: the entry is wrapped in quotes and dropped into the Google query as an exact-phrase term. A plain term matches indexed text and works well; a full profile URL rarely appears as literal text on an indexed page, so it usually returns nothing.

Example input

{
"keywords": ["marketing", "operations", "revenue"],
"country": "United Kingdom (+44)",
"platform": "Linkedin",
"maxPhoneNumbers": 50,
"engine": "legacy",
"seniorityLevels": ["c_level", "vp", "director"],
"jobFunctions": ["sales", "marketing"],
"proxyConfiguration": {
"useApifyProxy": true
}
}

⬆️ Output

Typed, normalised JSON with a consistent shape across runs — the same eleven keys on every row, never omitted and never null. Rows are pushed live as each results page is parsed, so the dataset fills while the run is still going. Export as JSON, CSV or Excel, or read the dataset through the Apify API.

Every row in the dataset is a kept lead, and every one is charged as a single row_result event. This Actor writes no header rows, no diagnostic rows, no accounting rows and no error rows — there is no isError, no errorReason, no status and no rowType marker, because there is nothing to filter out. Failures live in the log, not in the data.

Filtered-out rows are never pushed, so they are never charged. The filter runs post-fetch, not server-side: Google has no seniority parameter, so the Actor fetches the page, parses it, classifies each result, and only then decides. A result that fails seniorityLevels or jobFunctions is logged as a dropped row with its classified band and URL, counted in the run summary, and discarded. It does not appear in the dataset in any form — so a tight filter costs you fetch time, not events.

A profile the classifier cannot place is not dropped by default. It is returned with seniorityLevel: "unknown" and jobFunction: "other", so nothing is lost. But if you set seniorityLevels to specific bands without including unknown, every unclassifiable lead disappears from your dataset with no trace in the output — likewise jobFunctions without other. Include unknown and other in your filter lists if you would rather review those leads than lose them.

One more cost note. The empty-page counter only advances when a page yields no parseable rows at all — not when its rows are all filtered out. A very narrow filter therefore keeps paging deep into Google's results, which lengthens a run without adding charged events.

Duplicates are removed per results page, not across the run. The URL-seen set is rebuilt for each SERP page, so the same linkedin.com URL can legitimately appear again on a later page of the same keyword, or under a different keyword that matched the same page — and each occurrence is its own charged row. Deduplicate on url yourself if uniqueness matters:

seen, unique = set(), []
for row in items:
if row["url"] not in seen:
seen.add(row["url"])
unique.append(row)

Example output

{
"platform": "Linkedin.com",
"keyword": "marketing",
"title": "Priya Raman - Head of Marketing at Northwind Analytics",
"description": "Head of Marketing at Northwind Analytics. Direct line +44 20 7946 0958. Demand generation, brand and lifecycle marketing for B2B SaaS teams ...",
"url": "https://uk.linkedin.com/in/priya-raman-marketing",
"phone_number": "+442079460958",
"country": "United Kingdom",
"dial_code": "+44",
"jobTitle": "Head of Marketing",
"seniorityLevel": "vp",
"jobFunction": "marketing"
}

A second row from the same run shows what an unclassifiable result looks like — the title carries no role word at all, so jobTitle is recovered from the snippet, the function matches on the word Operations, and the seniority falls through to unknown:

{
"platform": "Linkedin.com",
"keyword": "operations",
"title": "Tomás Herrera",
"description": "Operations · Bristol · Contact: 07700 900412 · Open to freelance and contract work",
"url": "https://uk.linkedin.com/in/tomas-herrera-4b81a2",
"phone_number": "+447700900412",
"country": "United Kingdom",
"dial_code": "+44",
"jobTitle": "Operations",
"seniorityLevel": "unknown",
"jobFunction": "operations"
}

How does it work?

For each keyword the Actor builds one Google query — site:linkedin.com "+44" "marketing" — and requests it through Apify's GOOGLE_SERP proxy over plain HTTP, ten results at a time by incrementing the start parameter. The proxy runs the search and returns the results HTML.

That HTML is parsed for Google's result wrapper blocks. Each block must contain an <h3> inside a link, the link must point at linkedin.com, and the block's visible text must contain a phone-shaped string that survives normalisation and starts with your country's dial code. Only then is a row built. The Actor then classifies the row's title and snippet into jobTitle, seniorityLevel and jobFunction, applies your optional filters, and pushes what survives.

Only publicly visible data is collected — this is Google's public index, with no LinkedIn account, cookie or session involved anywhere. And because the eleven output keys are defined by the Actor rather than by Google's markup, a Google layout change can affect coverage but never your field names or types.

Integrations

The Actor is a standard Apify Actor, so it works with anything that can call the Apify API or consume a dataset.

Calling the Actor from Python

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("<YOUR_USERNAME>/linkedin-profile-seniority-title-filter").call(run_input={
"keywords": ["marketing", "operations"],
"country": "United Kingdom (+44)",
"maxPhoneNumbers": 50,
"seniorityLevels": ["c_level", "vp", "director"],
"jobFunctions": ["sales", "marketing"],
"proxyConfiguration": {"useApifyProxy": True},
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row["seniorityLevel"], row["jobFunction"], row["phone_number"], row["url"])

Works in Go, Ruby, Node.js, cURL — any language that can make an HTTP request. Every row has the same eleven keys, so no branching or presence checks are needed.

No-code tools (n8n, Make, LangChain)

In n8n, use the Apify node — or an HTTP Request node pointed at the Apify run endpoint with your token — and pass the same JSON input shown above; a Remove Duplicates node keyed on url handles the cross-page duplicates described in the Output section. In Make, the Apify module supports run-and-wait, so a weekly sweep of your keyword list can feed a Google Sheets, Airtable or CRM step directly, filtered on seniorityLevel. In LangChain, wrap the run endpoint as a tool and hand the rows straight to the model, since they are already flat JSON. Apify schedules and webhooks cover recurrence and completion triggers.

Collecting publicly indexed information is generally treated as permissible where no login is bypassed, and this Actor reads only Google's public index. But the output is personal data about identifiable individuals, and filtering it by seniority deliberately produces a targeted list of named people — which changes your obligations rather than removing them.

Every substantive field here relates to a person: title typically carries their name and headline, url points to their profile, description is text about them, phone_number is a contact point, and jobTitle, seniorityLevel and jobFunction are inferences you have generated about their role. Under GDPR and UK GDPR the moment you store a row you are a data controller, and you need a lawful basis before storing or reusing it. If you rely on legitimate interest, that requires a documented balancing test, not an assumption — and a list built specifically to reach senior decision-makers is exactly the scenario where such a test is scrutinised. Because the data was not collected from the individuals themselves, Article 14 transparency duties apply: you have to tell people you hold their data, where it came from and what you are doing with it. Data minimisation means keeping only the fields you actually use, retention means deciding in advance how long you keep them, and subject requests — access, rectification, objection, erasure — must be answerable. Under CCPA/CPRA, California residents have parallel access, deletion and opt-out rights.

Note also that seniorityLevel, jobFunction and jobTitle are the Actor's inferences, not verified facts. An inaccurate inference about a person is still that person's personal data, and the right to rectification covers it.

Using a number is regulated separately, and more strictly, than collecting it. Direct marketing in the EU and UK falls under ePrivacy and PECR; calling and texting in the United States is governed by its own federal and state rules, including do-not-call registries, with CAN-SPAM covering commercial email. "It was public" is not a lawful basis for contacting someone. Screen against suppression and do-not-call lists before any outreach, and consult legal counsel if your use case involves bulk storage of personal data. This Actor gives you data; it does not give you permission to contact anyone.

❓ Frequently asked questions

What fields does this LinkedIn phone number scraper return?

Eleven, on every row: phone_number, url, title, description, jobTitle, seniorityLevel, jobFunction, keyword, country, dial_code and platform. The five you will use most are phone_number, url, seniorityLevel, jobFunction and description. See the data fields table above for what each one contains, and the classification section for how the derived three are produced.

Where does seniorityLevel come from — is it a LinkedIn field?

No. It is not published by LinkedIn, it is not a LinkedIn search filter, and it is not read from the profile. The Actor runs an ordered set of regular expressions over the Google result's title, falling back to the snippet, and returns the first band that matches. The complete rule set is printed in the classification section above. Treat it as a keyword heuristic over free text: Head of Growth lands in vp, Growth Lead lands in manager, and a title with no role word at all lands in unknown.

Does it require a LinkedIn account or login?

No — and it cannot use one. The Actor never sends a request to LinkedIn; it queries Google's index of linkedin.com pages through Apify's GOOGLE_SERP proxy. There is no cookie, no li_at, no session and no LinkedIn account to get flagged. The only credential involved is your Apify token.

What happens if a keyword returns no profiles at the requested seniority?

The run continues and finishes successfully. There is no error row and no diagnostic row — a keyword that yields nothing simply contributes no rows, and the Actor moves on to the next one. In code, treat an empty result set as "no indexed match at that band", not as a failure: if not items: ....

Four causes account for almost all empty results. Google may have no indexed linkedin.com page containing both your keyword and the dial-code digits in visible text. Your seniorityLevels or jobFunctions selection may be dropping everything that was found — the run log prints the dropped count plus a per-band breakdown of what was kept, so check there first. The keyword may be too narrow to appear in indexed text. Or you may have picked one of the twelve hyphenated-dial-code countries listed in the Input section, where the dial code never reaches the query.

Note that a filter which drops everything does not end the keyword early: the Actor keeps paging as long as pages contain parseable results, so a very narrow filter produces a longer run rather than a faster empty one.

Are profiles that do not match any seniority band dropped?

Only if you make them. By default they are returned with seniorityLevel: "unknown" and jobFunction: "other". But if you set seniorityLevels to specific bands and leave unknown out of the list, those leads are filtered out before they reach the dataset and there is no marker anywhere in the output to tell you it happened — only a line in the log and a count in the run summary. Add unknown (and other for functions) to your filter lists to keep them.

How many phone leads can I extract in one run?

maxPhoneNumbers accepts 1 to 10,000 and applies per keyword, counted after the filters, so twenty keywords at 100 each is a 2,000-row target in one run. Whether you reach it depends on Google's index and on how tight your filters are. A keyword stops early once three consecutive results pages return nothing parseable. Note also that duplicates are only removed within a single results page, so a large run can contain the same url more than once — deduplicate if you need unique contacts.

Can I scrape multiple keywords or countries at once?

Multiple keywords, yes — keywords is a list and each entry is searched in turn within one run, with the originating term stamped on every row as keyword. Countries, no: country is a single string, so one run covers one dial code. Run the Actor once per country with the same keyword list; because country and dial_code are on every row, the datasets merge into one multi-region table without any post-processing.

Does it work with Claude, ChatGPT and other AI agent tools?

Yes. It is callable as a standard HTTP-triggered run through the Apify API, so LangChain, CrewAI, n8n or a hand-written tool definition can invoke it and receive typed JSON with no parsing step. Because the row shape is fixed at eleven keys, an agent tool schema for it is trivial to write.

How does it compare to other LinkedIn phone number scrapers?

Checked on the Apify Store on 25 July 2026. api-empire/linkedin-profile-phone-number-scraper and scraper-engine/linkedin-profile-phone-number-scraper publish byte-identical READMEs, and both document six inputs — keywords, country, maxPhoneNumbers, platform, engine, proxyConfiguration — and eight output fields: platform, keyword, title, description, url, phone_number, country, dial_code. Those are the same eight this Actor returns, without the three classification keys. Neither listing documents a seniority filter, a job-function filter, a derived job-title field or a per-band run summary. Their shared example input uses "country": "Global", which is not one of the 194 enum values in this Actor's schema, and their example output shows "dial_code": "Auto-detected". weighty_marshmallow/phone-number-enrichment-from-linkedin-profile-url-80-coverage could not be retrieved when checked on the same date, so nothing about its behaviour is documented here.

The closest sibling is LinkedIn Phone Number Scraper: Lead Scoring in this same account, which uses the same Google-SERP approach and the same required keywords and country pair but replaces fixed bands with a computed leadScore, a minLeadScore threshold and a titleKeywords boost list. Use that one when you want ranking; use this one when you want categorical labels you can group and count on.

What happens when Google changes its layout or anti-bot system?

The scraper is maintained, and the output schema is defined by the Actor rather than by Google's markup, so your eleven field names and types stay put regardless. Layout changes affect coverage, not shape. Block detection deliberately avoids brittle body-text tokens like captcha and /sorry/, which appear on healthy pages, and keys on HTTP status plus the genuine interstitial's own copy instead.

Can I use it without managing proxies or browser infrastructure?

Yes. The Actor selects Apify's GOOGLE_SERP proxy group itself, requests a fresh proxy URL after a failed attempt, and rotates user-agents and accept-language headers per attempt. You never create a proxy account, pick a group or rotate an IP — and there is no browser to run, because only HTML is fetched and parsed. It does not solve CAPTCHAs, and it makes no claim to: a page that stays blocked after three attempts is counted as an empty page and the run carries on.

Which fields work best for AI training data and RAG indexing?

For RAG indexing, description carries the most information per record — it is the natural-language snippet the number appeared in — with title and jobTitle as shorter text fields, and seniorityLevel, jobFunction, keyword and dial_code as ready-made categorical metadata filters. For training data or feature extraction, phone_number, dial_code, country, platform, seniorityLevel and jobFunction are the most structurally consistent, since all six are normalised or drawn from a closed value set rather than passed through. Every value is a typed string and always present, so no imputation or presence-checking pass is needed. Two caveats: seniorityLevel and jobFunction are heuristic labels, so do not train on them as ground truth without review, and the personal-data obligations in the legal section apply before contact records go into a training set or a shared vector store.

Scraper NameWhat it extracts
LinkedIn Phone Number Scraper: Lead ScoringThe same SERP approach with a computed lead score and title-keyword boost instead of fixed bands
Linkedin Profile Phone Number ScraperThe same eight scraped keys without the classification fields or filters
LinkedIn Profile Company Enrichment ScraperFull profile and company records from the url values this Actor returns
LinkedIn Jobs Scraper With Salary Range FiltersJob listings with pay ranges, contract type and experience level
LinkedIn Post Comments ScraperComment text, commenter identity and reaction counts under a post
Instagram Phone Number ScraperThe same Google-SERP technique applied to instagram.com pages

💬 Your feedback

Found a bug, a title the classifier bands wrongly, or a keyword and country pair that returns nothing when you expect results? Open an issue on the Actor's Issues tab. Reports that include the exact input JSON and the offending title string are the fastest to reproduce and fix.