# LinkedIn Smart Job Description & Recruiter Extractor (`fanndev/linkedin-job-recruiter-extractor`) Actor

Read LinkedIn job postings in full and turn each description into structured data: the tech stack named, certifications and clearances demanded, the pay range, seniority and the applicant count. No login required. Recruiter profiles are added when you supply your own li\_at cookie.

- **URL**: https://apify.com/fanndev/linkedin-job-recruiter-extractor.md
- **Developed by:** [Faisal Ahdan naufal](https://apify.com/fanndev) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.80 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## LinkedIn Smart Job Description & Recruiter Extractor

Most job scrapers give you a title, a company and a link. The interesting part of a job posting is the 5,000 words underneath that, and this actor reads them: which technologies the employer named, which certifications and clearances they demand, what they will pay, and how many people have already applied.

**No login is needed.** LinkedIn's public jobs API answers anonymous requests, paginates cleanly and carries the full description. This is the one LinkedIn surface that is genuinely open, and it is why this actor can run without putting any account at risk.

### What it pulls out of the description

```json
{
  "jobTitle": "Senior Software Engineer - Python",
  "companyName": "Venmo",
  "location": "San Jose, CA",
  "workplaceType": "HYBRID",
  "employmentType": "Full-time",
  "industries": "Financial Services",
  "techStack": ["AWS", "Azure", "CI/CD", "Django", "Flask", "GCP", "Git", "Kafka", "NoSQL", "Python"],
  "certifications": [],
  "salaryMin": 143500.0,
  "salaryMax": 212850.0,
  "salaryCurrency": "$",
  "salaryPeriod": "annually",
  "salarySource": "parsed-from-description",
  "applicants": 114,
  "descriptionLength": 8110
}
```

`techStack` is matched as whole words against a curated list, not by substring search. That distinction is the difference between a usable dataset and a noisy one: a naive matcher turns every "R\&D" into a hit for **R** and every "Going to market" into a hit for **Go**. Longer terms also consume their matches first, so "React Native" is not double-counted as "React" — while "SQL and PostgreSQL" correctly reports both.

`salaryMin` / `salaryMax` handle both thousand-separator conventions, so `$120,000` and `Rp 8.000.000` parse equally well, and `salaryPeriod` is kept separately because a bare `45,000` means wildly different things per hour and per year. Where LinkedIn renders an employer-provided pay block, that wins over anything mined from the prose, and `salarySource` tells you which you got.

### The filter behaviour you need to know

This was measured, not assumed:

- **`datePosted` really filters.** It maps to LinkedIn's own `f_TPR` and is applied server-side.
- **Workplace, experience level and job type do not.** LinkedIn's public endpoint accepts those parameters and then ignores them — passing `f_WT=2` (remote) returns results identical to passing nothing, same first ten ids, every time.

Rather than offer filters that quietly do nothing, this actor applies those three **client-side**, against each posting's own criteria grid and description, after the detail page is read. They work — but a filtered run still costs one request per posting examined, so filtering does not make a run cheaper.

Pagination is ten results per page and the endpoint stops serving anything past roughly **975 postings per query** (1,000 returns HTTP 400). Coverage comes from more queries, not a bigger cap.

### The recruiter field, and when it is empty

LinkedIn does not render the "Meet the hiring team" block to logged-out visitors. Checked across eight postings: none carried it. So `recruiterName` and `recruiterProfileUrl` are filled **only** when you supply your own `li_at` cookie, and stay empty otherwise.

If you want them: sign in to LinkedIn in your browser, open DevTools → Application → Cookies → `www.linkedin.com`, copy the **value** of `li_at`, paste it into `sessionCookie`. This actor never asks for your password and never signs in on your behalf.

**The honest risk:** scraping while signed in breaches LinkedIn's User Agreement and accounts do get restricted for it. Use a residential proxy in your own country (a datacenter exit IP is the mismatch that flags sessions), keep `maxJobs` modest, and use an account you would not mind losing. Everything else in this actor works without a cookie — the recruiter field is the only thing you are trading that risk for.

### Input

```json
{
  "searchQueries": ["python developer", "data engineer"],
  "location": "United States",
  "datePosted": "pastWeek",
  "requiredTech": ["Kubernetes"],
  "onlyWithSalary": true,
  "maxJobs": 200
}
```

You can also skip search entirely and pass `jobUrls` — full `/jobs/view/` links, `?currentJobId=` links or bare numeric ids.

### Who this is for

- **Job aggregators** — structured postings without maintaining a parser.
- **Recruitment agencies** — which employers are hiring for what, and at what pay.
- **Bootcamps and course designers** — `techStack` aggregated across a few thousand postings is a direct read on what the market is actually asking for, which is a better curriculum input than a survey.
- **Candidates** — `applicants` tells you how contested a role is before you spend an hour on the application.

### Rate gating

LinkedIn answers request bursts with **HTTP 999** and a ~1.5 KB stub. It is per-IP and per-burst, not a TLS fingerprint gate — five browser fingerprints all returned 200 on the same page. The actor sleeps a random 2–8 seconds between requests, backs off exponentially when gated, and rotates fingerprints. If the log shows rate-gate warnings, raise both delay bounds first.

### No email addresses

`resolveCompanyDomain` gives you the employer's web domain. It does not invent `first.last@domain.com`. A guessed address that bounces costs you your sending domain's reputation, which is far more expensive than a missing column — put the name and domain through a verification tool that actually checks.

### Related actors

- **LinkedIn Post Engagers Scraper** — who engaged with a post, with their comments.
- **LinkedIn Top Voice & Competitor Activity Monitor** — content strategy, no cookie needed.

# Actor input Schema

## `searchQueries` (type: `array`):

Job titles or keywords to search, one per line. Each query is paginated ten results at a time and stops returning anything past roughly 975 postings, so coverage comes from using more queries rather than raising the per-query cap.

## `location` (type: `string`):

Free-text location applied to every query, exactly as you would type it on LinkedIn - 'United States', 'Jakarta, Indonesia', 'Remote'. Leave empty for worldwide.

## `jobUrls` (type: `array`):

Specific postings to read, instead of or alongside a search. Full /jobs/view/ URLs, ?currentJobId= links and bare numeric ids all work.

## `startUrls` (type: `array`):

The same job links in the request-list format, for callers that already keep one.

## `datePosted` (type: `string`):

The one filter LinkedIn's public jobs endpoint actually honours - it is applied server-side, before results are returned.

## `workplaceType` (type: `string`):

Applied AFTER each description is read, not by LinkedIn. LinkedIn's public endpoint accepts its own workplace filter and then ignores it - passing it returns results identical to passing nothing - so this actor infers the workplace type from the description text and filters on that instead. It means filtered runs still cost one request per posting read.

## `employmentType` (type: `string`):

Also applied client-side, matched against the posting's own criteria grid. Same reason as above.

## `seniorityLevel` (type: `string`):

Client-side filter against the posting's stated seniority.

## `requiredTech` (type: `array`):

Keep only postings whose description names at least one of these, matched against the parsed tech stack rather than raw text - so 'R' will not match 'R\&D' and 'React' will not match 'React Native' alone.

## `onlyWithSalary` (type: `boolean`):

Keep only postings where a pay range was found. Be aware most postings state none: in a sample of eight, two carried an employer-provided range.

## `includeFullDescription` (type: `boolean`):

Keep the complete description on each row. Turn off for a much smaller dataset when you only want the parsed fields.

## `includeRecruiter` (type: `boolean`):

Requires a session cookie below. LinkedIn does not render the 'Meet the hiring team' block to logged-out visitors - none of eight sampled postings carried it - so without a cookie the recruiter fields stay empty.

## `sessionCookie` (type: `string`):

Optional, and only needed for recruiter details. Copy the li\_at cookie value from a browser where you are already signed in to LinkedIn (DevTools > Application > Cookies > www.linkedin.com > li\_at). This actor never asks for your password and never signs in for you. Scraping while signed in breaches LinkedIn's User Agreement and can get an account restricted - use one you are willing to risk.

## `resolveCompanyDomain` (type: `boolean`):

Opens each employer's LinkedIn page once to read their website, giving you the company domain for mail-merge work. Cached per company, so forty employers cost forty extra requests regardless of how many postings they have. No email address is ever guessed from it - a fabricated address that bounces costs you your sending reputation.

## `maxJobsPerQuery` (type: `integer`):

Cap per search query, before detail pages are read. LinkedIn stops serving results past roughly 975 per query.

## `maxJobs` (type: `integer`):

Overall ceiling across every query and URL after de-duplication. Each posting costs one detail request on top of the search requests.

## `minDelaySeconds` (type: `integer`):

LinkedIn answers request bursts with HTTP 999 and a near-empty page. A randomised gap between requests is what keeps a run under that gate; this is the floor of that range.

## `maxDelaySeconds` (type: `integer`):

The ceiling of the randomised gap. Raise both bounds if the log shows rate-gate warnings.

## `exportFormats` (type: `array`):

Besides the dataset, write ready-made files into this run's key-value store.

## `proxyConfiguration` (type: `object`):

Genuinely optional here: the public jobs endpoint has no WAF and no IP block, and Apify's datacenter range reaches it. Worth turning on for long runs, for results localised to another country, and always when you supply a session cookie.

## Actor input object example

```json
{
  "searchQueries": [
    "python developer",
    "data engineer"
  ],
  "location": "United States",
  "datePosted": "",
  "workplaceType": "",
  "employmentType": "",
  "seniorityLevel": "",
  "onlyWithSalary": false,
  "includeFullDescription": true,
  "includeRecruiter": false,
  "resolveCompanyDomain": false,
  "maxJobsPerQuery": 50,
  "maxJobs": 200,
  "minDelaySeconds": 2,
  "maxDelaySeconds": 8,
  "exportFormats": [],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One JOB row per posting with its parsed tech stack, certifications and pay range, plus NOTICE and ERROR rows.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "python developer",
        "data engineer"
    ],
    "location": "United States"
};

// Run the Actor and wait for it to finish
const run = await client.actor("fanndev/linkedin-job-recruiter-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": [
        "python developer",
        "data engineer",
    ],
    "location": "United States",
}

# Run the Actor and wait for it to finish
run = client.actor("fanndev/linkedin-job-recruiter-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "python developer",
    "data engineer"
  ],
  "location": "United States"
}' |
apify call fanndev/linkedin-job-recruiter-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fanndev/linkedin-job-recruiter-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/YeMjx35KHBzoPfeXW/builds/Ewi64jQaJpAC2dC3z/openapi.json
