# Job Post Enricher: Salary, Seniority, Remote, Visa & Skills (`nerolabs/job-post-enricher`) Actor

Adds AI fields to job posts from any LinkedIn, Indeed or career-site dataset, CSV or Google Sheet: yearly salary in one currency, seniority, remote or hybrid, visa sponsorship, skills, years of experience. Never invents pay. Charged per enriched job. Agent-ready: x402, MCP.

- **URL**: https://apify.com/nerolabs/job-post-enricher.md
- **Developed by:** [Adam Pearce](https://apify.com/nerolabs) (community)
- **Categories:** Jobs, AI, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.10 / 1,000 job enricheds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Job Post Enricher: Salary, Seniority, Remote, Visa & Skills

**Scraped 5,000 jobs from LinkedIn or Indeed and now need to know which ones pay over $120k, are remote, sponsor visas and want Python?** Point this Actor at the scraper's dataset (or a CSV or Google Sheet) and every job comes back with clean, sortable fields added next to your original columns:

- **Salary as yearly min and max in one currency** (USD, GBP, EUR or any ECB currency), plus the pay exactly as written, its period (hour, day, week, month, year), where it was found and a confidence flag. **Never invented:** if the job states no pay, the salary is empty, and every number is checked against the job's own text before it is kept.
- **Seniority**: intern, entry, junior, mid, senior, lead, manager, director, executive
- **Work mode**: remote, hybrid or onsite, plus remote scope ("US only", "3 days in office")
- **Visa sponsorship**: yes, no or unknown, only when the job says so, with the sentence quoted
- **Contract type**: full time, part time, contract, temporary, internship, apprenticeship, freelance
- **Job category**: software engineering, data and AI, sales, healthcare, skilled trades and 20 more
- **Required and nice-to-have skills**, years of experience (min and max) and a one-line summary
- Optional, same price: stated benefits, minimum education and languages required

It reads the output of any jobs scraper without setup. Title, description, company, location and salary columns are detected automatically, including nested ones such as Indeed's `description.text` and `baseSalary`. Tested on LinkedIn Jobs, Indeed (US and UK) and Greenhouse career-site data.

### Who uses it

- **Job boards and aggregators**: add salary, remote and seniority filters to scraped listings without hand-tagging.
- **Recruiters and staffing agencies**: find every senior, remote, visa-sponsoring role in a scrape in seconds.
- **Salary and labour-market research**: yearly pay in one currency across hourly, daily and monthly postings.
- **Job seekers and career coaches**: shortlist roles by pay, work mode and visa sponsorship.
- **Sales teams selling to hiring companies**: see which companies hire for which skills and levels.
- **AI agents**: a jobs scraper run in, a clean structured table out, in one call.

### How to use it

1. Run any jobs scraper on Apify (for example a LinkedIn Jobs or Indeed scraper), then pick its dataset in **Dataset**. Or paste a CSV, Excel, JSON or Google Sheet link into **File or Google Sheet URL**, or paste whole job posts into **Job post texts**.
2. Choose the **Salary currency** for the yearly columns (default USD, or `original` to keep each job's own currency).
3. Run. 100 jobs take about 2 minutes and 1,000 jobs about 20 to 25 minutes. Results stream into the dataset as they are ready; a summary (salary coverage, seniority and work mode counts, top skills) is saved as `OUTPUT`.

Chain it after a scraper with an Apify integration or webhook ("run this Actor when the scraper finishes, with its dataset ID") so every scrape arrives enriched.

### Sample output

One job from the default example (a public Stripe posting), with its description column left out here:

```json
{
  "title": "Senior Software Engineer, Backend",
  "companyName": "Stripe",
  "location": "Seattle, WA",
  "jobUrl": "https://stripe.com/jobs/search?gh_jid=8230952",
  "enrichStatus": "ok",
  "enrichDetail": null,
  "jobSummary": "Senior Software Engineer, Backend available in Seattle, WA with 40 hours/week and 50% telecommuting option.",
  "seniority": "senior",
  "workMode": "hybrid",
  "remoteScope": null,
  "contractType": "unknown",
  "jobCategory": "software_engineering",
  "salaryMinYearly": 206090,
  "salaryMaxYearly": 285600,
  "salaryYearlyCurrency": "USD",
  "salaryStatedText": "Salary: $206,086.00 - $285,600.00/yr.",
  "salaryStatedMin": 206086,
  "salaryStatedMax": 285600,
  "salaryStatedCurrency": "USD",
  "salaryStatedPeriod": "year",
  "salarySource": "description",
  "salaryConfidence": "high",
  "visaSponsorship": "unknown",
  "visaEvidence": null,
  "yearsExperienceMin": 5,
  "yearsExperienceMax": 5,
  "requiredSkills": ["Java", "Scala", "Python", "Distributed systems", "Big Data", "Spark", "SQL", "Testing (JUnit, Mockito, TestNG)", "Leadership"],
  "preferredSkills": []
}
```

An hourly UK electrician job reads `salaryStatedText: "Pay: £17.00-£19.00 per hour"`, `salaryStatedPeriod: "hour"`, and with `salaryCurrency: "GBP"` gives `salaryMinYearly: 35360` (17 x 40 hours x 52 weeks). A part-time job that says "£12.60 an hour, 20 hours a week" uses its own 20 hours: `salaryMinYearly: 13100`.

### Accuracy

Hand-checked on 30 real jobs (10 LinkedIn, 10 Indeed, 10 Greenhouse; 17 with a stated salary, 13 without), across repeated runs on 03-10-2026:

- **Salary: right on every job the AI answered, in every test run** (correct amount, period and currency, or correctly empty), including a "From 45k a year" line, a weekly travel-nurse wage, a daily contractor rate, an Italian base-and-OTE posting and a job with two location-based ranges.
- **Seniority: 27 to 30 of 30** (90 to 100%) across six runs. Misses are close calls such as "Senior Engineering Manager" read as senior or lead instead of manager.
- **Work mode**: nursing, electrician and trades jobs read as onsite 10 of 10 in five of six runs.
- **On 200 more real jobs** (100 LinkedIn, 100 Indeed US and UK): in the final test run all 118 jobs with a visible pay line came back with their salary (97% before the last fixes), and **no salary was invented**: every returned amount appears in the job's own text.

It is an AI reading text, so spot-check a sample before relying on fields in bulk. `salaryConfidence` is `low` when the pay looks unusual (for example a job board labelling an hourly rate as weekly), when the currency had to be guessed or the AI's quote does not match the text, and `enrichDetail` says why.

### Pricing

Pay per event, no subscription:

| Event | Price |
|---|---|
| Job enriched (one job returned with the AI fields) | $0.003 ($3 per 1,000 jobs), less on Bronze, Silver and Gold plans |
| CSV or Excel export file | $0.01 each, only if requested |
| Webhook delivery | $0.02, only on a 2xx response |

**Examples:** enriching a 1,000-job LinkedIn scrape costs about $3. A daily job board refresh of 500 new jobs costs about $1.50 a day. Rows with no title or description, and jobs where the AI step fails, are never charged. Set the run's maximum charge to cap spend; the Actor stops cleanly at the limit.

### Input

| Field | What it does |
|---|---|
| `datasetId` | Apify dataset of jobs (any scraper's output) |
| `fileUrl` | CSV, TSV, Excel, JSON or JSON Lines link, or a Google Sheet shared as "anyone with the link" |
| `jobTexts` | Whole job posts pasted as plain text |
| `data` | Rows as inline JSON |
| `salaryCurrency` | Currency for the yearly columns, default `USD`; `original` keeps each job's own |
| `hoursPerWeek` | Used for hourly pay, default 40 |
| `includeExtras` | Adds `statedBenefits`, `educationLevel`, `languagesRequired` |
| `titleField`, `descriptionField`, `companyField`, `locationField`, `salaryField` | Only if auto-detection picks the wrong column; nested fields use a dot (`description.text`) |
| `keep`, `keepOriginalFields`, `exportFormats`, `outputDatasetName`, `webhookUrl` | Output options |

### FAQ

**Does it guess salaries for jobs that do not list one?** No. A salary is filled only when the job states an amount ("competitive salary" stays empty), and each number must appear in the job's own text. Bonuses, sign-on bonuses, referral fees and allowances are not counted as salary.

**Which scrapers does it work with?** Any that output a title or description column: LinkedIn Jobs, Indeed, Glassdoor, ZipRecruiter, Seek, Reed, Google Jobs and career-site scrapers for Greenhouse, Lever, Ashby, Workday and others. The run log names the columns it used.

**What about jobs in other languages?** The AI reads job posts in most languages; output fields and summaries are in English.

**How is pay converted to yearly?** Hourly x hours per week (default 40) x 52, daily x 260, weekly x 52, monthly x 12. Currency uses the European Central Bank reference rate of the day, recorded in the `OUTPUT` summary. The stated pay is always kept as written too.

**Is my data stored?** Only the job text is sent to the AI model (OpenAI) for the run; nothing is kept afterwards.

**Is it agent-ready?** Yes. Pay per event, works through Apify's MCP server and x402 payments, with typed fields and a summary record an agent can read.

If this saved you reading job posts by hand, a short review on the Apify Store helps a lot and is read personally. Questions or a scraper whose columns are not picked up? Open an issue on the Issues tab and it will be answered the same day.

# Actor input Schema

## `datasetId` (type: `string`):

An Apify dataset of job posts, for example the output of a LinkedIn Jobs, Indeed, Glassdoor or career-site (Greenhouse, Lever, Ashby, Workday) scraper run. Every original column is kept and the AI fields are added alongside. Title, description, company, location and salary columns are found automatically. Use the picker rather than typing an ID.

## `fileUrl` (type: `string`):

A public link to a CSV, TSV, Excel, JSON or JSON Lines file with one job per row (an ATS export, a job board feed, a sheet of scraped jobs). A normal Google Sheets link works: share it as 'Anyone with the link can view'.

## `fileFormat` (type: `string`):

Leave on 'Detect automatically' unless the link has no file extension and the server reports the wrong content type.

## `sheetName` (type: `string`):

Which sheet to read from an Excel workbook. Defaults to the first sheet.

## `jobTexts` (type: `array`):

A plain list of whole job posts pasted as text, one per entry, for a quick run with no file. Use a dataset, file or Google Sheet to keep your own columns alongside the AI fields.

## `data` (type: `array`):

Job rows as inline JSON, an alternative to a dataset or file. Each object needs a title or description field.

## `salaryCurrency` (type: `string`):

The currency for the yearly salary columns (salaryMinYearly, salaryMaxYearly), as a three-letter code: USD, GBP, EUR, CAD, AUD, INR and any other European Central Bank currency. Converted at the day's ECB reference rate. Enter 'original' to keep each job's own currency. The stated pay is always kept as written as well.

## `hoursPerWeek` (type: `integer`):

Used only to turn hourly pay into a yearly figure (hourly x hours x 52) when the job does not state its own weekly hours; stated hours ("20 hours a week") win. Daily pay uses 260 days a year, weekly 52 weeks, monthly 12 months.

## `includeExtras` (type: `boolean`):

Also return statedBenefits (up to 10), educationLevel (minimum stated) and languagesRequired. Same price per job; adds a little run time.

## `titleField` (type: `string`):

The column holding the job title. Leave empty to detect it. Nested fields use a dot, e.g. 'job.title'.

## `descriptionField` (type: `string`):

The column holding the job description text or HTML, e.g. 'descriptionText', 'description' or 'description.text'. Leave empty to detect it.

## `companyField` (type: `string`):

The column holding the company name, e.g. 'companyName' or 'employer.name'.

## `locationField` (type: `string`):

The column holding the job location, e.g. 'location' or 'location.city'.

## `salaryField` (type: `string`):

The column(s) holding the posted salary, comma separated, e.g. 'salary' or 'baseSalary,salaryText'. The description is always read for pay too.

## `keep` (type: `string`):

Filtering happens after the jobs are enriched, so it does not make a run cheaper.

## `keepOriginalFields` (type: `boolean`):

Keep every column from the input row next to the AI fields, so results line up with the scraper's own data (job URL, posted date, applicants). Turn off for the AI fields plus title, company and location only.

## `exportFormats` (type: `array`):

Also write the results as a downloadable CSV or Excel file, linked from the run's output. Lists (skills) are joined with semicolons so they fit a spreadsheet.

## `outputDatasetName` (type: `string`):

Also append every kept row to a named dataset that persists across runs, building one growing jobs table. Not charged again.

## `webhookUrl` (type: `string`):

POST the run summary (counts, top skills, export links) to this URL when the run finishes, for Slack, Zapier, Make, n8n or your own API. Charged only on a confirmed 2xx response.

## `concurrency` (type: `integer`):

How many AI requests run in parallel (4 jobs each). The default 4 handles about 1,000 jobs in 20 to 25 minutes; higher values can hit the AI service's rate limit, which the Actor then waits out.

## `maxItems` (type: `integer`):

A hard ceiling on how many rows are read from the input, as a safety net on a large dataset. The run's own maximum charge limit also stops it cleanly.

## Actor input object example

```json
{
  "fileUrl": "https://nerolabs-samples.nerolabs.workers.dev/sample-jobs.json",
  "fileFormat": "auto",
  "salaryCurrency": "USD",
  "hoursPerWeek": 40,
  "includeExtras": false,
  "keep": "all",
  "keepOriginalFields": true,
  "exportFormats": [],
  "concurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

Every original job row with salary, seniority, work mode, visa sponsorship, contract type, category, skills and summary added.

## `summary` (type: `string`):

Fields used, jobs enriched, salary coverage, seniority and work mode counts, top skills, export links and warnings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "fileUrl": "https://nerolabs-samples.nerolabs.workers.dev/sample-jobs.json",
    "salaryCurrency": "USD"
};

// Run the Actor and wait for it to finish
const run = await client.actor("nerolabs/job-post-enricher").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "fileUrl": "https://nerolabs-samples.nerolabs.workers.dev/sample-jobs.json",
    "salaryCurrency": "USD",
}

# Run the Actor and wait for it to finish
run = client.actor("nerolabs/job-post-enricher").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "fileUrl": "https://nerolabs-samples.nerolabs.workers.dev/sample-jobs.json",
  "salaryCurrency": "USD"
}' |
apify call nerolabs/job-post-enricher --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nerolabs/job-post-enricher"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qxXugkobskqxtmMwC/builds/26lMv8TjqTyfm7Rds/openapi.json
