# Job Search API: Jobs from Company Career Sites (`digital_influx/job-search`) Actor

Search open jobs straight from 20,000+ company career sites (Greenhouse, Ashby, Lever, Personio, Recruitee, Teamtailor, Gem) by title, description keywords, location, remote, seniority, employment type, salary and visa. Newest first; schedule it for only new jobs.

- **URL**: https://apify.com/digital\_influx/job-search.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** Jobs, Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.00 / 1,000 job saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Job Search API: Jobs from Company Career Sites

Search open jobs **straight from company career sites**, the way LinkedIn and Indeed get them, but at the source. Type a job title, pick locations, seniority, workplace, employment type and dates, search inside the descriptions, and get the newest matching jobs from **20,000+ companies** that hire through Greenhouse, Ashby, Lever, Personio, Recruitee, Teamtailor and Gem.

- **Checked live on every run.** A daily index of every open job picks the companies that have a matching job, and the Actor reads those companies' public job boards live when you run it. No expired jobs, no reposts: if it is in the results, it is open right now.
- **Only new jobs, for schedules.** Turn on *Only new jobs* and schedule the search daily: each run returns just the postings earlier runs did not return, and you pay only for those.
- **Salary, seniority, visa and experience.** Pay ranges are read into numbers even when they are only written in the description, the level comes from the title, and the description tells whether the job sponsors visas and how many years of experience it asks for.
- **Direct from the employer.** Every job links to the company's own posting and application form. No recruiters, no aggregator duplicates.
- **No LinkedIn, no login, no proxies.** Only the public job-board APIs these platforms serve to the companies' own careers pages.

### Use it for

- **Job boards and newsletters:** "remote senior data engineer jobs posted this week" from thousands of companies, in one run, and only the new ones every morning.
- **Job seekers and career coaches:** a daily search that only returns new, real openings at the right level, with the salary.
- **Recruiters and sales teams:** which companies are hiring for the roles you place or sell to.
- **A daily feed of every new job:** leave the title empty, set *Posted within* to 1 day, *Max jobs* to 5,000 and *Only new jobs* on, and schedule it daily.
- **Market research:** hiring volume, seniority mix, remote share and salary ranges by role.
- **International talent:** keep the jobs that say they sponsor visas.
- **AI agents:** plain JSON with plain-text descriptions. Agents can call it through the Apify MCP server.

### Input

| Field | What it does |
|---|---|
| Job title keywords | Whole-word match on the title: `intern` does not match `internal`. |
| Exclude title keywords | e.g. `manager`. |
| Locations | `Berlin`, `Germany`, `DE`, `London`. A two-letter code matches as a whole word or the country code, so `US` does not match "Dusseldorf". `remote` also matches jobs flagged remote. |
| Remote only | Only fully remote jobs. |
| Seniority | Intern, entry, not stated (usually mid-level), senior, lead, director, executive. Read from the title, so it costs no extra time. |
| Posted within (days) | Exact dates from each platform. |
| Workplace type | Remote, hybrid, on-site, or not stated. Greenhouse jobs rarely state it (85% of the 195,000 Greenhouse jobs in our index), unless the location says Remote, Hybrid or On-site: add *Not stated* to keep them. |
| Employment type | Full-time, part-time, contract, temporary, internship, or not stated (Greenhouse lists none, so its jobs are *Not stated* unless the title says so, as in "Intern" or "(Contract)"). |
| Description keywords / Exclude description keywords | Whole-word match on the full description: `Kubernetes`, `GDPR`; exclude `clearance`. |
| Only jobs that state a salary | A pay range from the job board or written in the description. |
| Only jobs that offer visa sponsorship | The description says the company sponsors visas. Jobs that say nothing are dropped. |
| Platforms | Greenhouse, Lever, Ashby, Workable, Recruitee, Personio, Teamtailor, Gem (default: all). |
| Only these companies / Exclude companies | Match on the company or board name. |
| Max jobs | Newest first, default 200. |
| Max jobs per company | Stops one big employer from filling the results. |
| Only new jobs | Skip jobs that earlier runs of the same search already saved. |
| Memory name | Optional: share one "only new" memory between tasks, or start a fresh one. |
| Read every company live | Skip the daily index and read every company in the directory (a few minutes). |
| Job description | Plain text (default), HTML, both or none. Fetched only for the jobs you keep. |

The description filters (keywords, salary, visa) read the full text of every job that passes the other filters, so those runs take longer (a search over 12 companies and 3,730 open jobs took about a minute). You still pay only for the jobs saved.

```json
{
  "keywords": ["data engineer", "analytics engineer"],
  "locations": ["remote", "Germany"],
  "seniority": ["senior", "lead"],
  "employmentTypes": ["fulltime", "unspecified"],
  "descriptionKeywords": ["dbt", "Airflow"],
  "postedWithinDays": 7,
  "onlyNew": true,
  "maxJobs": 100
}
```

### Output

One item per job, the same schema as our ATS Jobs Scraper:

```json
{
  "id": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "platform": "ashby",
  "companySlug": "ramp",
  "company": "Ramp",
  "title": "Security Engineer, Cloud",
  "department": "Engineering",
  "location": "New York, NY (HQ)",
  "locations": ["New York, NY (HQ)", "Remote (US)"],
  "country": "US",
  "remote": true,
  "workplaceType": "hybrid",
  "employmentType": "Full time",
  "seniority": null,
  "postedAt": "2026-04-07T17:12:35.753Z",
  "url": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "applyUrl": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245/application",
  "salary": { "min": 211400, "max": 290600, "currency": "USD", "period": "year", "text": "$211.4K - $290.6K", "source": "platform" },
  "yearsOfExperience": null,
  "visaSponsorship": null,
  "descriptionText": "ABOUT RAMP\n\nRamp is building...",
  "scrapedAt": "2026-09-24T20:00:00.000Z"
}
```

- `salary.source` is `platform` when the job board publishes the pay, or `description` when it was read from the text (US pay-transparency ranges, "€50.000 – 60.000", "$45/hr"). Funding rounds, revenue and deal sizes are not taken for pay; when in doubt the field stays `null`. In our test on 469 real jobs, 382 came out with a salary, 262 of them from the description alone.
- `seniority`: `intern`, `entry`, `senior`, `lead` (lead, staff, principal), `director` (director, head of), `executive` (VP, C-level), or `null` when the title names no level.
- `payRanges` (Greenhouse jobs): the pay ranges the company set in Greenhouse, one per zone or level (`[{ "title": "Zone A", "min": 159000, "max": 254000, "currency": "USD" }]`), or `null` when it publishes none. They come from the job page, read with the description. When the text states no salary, `salary` is the first range.
- `yearsOfExperience`: the minimum the description asks for. `visaSponsorship`: `true` if it offers sponsorship, `false` if it says there is none, `null` if it does not say.
- Experience, visa and a salary written in the text come from the description. With *Job description: none* they are still filled for Ashby and Lever jobs (their text comes with the list) and for every job when a description filter is on; Greenhouse jobs otherwise leave them empty.

The run's **OUTPUT** record says how many company boards were read, how many jobs were open, how many matched, how many were skipped as already seen, and which boards could not be read.

### Pricing

Pay per event: **one `search` event per run** (it checks the index of every open job and reads the matching companies live) **plus one `job` event per job saved**. Filters cost nothing, and with *Only new jobs* the jobs you already have are not saved again, so they are not charged again. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

### Only new jobs: how it works

- The memory is a list of the jobs this search saved, kept in a key-value store named `job-search-state` in your own Apify account, for 180 days.
- Each set of filters has its own memory, so two different searches never hide each other's jobs. The date window and the caps are not part of it: widening *Posted within* keeps the same memory.
- Only jobs that were really saved are remembered. If a run stops at your spending limit, the rest come next time.
- To start over, give the search a new *Memory name*, or delete the record from the store.

### Good to know

- **Which companies?** The directory holds companies with at least one open job: on 2026-09-28, 20,950 companies (5,801 on Greenhouse, 3,565 on Personio, 3,396 on Ashby, 2,784 on Teamtailor, 2,566 on Recruitee, 1,886 on Lever, 952 on Gem) with 426,490 open jobs, many of them in Europe. Workable is supported, but its servers limit how fast boards can be checked, so it has no companies in the directory yet. It is built from public web crawl data and every board is checked against its live API; it is rebuilt every month, so new companies are added and closed boards removed. The run's OUTPUT record shows the date of the directory used.
- **Not every company in the world:** companies on Workday, iCIMS, Taleo and other systems are not in this search. For a list of companies you already know, including Workday, use our **ATS Jobs Scraper**, which takes careers page URLs.
- **How a search works.** It first runs your title, place, date, seniority, workplace and employment filters on a daily index of every open job in the directory, to find the companies that have a matching job. It then reads only those companies live, 20 at a time on each job board platform, and the results come from that live read. When more companies match than the search can keep (Max jobs, newest first), only the ones with the newest matching jobs are read. A company whose first matching job was posted after the index was built appears once the index includes it (it is rebuilt every day). *Read every company live* reads the whole directory instead, which takes a few minutes. A company that could not be read when the index was built is skipped until the next one. The OUTPUT record shows the date of the index and how many companies were skipped, and why. With description filters, Greenhouse jobs are judged on their own job pages, read newest first until *Max jobs* is reached. Lever asks for one request per second and the Actor keeps to it.
- It runs with 1 GB of memory, the minimum: the daily index of about 400,000 open jobs is loaded in memory, and Apify gives CPU in proportion to memory.
- *Job title keywords* match titles only, so "python" finds "Python Developer" but not a backend role that mentions Python only in the text. For that, use *Description keywords*.

### Support

A company missing from the directory, or a job that looks wrong? Open an issue on the Actor's Issues tab. Issues get an answer within a few days.

# Actor input Schema

## `keywords` (type: `array`):

Keep jobs whose title contains any of these words (whole-word match: 'intern' does not match 'internal'). Example: data engineer, analytics engineer.

## `excludeKeywords` (type: `array`):

Drop jobs whose title contains any of these words, e.g. senior, manager, intern.

## `locations` (type: `array`):

Keep jobs whose location or country code contains one of these texts: Berlin, Germany, DE, London. The word 'remote' also matches jobs flagged remote.

## `remoteOnly` (type: `boolean`):

Keep only jobs that offer a fully remote option.

## `seniority` (type: `array`):

Keep only these levels, read from the job title: "Senior Data Engineer" is senior, "Head of Sales" is director. "Not stated" keeps titles without a level word, which are mostly mid-level roles. Empty = all levels.

## `postedWithinDays` (type: `integer`):

Keep only jobs published in the last N days.

## `workplaceTypes` (type: `array`):

Keep only these workplace types, as the company set them on its job board. Greenhouse jobs rarely state it (85% of the 195,000 in our index), unless the location says Remote, Hybrid or On-site: add "Not stated" to keep them. Empty = all.

## `employmentTypes` (type: `array`):

Keep only these employment types. Ashby, Lever and Workable give the type; Greenhouse lists none, so its jobs count as "Not stated" unless the title says intern, part-time or (contract). Empty = all.

## `descriptionKeywords` (type: `array`):

Keep jobs whose description contains any of these words (whole-word match), e.g. Python, Kubernetes, GDPR. Description filters read the full text of every candidate job, so the run takes longer.

## `excludeDescriptionKeywords` (type: `array`):

Drop jobs whose description contains any of these words, e.g. clearance, relocation.

## `onlyWithSalary` (type: `boolean`):

Keep only jobs with a salary given by the job board or written in the description (for example a US pay-transparency range).

## `onlyVisaSponsorship` (type: `boolean`):

Keep only jobs whose description says the company sponsors visas. Jobs that say nothing about it are dropped.

## `platforms` (type: `array`):

Search only these applicant tracking systems. Empty = all.

## `companies` (type: `array`):

Search only companies whose name or board name contains one of these texts. Empty = all companies in the directory.

## `excludeCompanies` (type: `array`):

Skip companies whose name or board name contains one of these texts.

## `maxJobs` (type: `integer`):

Newest first. The search reads every matching company, then keeps this many.

## `maxJobsPerCompany` (type: `integer`):

Keeps one big employer from filling the results. 0 or empty = no limit.

## `onlyNew` (type: `boolean`):

Save only jobs that earlier runs of this same search did not save, so a daily or hourly schedule gives you just the new postings and you pay only for those. The memory is kept for 180 days in a key-value store named "job-search-state" in your account.

## `stateKey` (type: `string`):

By default each set of filters has its own memory. Give a name to share one memory between tasks, or a new name to start over. Letters, digits, dot, dash and underscore.

## `fullLiveScan` (type: `boolean`):

Off: a daily index of every open job picks the companies that have a matching job, and only those are read live, which makes a search much faster. A company whose first matching job appeared after the index was built is found once the index includes it (it is rebuilt every day). On: every company in the directory is read live, which takes a few minutes.

## `descriptionFormat` (type: `string`):

Fetched only for the jobs you keep. 'None' is fastest.

## `maxConcurrency` (type: `integer`):

1 to 40.

## Actor input object example

```json
{
  "keywords": [
    "data engineer"
  ],
  "remoteOnly": false,
  "postedWithinDays": 14,
  "onlyWithSalary": false,
  "onlyVisaSponsorship": false,
  "maxJobs": 100,
  "onlyNew": false,
  "fullLiveScan": false,
  "descriptionFormat": "text",
  "maxConcurrency": 20
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "data engineer"
    ],
    "postedWithinDays": 14,
    "maxJobs": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/job-search").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["data engineer"],
    "postedWithinDays": 14,
    "maxJobs": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/job-search").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "data engineer"
  ],
  "postedWithinDays": 14,
  "maxJobs": 100
}' |
apify call digital_influx/job-search --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/job-search"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XMN4lcZTo3kRgDdYn/builds/1R6tvlqyMouvh4x8J/openapi.json
