# Greenhouse.io Jobs Scraper (`hoholabs/greenhouse-scraper`) Actor

Every open job from any company's Greenhouse job board, with salary ranges, departments and offices. Filter by location, keyword or department. No API key required.

- **URL**: https://apify.com/hoholabs/greenhouse-scraper.md
- **Developed by:** [Hoho](https://apify.com/hoholabs) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Greenhouse.io Jobs Scraper

Get every open job from any company's [Greenhouse](https://www.greenhouse.com) job board: titles, locations, departments, offices, **salary ranges**, and optionally full descriptions. List as many companies as you like in one run, and filter by location, job title or department. No API key required.

***

### Why this scraper?

- **Thousands of companies hire through Greenhouse.** Discord, Airbnb, Figma, Databricks, Coinbase, Dropbox, GitLab, Robinhood, Stripe and many more. One run can cover your whole target list.
- **Salary ranges included.** When a company publishes pay (most US companies now do), you get `min`, `max` and `currency` per range, already converted from cents.
- **Filters that work.** Location, title keyword and department. Each is a case-insensitive "contains" match, so `New York` also matches `New York City, New York` and `Remote - New York`.
- **Built for scheduled runs.** A company that has left Greenhouse is skipped and logged; it doesn't fail the rest of your list.
- **Paste names or URLs.** `discord`, `Discord`, `https://job-boards.greenhouse.io/discord` and `boards.greenhouse.io/discord/jobs/123` all work.
- **Cheap:** $1 per 1,000 jobs.

***

### What you can fetch

| Mode | Description |
|------|-------------|
| `jobs` | Every open job at the companies you list, optionally filtered |

#### How to find a company's board name

Greenhouse is the hiring software behind many companies' careers pages, not a job site. Each company has its own board at `job-boards.greenhouse.io/<board-name>` (older boards use `boards.greenhouse.io/<board-name>`). The board name is usually the company name in lowercase (`discord`, `airbnb`, `figma`). If a company's careers page links to a Greenhouse address, the board name is the part after `greenhouse.io/`. Some companies show Greenhouse jobs on their own site (Stripe forwards to stripe.com/jobs), but their board name still works here.

***

### Usage

#### All jobs at one company

```json
{
  "companies": ["discord"]
}
```

#### Several companies, New York only

```json
{
  "companies": ["databricks", "figma", "airbnb", "coinbase"],
  "location": "New York"
}
```

#### Engineering roles, with full descriptions

```json
{
  "companies": ["gitlab", "dropbox"],
  "department": "engineering",
  "includeDescription": true
}
```

#### Remote sales roles

```json
{
  "companies": ["databricks", "robinhood"],
  "keyword": "sales",
  "location": "remote"
}
```

***

### Input fields

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `companies` | array of strings | `["discord"]` | Greenhouse board names or Greenhouse job-board URLs |
| `location` | string | — | Keep jobs whose location or office contains this text (e.g. `New York`, `London`, `Remote`) |
| `keyword` | string | — | Keep jobs whose title contains this text (e.g. `engineer`) |
| `department` | string | — | Keep jobs in a department whose name contains this text (e.g. `sales`) |
| `includeDescription` | boolean | `false` | Add the full job description as HTML (`description`) and plain text (`descriptionText`) |
| `limit` | integer | `0` | Max results: stop after this many jobs. `0` returns every matching job. Companies are fetched in the order you list them, and once enough jobs match, the rest aren't fetched. |

All text filters ignore case. You can combine them: a job must match every filter you set.

***

### Output fields

One dataset row per job:

| Field | Type | Description |
|-------|------|-------------|
| `id` | integer | Greenhouse job ID |
| `title` | string | Job title |
| `company` | string | The board name you asked for (e.g. `discord`) |
| `companyName` | string | Company display name (e.g. `Discord`) |
| `location` | string | Location as the company wrote it (e.g. `San Francisco Bay Area`) |
| `departments` | array | Department names (e.g. `["Policy"]`) |
| `offices` | array | Offices, each `{ name, location }` |
| `url` | string | Job posting / apply link |
| `firstPublished` | string | When the job was first posted (ISO 8601) |
| `updatedAt` | string | When the job was last updated (ISO 8601) |
| `requisitionId` | string | The company's internal requisition ID, when set |
| `applicationDeadline` | string | Application deadline, when set (usually empty) |
| `language` | string | Posting language code (e.g. `en`) |
| `salary` | array | Salary ranges, each `{ label, min, max, currency }`. Empty when the company doesn't publish pay |
| `description` | string | Full description as HTML (only with `includeDescription`) |
| `descriptionText` | string | Full description as plain text (only with `includeDescription`) |

Example row:

```json
{
  "id": 8806482002,
  "title": "Commercial Policy Lead",
  "company": "discord",
  "companyName": "Discord",
  "location": "San Francisco Bay Area",
  "departments": ["Policy"],
  "offices": [{ "name": "San Francisco, CA", "location": "San Francisco, California, United States" }],
  "url": "https://job-boards.greenhouse.io/discord/jobs/8806482002",
  "firstPublished": "2026-09-15T12:52:11-04:00",
  "updatedAt": "2026-09-15T12:52:11-04:00",
  "requisitionId": "R-107406",
  "applicationDeadline": null,
  "language": "en",
  "salary": []
}
```

> **Note:** these are the exact keys on each dataset row. In the Apify console's
> field list you'll see `jobId`, `jobTitle` and `jobDescription` instead. Apify's
> dataset-schema format reserves `id`, `title` and `description`, so those three
> are aliased for display only. The data itself uses `id` / `title` / `description`.

***

### Good to know

- **You name the companies.** Greenhouse has no search across all companies, so this actor can't answer "every Greenhouse job in New York". It gets every job at the companies you list, and the filters narrow that down.
- **Unknown companies are skipped.** If a board name doesn't exist, the run logs it and carries on with the rest. If *none* of the names exist, the run fails with a message listing them.
- **Salary is only as good as the posting.** Some companies publish several ranges per job (by US pay zone, for example); you get all of them, each with the company's own label.
- **Filters are plain text matches.** `NYC` won't match `New York`. Use the wording companies actually use.

***

### Use cases

- **Sales and recruiting signals:** watch target accounts and see when they start hiring for a team
- **Job boards and newsletters:** fill a niche board from the careers pages of hand-picked companies
- **Salary research:** collect published pay ranges by role, company and location
- **Job search:** track openings at the companies you want to work for, on a daily schedule

***

### Disclaimer

This actor is not affiliated with, endorsed by, or sponsored by Greenhouse Software, Inc. or any company whose jobs it returns. It reads the public Greenhouse Job Board API, which companies use to publish their openings. Use the data responsibly and in line with applicable laws.

# Actor input Schema

## `queryType` (type: `string`):

What to fetch: get open jobs for companies.

## `companies` (type: `array`):

Greenhouse board names, one per line — the part after greenhouse.io/ in the company's job-board URL (e.g. "discord" from job-boards.greenhouse.io/discord). Full Greenhouse URLs are accepted too. Companies without a Greenhouse board are skipped and logged.

## `location` (type: `string`):

Only keep jobs whose location or office contains this text (case-insensitive), e.g. "New York", "London" or "Remote". Leave blank for all locations.

## `keyword` (type: `string`):

Only keep jobs whose title contains this text (case-insensitive), e.g. "engineer". Leave blank for all titles.

## `department` (type: `string`):

Only keep jobs in a department whose name contains this text (case-insensitive), e.g. "sales". Leave blank for all departments.

## `includeDescription` (type: `boolean`):

Add the full job description (HTML and plain text) to every job. Makes each row much larger.

## `limit` (type: `integer`):

Stop after this many results. Leave at 0 to get all results.

## Actor input object example

```json
{
  "queryType": "jobs",
  "companies": [
    "discord"
  ],
  "includeDescription": false,
  "limit": 0
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("hoholabs/greenhouse-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("hoholabs/greenhouse-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call hoholabs/greenhouse-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hoholabs/greenhouse-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bzzaeXRw85nUwKZcX/builds/2DKbjbzMx6dHtg5K6/openapi.json
