# Greenhouse Jobs Scraper & API (`miladamirzadeh/greenhouse-jobs-scraper`) Actor

Free Greenhouse jobs scraper: pay only for platform usage. Extract every open job from one or many Greenhouse boards with titles, locations, departments, salary ranges, remote/hybrid type, descriptions, publish dates, and application questions. Filter by department, location, keywords, or date.

- **URL**: https://apify.com/miladamirzadeh/greenhouse-jobs-scraper.md
- **Developed by:** [Milad Amirzadeh](https://apify.com/miladamirzadeh) (community)
- **Categories:** Jobs, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Greenhouse Jobs Scraper & API

Collect every open job from one or many Greenhouse job boards in a single run: titles, locations, departments, offices, salary ranges, remote/hybrid type, full descriptions, publish dates, and application questions. Use the results for job-board feeds, recruiting research, salary benchmarking, hiring-trend tracking, and scheduled employer snapshots, with no login or API key.

The Actor is free to use: you pay only for Apify platform usage, which is typically a fraction of a cent per run.

Greenhouse is the applicant tracking system behind the careers pages of companies such as Airbnb, Stripe, Figma, Anthropic, Databricks, Cloudflare, and thousands more. This Actor reads Greenhouse's public Job Board API directly over HTTP. It runs without a browser or proxy and finishes most companies in a few seconds.

### What you get

- All open jobs for one or many companies per run, from a slug (`stripe`) or any board URL
- Specific job postings by ID or URL (`job_details` mode)
- **Salary ranges** (min, max, currency) whenever the employer publishes pay transparency data
- **Workplace type** (Remote, Hybrid, On-site) detected from employer fields and location text
- Departments, offices, requisition ID, custom employer fields, and language
- Full job description as HTML **and** clean plain text
- First-published and last-updated timestamps
- Optional application form questions
- Filters for department, location, keywords, and publish or update date
- US and EU Greenhouse boards (`job-boards.greenhouse.io` and `job-boards.eu.greenhouse.io`)
- A per-company run summary showing open, matched, and saved counts, plus any slugs that were not found

### Quick start

#### All jobs from several companies

```json
{
  "mode": "company_jobs",
  "companySlugs": ["stripe", "figma", "https://job-boards.greenhouse.io/airbnb"],
  "maxJobsPerCompany": 9999
}
```

#### Filtered: recent remote engineering roles mentioning AI

```json
{
  "mode": "company_jobs",
  "companySlugs": ["stripe", "figma", "airbnb"],
  "filterDepartment": "Engineering",
  "filterLocation": "Remote",
  "jobKeywords": ["AI", "machine learning", "backend"],
  "firstPublishedAfter": "2026-08-01",
  "maxJobsPerCompany": 200
}
```

#### Specific jobs, with application questions

```json
{
  "mode": "job_details",
  "jobIds": [
    "airbnb/8187190",
    "https://job-boards.greenhouse.io/airbnb/jobs/8232207",
    "https://job-boards.eu.greenhouse.io/parloa/jobs/4925162101"
  ],
  "includeQuestions": true
}
```

#### How to find a company's slug

Open the company's Greenhouse board. The slug is the path segment right after `greenhouse.io/`:

| Board URL | Slug |
| --- | --- |
| `https://job-boards.greenhouse.io/stripe` | `stripe` |
| `https://boards.greenhouse.io/airbnb/jobs/8187190` | `airbnb` |
| `https://job-boards.eu.greenhouse.io/parloa` | `parloa` |
| `https://boards.greenhouse.io/embed/job_board?for=figma` | `figma` |

You can paste any of these URLs directly. The Actor extracts the slug for you.

### Input reference

| Field | Type | Default | Behavior |
| --- | --- | --- | --- |
| `mode` | string | `company_jobs` | `company_jobs` lists all open jobs for each company. `job_details` fetches specific jobs. |
| `companySlugs` | array | `["airbnb"]` | Board slugs or board URLs. Used in `company_jobs` mode. Duplicates are removed. |
| `jobIds` | array | `[]` | `slug/jobId`, `slug:jobId`, or Greenhouse job URLs. Used in `job_details` mode. |
| `includeContent` | boolean | `true` | Save the description as HTML (`content`) and plain text (`descriptionText`). |
| `includeQuestions` | boolean | `false` | Fetch application form questions. Adds one small request per job. |
| `filterDepartment` | string | empty | Case-insensitive partial match against department names. |
| `filterLocation` | string | empty | Case-insensitive partial match. Place names (`London`, `Germany`) match the job location and office names. Workplace terms (`Remote`, `Hybrid`, `On-site`) match the job location and the detected `workplaceType`, not office names, because employers often use labels like "Remote India" for office-based teams. |
| `jobKeywords` | array | `[]` | Up to 20 keywords or phrases. A job matches when **any** of them appears as a whole word in its title or description. |
| `firstPublishedAfter` | string | empty | `YYYY-MM-DD` or relative (`7 days`). Inclusive, compared as UTC dates. |
| `updatedAfter` | string | empty | Same format as above, applied to the last-updated timestamp. Useful for daily incremental runs. |
| `maxJobsPerCompany` | integer | `500` | Maximum jobs saved per company after filtering, newest first. Use `9999` for all. |
| `proxyConfiguration` | object | off | Optional. The API is public, so a proxy is not required. |

All filters that you set must match (AND). Filters apply only in `company_jobs` mode.

Keyword matching uses whole words, so `AI` matches "AI Engineer" and "Applied AI" but not "maintain" or "email". Add every form you want to catch, such as `engineer` and `engineering`.

### Output

Each saved job looks like this:

```json
{
  "jobId": 8187190,
  "title": "Senior Staff Software Engineer, Tech Foundations",
  "companyName": "Airbnb",
  "companySlug": "airbnb",
  "location": "Remote",
  "workplaceType": "Remote",
  "departments": [{ "id": 77, "name": "Software Engineering" }],
  "departmentNames": ["Software Engineering"],
  "offices": [{ "id": 63868, "name": "United States", "location": null }],
  "url": "https://job-boards.greenhouse.io/airbnb/jobs/8187190",
  "absoluteUrl": "https://careers.airbnb.com/positions/8187190?gh_jid=8187190",
  "salaryMin": 244000,
  "salaryMax": 305000,
  "salaryCurrency": "USD",
  "payRanges": [
    { "title": "Pay Range", "min": 244000, "max": 305000, "currency": "USD", "description": "Our job titles may span more than one career level..." }
  ],
  "content": "<div class=\"content-intro\"><p>Airbnb was born in 2007...</p></div>",
  "descriptionText": "Airbnb was born in 2007 when two hosts welcomed three guests...",
  "metadata": [{ "name": "Workplace Type", "value": "Remote" }],
  "firstPublished": "2026-09-18T17:33:00-04:00",
  "updatedAt": "2026-09-24T18:49:00-04:00",
  "applicationDeadline": null,
  "language": "en",
  "requisitionId": "ONE",
  "internalJobId": 3542815,
  "questions": null,
  "scrapedAt": "2026-09-26T11:21:13+00:00"
}
```

| Field | Meaning |
| --- | --- |
| `jobId` | Greenhouse job post ID. |
| `title` | Job title. |
| `companyName` / `companySlug` | Company name from Greenhouse and the board slug. |
| `location` | Location text shown on the posting. |
| `workplaceType` | `Remote`, `Hybrid`, `On-site`, or `null`. Detected from employer fields, then from location text. |
| `departments` / `departmentNames` | Departments as objects and as a simple list of names. |
| `offices` | Offices as `{id, name, location}`. |
| `url` | Hosted Greenhouse job page. |
| `absoluteUrl` | The job on the company's own careers site when it has one. |
| `salaryMin` / `salaryMax` / `salaryCurrency` | First published pay range, in currency units (not cents). `null` when not published. |
| `payRanges` | All published pay ranges, for example separate ranges per region. |
| `content` / `descriptionText` | Description as HTML and as plain text. `null` when `includeContent` is off. |
| `metadata` | Custom employer fields such as cost center or employment type. |
| `firstPublished` / `updatedAt` | ISO 8601 timestamps with the employer's time zone offset. |
| `applicationDeadline` | Deadline when the employer set one. |
| `language` / `requisitionId` / `internalJobId` | Posting language, employer requisition ID, and the Greenhouse job ID shared by all postings of the same role. |
| `questions` | Application questions as `{label, required, description, fields}` when `includeQuestions` is on. |
| `scrapedAt` | When the job was collected (UTC). |

Results appear in the default dataset. Use the **Job overview** view for compact metadata, salaries, and links, or the **Job descriptions** view for text.

#### Run summary

Each run also writes an `OUTPUT` record to the default key-value store. It shows every target's status (`ok`, `not_found`, or `error`) and its open, matched, and saved job counts, so you can spot mistyped slugs without reading the log. When no target succeeds, the run is marked as failed.

### API and exports

Run the Actor from your application with the standard Apify API:

```bash
curl -X POST \
  "https://api.apify.com/v2/acts/miladamirzadeh~greenhouse-jobs-scraper/run-sync-get-dataset-items?token=<APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{
    "companySlugs": ["stripe", "figma"],
    "filterLocation": "Remote",
    "maxJobsPerCompany": 50
  }'
```

You can also use the generated Python, JavaScript, CLI, OpenAPI, and MCP examples in the Actor's **API** tab. Datasets export to JSON, CSV, Excel, XML, and HTML.

### Recurring workflows

Save a tested input as an Apify Task and schedule it daily or weekly. For daily runs, set `updatedAfter` to `1 day` so each run saves only postings that are new or changed since the day before.

Typical uses:

- Job-board and aggregator ingestion
- Salary benchmarking across companies that publish pay ranges
- Tracking hiring volume by department or location
- Alerts for new roles at target companies, through Apify integrations such as Slack, Google Sheets, Make, or Zapier

Each run produces an independent dataset. The Actor does not keep history between runs.

### Practical limitations

- You need each company's Greenhouse slug. Greenhouse has no public directory of all boards.
- Only public job boards are supported. Internal or employee-only boards are not available through the public API.
- Salary fields are filled only when the employer publishes pay transparency data in Greenhouse. Some employers put pay only in the description text, and those jobs have `null` salary fields.
- `workplaceType` is best-effort. It uses employer metadata and location text and does not read the description.
- Keyword and location filters are literal text matches, not semantic search.
- Filters apply only in `company_jobs` mode.
- Closed jobs return `not_found` in `job_details` mode.

### Pricing

The Actor is **free**. There is no per-run or per-result fee; you pay only for the Apify platform usage (compute units) of your runs, which the Apify free plan's monthly credit covers for most users.

Runs are cheap because the Actor makes plain HTTP requests with no browser or proxy. It runs at 512 MB by default, and most companies finish in a few seconds. For example, the full Databricks and Anthropic boards (about 1,500 jobs) took under a minute at 512 MB, which is less than 0.01 compute units. Turning on `includeQuestions` adds one request per job and makes runs a little longer.

### Troubleshooting and support

If a run returns no jobs:

1. Open the `OUTPUT` record. A `not_found` status means the slug is wrong or the company does not use Greenhouse.
2. Open the company's board in a browser and copy the slug from the URL.
3. Remove the filters temporarily to rule out an overly narrow match.

For support, open an Actor issue with your input (without tokens), the run link, and what you expected to see.

# Actor input Schema

## `mode` (type: `string`):

<b>Company jobs</b> lists every open job on the given company boards. <b>Job details</b> fetches specific job postings by ID or URL.

## `companySlugs` (type: `array`):

Greenhouse board slugs such as <code>stripe</code> or <code>figma</code>, or full board URLs such as <code>https://job-boards.greenhouse.io/stripe</code>. The slug is the path segment right after <code>greenhouse.io/</code>. Used in <b>Company jobs</b> mode.

## `jobIds` (type: `array`):

Specific jobs as <code>slug/jobId</code> (for example <code>airbnb/7649441</code>) or full Greenhouse job URLs from boards.greenhouse.io or job-boards.greenhouse.io. Used in <b>Job details</b> mode.

## `includeContent` (type: `boolean`):

Save the full job description as HTML (<code>content</code>) and plain text (<code>descriptionText</code>). Turn off for smaller, faster exports.

## `includeQuestions` (type: `boolean`):

Fetch each job's application form questions. Adds one small request per job, so large boards take longer.

## `filterDepartment` (type: `string`):

Keep only jobs whose department name contains this text (case-insensitive), e.g. <code>Engineering</code>. Company jobs mode only.

## `filterLocation` (type: `string`):

Keep only jobs whose location contains this text (case-insensitive). Place names such as <code>London</code> also match office names. <code>Remote</code>, <code>Hybrid</code>, and <code>On-site</code> match the job location and detected workplace type.

## `jobKeywords` (type: `array`):

Keep jobs whose title or description contains <b>any</b> of these keywords (case-insensitive substring match). Up to 20 keywords.

## `firstPublishedAfter` (type: `string`):

Keep jobs first published on or after this UTC date. Jobs without a publish date are excluded when set.

## `updatedAfter` (type: `string`):

Keep jobs updated on or after this UTC date. Useful for daily incremental runs.

## `maxJobsPerCompany` (type: `integer`):

Maximum number of matching jobs saved per company, newest first. Set a high value such as 9999 to save every open role.

## `proxyConfiguration` (type: `object`):

Optional. The Greenhouse Job Board API is public, so a proxy is not required.

## Actor input object example

```json
{
  "mode": "company_jobs",
  "companySlugs": [
    "airbnb",
    "figma"
  ],
  "jobIds": [],
  "includeContent": true,
  "includeQuestions": false,
  "jobKeywords": [],
  "maxJobsPerCompany": 500,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `overview` (type: `string`):

No description

## `details` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companySlugs": [
        "airbnb",
        "figma"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("miladamirzadeh/greenhouse-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "companySlugs": [
        "airbnb",
        "figma",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("miladamirzadeh/greenhouse-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companySlugs": [
    "airbnb",
    "figma"
  ]
}' |
apify call miladamirzadeh/greenhouse-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,miladamirzadeh/greenhouse-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TJiyy3dMeXpoaUWt5/builds/QUczmDMqR7bSMls9r/openapi.json
