# Greenhouse Jobs Scraper | Pay Only for New Jobs (`khassinx/greenhouse-jobs-scraper`) Actor

Scrape job listings from 313 companies hiring now on public Greenhouse job boards. Deduplicated across runs: you are charged only for postings you have not received before, so a daily schedule costs a fraction of a full re-scrape. Filter by company or by tag, or pass your own target list.

- **URL**: https://apify.com/khassinx/greenhouse-jobs-scraper.md
- **Developed by:** [KHASSINX LLC](https://apify.com/khassinx) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.30 / 1,000 new job postings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Greenhouse Jobs Scraper | Pay Only for New Jobs

Job listings from **313 companies** hiring now on public **Greenhouse** job boards.

### Pay only for new jobs

Most job scrapers charge you for every posting they return, every run. This one keeps
track of what it already delivered, so a run that finds nothing new costs almost nothing.
Schedule it daily and you pay for the change, not for the whole board.

### Input

| field | what it does |
|---|---|
| `catalog_tags` | keep only companies with these tags (`tech`, `fintech`, `health`, …) |
| `catalog_limit` | cap how many companies to read |
| `targets` | your own list of `{"source": "greenhouse", "company": "<slug>"}` — overrides the catalog |
| `max_postings` | ceiling per run. Defaults to 2500 so a first run never surprises you with a bill |

Leave everything empty and it reads the full Greenhouse catalog, up to the cap.

### Output

One record per posting: company, title, location, department, URL, and the date it was
first seen. Straight to a dataset — export as JSON, CSV or Excel, or read it from the API.

### Only public boards, only what the terms allow

This actor reads the job boards companies publish on purpose. Boards whose terms forbid
automated access are not supported and will not be added — Workday and SmartRecruiters
were both evaluated and left out for that reason.

# Actor input Schema

## `max_postings` (type: `integer`):

Safety cap on how many postings a single run delivers — and therefore what it costs you. Defaults to 2500 (about $5.00). Nothing is lost when the cap is hit: the next run picks up exactly where this one stopped, and you are never charged twice for the same posting.

## `catalog_tags` (type: `array`):

Only fetch companies matching these tags. Leave empty to fetch the whole catalog. Examples: tech, fintech, ai, data, design, saas, devtools, infra, consumer, retail, health, education, media, gaming, remote.

## `catalog_sources` (type: `array`):

Which ATS to read. Defaults to Greenhouse. Leave as is unless you also want postings from other supported boards.

## `catalog_limit` (type: `integer`):

Caps how many companies are queried in a single run. Leave empty for no cap.

## `targets` (type: `array`):

Query exact companies instead of the built-in catalog. Each item: { "source": "lever" | "greenhouse" | "ashby", "company": "board-token" }. When set, the catalog filters above are ignored.

## Actor input object example

```json
{
  "max_postings": 2500,
  "catalog_tags": [
    "fintech"
  ],
  "catalog_sources": [
    "greenhouse"
  ],
  "catalog_limit": 5,
  "targets": [
    {
      "source": "greenhouse",
      "company": "stripe"
    },
    {
      "source": "lever",
      "company": "palantir"
    }
  ]
}
```

# Actor output Schema

## `postings` (type: `string`):

Every job posting found, normalized to one shape regardless of which ATS it came from. Postings already delivered in a previous run are not included and are not charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "max_postings": 2500,
    "catalog_tags": [
        "fintech"
    ],
    "catalog_sources": [
        "greenhouse"
    ],
    "catalog_limit": 5,
    "targets": [
        {
            "source": "greenhouse",
            "company": "stripe"
        },
        {
            "source": "lever",
            "company": "palantir"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("khassinx/greenhouse-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "max_postings": 2500,
    "catalog_tags": ["fintech"],
    "catalog_sources": ["greenhouse"],
    "catalog_limit": 5,
    "targets": [
        {
            "source": "greenhouse",
            "company": "stripe",
        },
        {
            "source": "lever",
            "company": "palantir",
        },
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("khassinx/greenhouse-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "max_postings": 2500,
  "catalog_tags": [
    "fintech"
  ],
  "catalog_sources": [
    "greenhouse"
  ],
  "catalog_limit": 5,
  "targets": [
    {
      "source": "greenhouse",
      "company": "stripe"
    },
    {
      "source": "lever",
      "company": "palantir"
    }
  ]
}' |
apify call khassinx/greenhouse-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,khassinx/greenhouse-jobs-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QrvGAXPVHgnB4T3nN/builds/7RsTrap63f87t4ZNW/openapi.json
