Greenhouse Jobs Scraper API & Hiring Monitor avatar

Greenhouse Jobs Scraper API & Hiring Monitor

Pricing

from $0.50 / 1,000 job results

Go to Apify Store
Greenhouse Jobs Scraper API & Hiring Monitor

Greenhouse Jobs Scraper API & Hiring Monitor

Greenhouse jobs scraper API for complete live company boards. Export jobs with salary, remote, seniority and country fields, or monitor verified new, updated, closed and reopened vacancies. Uses the official public Job Board API; no login, browser or proxy.

Pricing

from $0.50 / 1,000 job results

Rating

0.0

(0)

Developer

Vadim Bezrukov

Vadim Bezrukov

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Scrape every live vacancy from any public Greenhouse board through the official Job Board API. Export jobs with flat salary, remote, seniority and country fields, or monitor a company watchlist for verified NEW, UPDATED, CLOSED and REOPENED vacancies. No login, browser, CAPTCHA or proxy is required.

Every run reads Greenhouse's official Job Board API at the moment you run it. There is no cached jobs database behind this Actor: if an employer removes a role this morning, it will not remain in your next snapshot. No login, browser, CAPTCHA or proxy is required.

Choose your workflow

GoalSettingsWhat you get
Export a live job board to CSV, JSON or Excelsnapshot + allOne clean row per current vacancy, plus a company summary
Find targeted roles across companiessnapshot + filtersOnly matching jobs, with salary, seniority, remote type and apply links
Track competitor or portfolio hiringmonitor + changesOnlyOnly jobs that moved, plus company, department and location deltas
Feed a CRM, dashboard or AI agentAny mode + schedule/webhookStable IDs, explicit statuses and machine-ready change records

Why teams choose this Actor

One request per company, whatever its size. Greenhouse returns a complete board in a single response, descriptions included. In the September 2026 release benchmark, an 853-vacancy board required one HTTP request. That is why descriptions cost no extra requests here and large boards stay fast.

Predictable result-linked pricing. Snapshot exports charge for delivered job rows. Monitoring charges for successfully verified companies plus only the changed job rows when you use changesOnly. Failed, partial and missing boards do not become billable results.

It never hands you a truncated board. Greenhouse reports how many jobs it is returning. If that count does not match the jobs actually delivered, the run reports FAILED for that company instead of writing a short list that looks complete. A silently truncated export is the worst failure mode for anything you run on a schedule, and it is the one this Actor is built to make impossible.

Flat columns, not nested JSON. salary_min, salary_max, salary_currency, salary_period, remote_type, seniority, location_cities, location_countries and days_since_published are derived on every row, so the output drops straight into a spreadsheet, a CRM or an agent prompt.

A real change monitor, not a diff you compute yourself. Switch mode to monitor and later runs return only NEW, UPDATED, CLOSED and REOPENED jobs, with the exact fields that changed and a per-company hiring delta.

Quick start: inspect one company

Paste board tokens, one per line. The token is the last part of a Greenhouse board URL: job-boards.greenhouse.io/stripe means the token is stripe. Full URLs work too.

If you do not know the token, open any vacancy on the employer's careers page and copy its boards.greenhouse.io/... or job-boards.greenhouse.io/... URL. You may paste the board URL, an individual job URL or just the first path segment. The Actor normalizes all three forms, so users do not need to guess or maintain a separate company-name resolver.

{
"boards": ["figma"],
"mode": "snapshot",
"outputMode": "all",
"includePayTransparency": true
}

That returns every current Figma vacancy plus one company summary. In the September 2026 release benchmark the same input returned 161 job rows. Once the output fits your workflow, add up to 100 companies in the same run.

The Actor also includes ready-made Task configurations for a full CSV/JSON export, competitor monitoring, remote senior engineering jobs with salary, and cross-company hiring comparison. Replace the sample board tokens with your own watchlist and save the Task before scheduling it.

Filtering

Ask for the roles you actually want instead of exporting everything and filtering afterwards. All filters combine with AND.

{
"boards": ["stripe", "anthropic", "databricks", "figma"],
"filters": {
"titleIncludes": ["engineer", "developer"],
"titleExcludes": ["intern", "manager"],
"seniorities": ["senior", "staff", "principal"],
"remoteTypes": ["remote"],
"countries": ["US", "GB"],
"withSalaryOnly": true,
"publishedAfter": "2026-08-01"
}
}
FilterMatches
titleIncludes / titleExcludesCase-insensitive substrings of the job title
locationIncludesThe raw location string or any parsed city
departmentsDepartment names exactly as the board publishes them
countriesISO 3166-1 alpha-2 codes parsed out of the location
remoteTypesremote, hybrid, onsite
senioritiesintern, entry, junior, senior, staff, principal, director, vp, executive
withSalaryOnlyOnly jobs that published a pay range
publishedAfterISO date or timestamp

Filters decide which rows are written, never what is observed. The monitoring snapshot and the company summary always describe the whole board, so narrowing a filter can never make a job look CLOSED and widening one can never make an old job look NEW. matched_jobs on the company summary tells you how many rows a run wrote; active_jobs always tells you the size of the board.

A derived value that is unknown is excluded rather than assumed. A job whose wording gives no remote signal has remote_type: null and will not appear under remoteTypes: ["onsite"].

Output

Each job row carries the source fields plus the derived columns:

FieldExampleNotes
title, company_name, job_idSenior Software Engineer
apply_urlhttps://job-boards.greenhouse.io/stripe/jobs/1234567The employer's own apply link
locationSan Francisco, CA | New York City, NYRaw, exactly as published
locations[{city: "San Francisco", region: "CA", country: "US"}, …]Every office, parsed
location_cities, location_countries["San Francisco"], ["US"]Flat, for filtering and grouping
remote_typeremotenull when the posting gives no signal
seniorityseniorFrom level words in the title; null when absent
employment_typefull_timeUsually null: Greenhouse titles rarely state it
salary_min, salary_max222800.0, 290000.0Currency units, not cents
salary_currency, salary_periodUSD, yearPeriod read from the range title
first_published, days_since_published2026-08-01T…, 29
departments, offices[{id, name}]As published
description_html, description_textSet includeDescription: true
change_type, changed_fieldsUPDATED, ["title", "pay_ranges"]Monitor mode

Boards write locations in several conventions at once.

San Francisco, CA | New York City, NY | Seattle, WA
, Atlanta, Georgia, Remote - California and San Francisco, CA • New York, NY • United States all parse into structured offices with an ISO country code. In the September 2026 release benchmark across five live boards and 1,963 vacancies, a country resolved for 99% of jobs that stated a location.

seniority, remote_type and employment_type are derived from wording with fixed, published rules rather than a model, so the same input always gives the same answer. "Manager" and "Lead" are deliberately not treated as seniority levels: on real boards "Customer Success Manager" and "APAC Tax Lead" are job functions, and reading them as levels both mislabels individual contributors and hides the level in "Senior Customer Success Manager".

Every record also carries source, a stable source_id (greenhouse:{board_token}:{job_id}), source_url, UTC scraped_at, schema_version and a SHA-256 fingerprint over the semantic fields. Derived columns are excluded from the fingerprint, so refining a heuristic here can never announce thousands of fake UPDATED jobs.

Statuses

StatusMeaningMonitoring state advanced?
SUCCESSComplete verified board responseYes, in monitor mode
NOT_FOUNDGreenhouse returned HTTP 404 for that boardNo
PARTIALA suspicious drop could not be consistently verifiedNo
FAILEDTransport error, exhausted 429/5xx retries, malformed JSON, or a job count that did not matchNo

A genuinely empty board is SUCCESS with active_jobs: 0. It is never confused with NOT_FOUND or FAILED.

Monitoring changes

{
"boards": ["stripe", "airbnb"],
"boardAliases": { "stripe": "account-1041", "airbnb": "account-2277" },
"mode": "monitor",
"outputMode": "changesOnly"
}

The first run records the full board as a baseline. Every later run returns only what moved:

  • NEW - a job that was not on the board last time
  • UPDATED - with changed_fields naming exactly what changed
  • CLOSED - gone from the board, rebuilt from stored state
  • REOPENED - a previously closed job that came back

Plus one company_summary per company with active_jobs_delta, new_jobs, closed_jobs and per-department and per-location deltas, ready to drop into a CRM trigger, a dashboard or an alert.

boardAliases maps a company to your own identifier. It is echoed into every result as external_id and is never sent to Greenhouse.

A failed, partial or rate-limited check never manufactures CLOSED jobs and never replaces the last known-good state, so a bad network minute cannot tell you that a company stopped hiring.

Pricing

The live Pricing tab is authoritative for the run you are about to start. Significant pricing changes have a notice period, so the rate currently charged can temporarily differ from the next configured model shown below. During that transition you pay the rate displayed on the live Pricing tab.

EventPriceCharged when
job-result$0.0008 Free/Bronze; $0.00065 Silver; $0.0005 Gold+Each job row written to the dataset
board-check$0.02A company returned a complete SUCCESS response in monitor mode only

Retries, duplicate aliases, NOT_FOUND, PARTIAL and FAILED checks are never charged. outputMode: all, the default, writes and charges every current job. Setting outputMode: changesOnly, which is the point of a monitoring schedule, writes and charges only the jobs that moved: UNCHANGED jobs are neither written nor charged. Company summaries and board-status rows are always free.

What that means in practice:

RunCost
Full export, the 853-job release benchmark$0.68
Full export, 10 boards averaging 200 jobs$1.60
Monitoring 100 companies, a quiet day$2.00
Monitoring 100 companies, 40 jobs moved$2.03
Daily monitoring of 100 companies, 30 days$60 plus baseline and changed job rows

Snapshot runs have no fixed per-company charge, so a small board is not made artificially expensive. The second monitoring run is where this Actor becomes especially cheap: a watchlist that reports "nothing changed" costs two cents per company, because you are billed for the verified check, not for re-exporting a board you already have.

Apify's synthetic apify-actor-start event adds $0.00005 per run at the default 512 MB. Platform usage is included in these event prices rather than billed separately. You can cap a run with maxTotalChargeUsd: once the remaining budget cannot pay for the next row, the Actor stops writing, reports what it skipped in RUN_SUMMARY, and does not advance the monitoring snapshot, so nothing that was dropped is missing from the next run's changes.

Input reference

FieldDefaultPurpose
boardsRequired1 to 100 Greenhouse board tokens or board URLs
modesnapshotsnapshot returns the board and touches no state; monitor compares with the last successful run
outputModeallall returns every job; changesOnly omits UNCHANGED jobs
filtersnoneNarrows which job rows are written. See above
boardAliasesnoneMap of board token to your own ID, echoed as external_id
includeDescriptionfalseAdds description HTML and text, and includes them in change detection. No extra requests
includePayTransparencytrueAdds published pay ranges and the flat salary columns. No extra requests

Scheduling

Save your watchlist as an Apify Task and pick a daily schedule in Console, or create one through the API. Keep the input JSON inside runInput.body as a string.

curl -X POST "https://api.apify.com/v2/schedules" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d @- <<'JSON'
{
"name": "daily-greenhouse-watchlist",
"isEnabled": true,
"cronExpression": "0 7 * * *",
"timezone": "UTC",
"actions": [{
"type": "RUN_ACTOR",
"actorId": "boyPuAVGr92b4EujA",
"runInput": {
"contentType": "application/json; charset=utf-8",
"body": "{\"boards\":[\"stripe\",\"airbnb\"],\"mode\":\"monitor\",\"outputMode\":\"changesOnly\"}"
},
"runOptions": { "build": "latest", "memoryMbytes": 512 }
}]
}
JSON

Webhooks

Push successful-run metadata, including the Dataset ID, to your own endpoint. The receiver fetches the change rows from that dataset.

curl -X POST "https://api.apify.com/v2/webhooks" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d @- <<'JSON'
{
"eventTypes": ["ACTOR.RUN.SUCCEEDED"],
"condition": { "actorId": "boyPuAVGr92b4EujA" },
"requestUrl": "https://example.com/hooks/hiring-signals",
"shouldInterpolateStrings": true,
"payloadTemplate": "{\"runId\":\"{{resource.id}}\",\"datasetId\":\"{{resource.defaultDatasetId}}\",\"itemsUrl\":\"https://api.apify.com/v2/datasets/{{resource.defaultDatasetId}}/items?clean=true\"}"
}
JSON

Run it from your own code

curl -X POST \
"https://api.apify.com/v2/acts/automa-flow~greenhouse-jobs-hiring-monitor/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"boards":["stripe","anthropic"],"filters":{"seniorities":["senior","staff"],"withSalaryOnly":true}}'

Use it from an AI agent through MCP

Apify's hosted MCP server can expose this public Actor as a tool with its input and output schemas. Add this server URL to an MCP-compatible client and sign in through Apify when prompted:

https://mcp.apify.com?tools=automa-flow/greenhouse-jobs-hiring-monitor

A concrete prompt is more reliable than asking an agent to invent the input:

Run automa-flow/greenhouse-jobs-hiring-monitor for the Stripe, Anthropic and
Figma Greenhouse boards. Return remote senior or staff roles with published
salary ranges, then group the results by company and country.

For recurring use, tell the agent to keep mode: monitor and outputMode: changesOnly, and to inspect company summaries separately from job rows. The explicit status and record_type fields let it distinguish a valid empty result from a failed source check.

Reliability and source health

The client uses bounded exponential backoff with jitter for 408, 429, selected 5xx responses, timeouts and connection resets, plus a run-wide retry budget. Concurrency is fixed at three and is not exposed as a Store knob.

One company failing never affects the others: 100 companies in means 100 answers out, each with its own status.

When a board falls sharply from its last successful job count, the Actor makes an independent verification request. It accepts closures only if both complete responses contain the same job IDs. Otherwise it emits PARTIAL, keeps the known-good state and produces no closure events.

Known limitations

  • Only public Greenhouse Job Board API data is supported. Private or internal boards and application submission are out of scope.
  • pay_transparency=true is documented by Greenhouse for individual jobs and was live-verified on the list endpoint. If Greenhouse removes list-level support, pay ranges may be empty until the Actor is reassessed; it will not silently add one request per job.
  • includeDescription=false minimizes output and stored state, but the source call still uses content=true because the departments and offices needed for hiring deltas are embedded with content.
  • employment_type is usually null. Greenhouse job titles rarely state it and the Actor does not guess.
  • Compact tombstones are retained in the named monitoring key-value store to detect REOPENED jobs. Delete that store to reset all baselines.
  • A board token that Greenhouse no longer serves is NOT_FOUND; the Actor does not search the web to guess a replacement.
  • Apify-container reachability is a release smoke gate. Direct local requests require no proxy; a future source change that requires CAPTCHA, browser or residential stealth triggers reassessment rather than an automatic bypass.

Source attribution and responsible use

Data comes from Greenhouse's official public Job Board API and remains subject to Greenhouse's and each hiring company's terms. This Actor collects public job and company data only. It does not request application questions, applicant records, demographic form answers or candidate information.

Run output is stored in the requesting user's Apify Dataset. In monitor mode, only compact last-successful job state and tombstones are retained in the requesting user's named key-value store. The Actor is not affiliated with Greenhouse Software.