Greenhouse Jobs Scraper & API - Pay Once Per Job avatar

Greenhouse Jobs Scraper & API - Pay Once Per Job

Pricing

from $0.96 / 1,000 jobs

Go to Apify Store
Greenhouse Jobs Scraper & API - Pay Once Per Job

Greenhouse Jobs Scraper & API - Pay Once Per Job

Read every open role on a Greenhouse job board without an API key: title, department, office, location, posting date, apply link and the full description on request. Name a memory and later runs return only the jobs added since. JSON, CSV, Excel or API, billed once per new job.

Pricing

from $0.96 / 1,000 jobs

Rating

0.0

(0)

Developer

Automation Craft

Automation Craft

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

9 days ago

Last modified

Share

Greenhouse Jobs Scraper & API - Pay Once Per Job

This Greenhouse jobs scraper turns any Greenhouse careers page into clean JSON, CSV or Excel: title, department, team, location, remote flag, employment type, posting date, apply link and (optionally) the full description. Paste the careers URL, press Start, done. No login, no API keys, no proxies, no code.

Built to be complete and fair:

  • Deduplicated, billed once. Duplicates inside a run are removed. Give your search a memory name and every later run returns only the jobs that appeared since; jobs you already received are skipped and never charged again. Schedule it daily or weekly and you have a change feed.
  • Filters run before billing. Keywords (must contain / must not contain), locations, remote only, employment type and an exact posting-date window. You only pay for jobs that pass.
  • Exact caps, fair share. Per-company and per-run caps are honored exactly (the top competitor's 3-star review is "I set maximum 500 but got 200"). When several companies share a cap, each gets a fair share and unused budget flows to the companies with more jobs.
  • Honest about what you get. Every run ends with a run-summary item: jobs per board, how many were filtered or already known, exactly what was charged, and how often each field was populated.

The same engine also reads Lever, Ashby, Workable and SmartRecruiters boards, so you can mix careers pages from other ATS platforms into the same run. Sibling listings for those platforms are linked at the bottom.

30-second start

  1. Under Companies / careers pages, paste one careers URL per line, for example https://boards.greenhouse.io/stripe (any of these forms work: boards.greenhouse.io/
  2. Optional: tick Include full job descriptions, add keyword / location / date filters.
  3. Give it a Memory name if you will run it again and only want new jobs.
  4. Start. Download JSON, CSV or Excel from the Dataset tab, or read it through the API.

Nothing is charged until a job that passed your filters is delivered. Prices and worked examples are in the cost section below.

What you get from Greenhouse

Every job record has the same fields whatever the ATS. Where a platform does not publish a field you get null; the Actor never guesses. This table is the exact contract:

FieldGreenhouseLeverAshbyWorkableSmartRecruiters
title, url, applyUrl, jobIdyesyesyesyesyes
departmentsyesyesyesyesyes
teamnoyesyesyes (function)yes (function)
location and allLocationsyes (+ offices)yesyes (+ secondary)yesone per posting
remote flagonly when the location says Remoteyesyesyesyes
employmentTypenoyesyesyesyes
compensationnowhen publishedstructured rangesnono
publishedAtyesyesyesyesyes
updatedAtyesnononono
descriptionHtml and descriptionTextyesyesyesyesyes (one extra request per job)
companyNameyesno (token only)no (token only)yesyes

Measured fill rates from a live run at build time (share of delivered jobs where the field was populated):

Field populatedGreenhouseLeverAshbyWorkableSmartRecruiters
companyName100%0%0%100%100%
departments100%100%100%100%76%
team0%100%100%57%100%
location100%100%100%100%100%
remote8%100%100%100%100%
employmentType0%100%100%100%100%
compensation0%0%100%0%0%
publishedAt100%100%100%100%100%
description100%100%100%100%100%

Output

One item per job, plus status items when a company could not be read and one run-summary item at the end. Nested values (departments, allLocations, compensation) for developers; flat fields for spreadsheets.

FieldMeaning
provider, company, companyNameWhich ATS, the board token you gave, and the company name when the platform publishes it.
jobId, url, applyUrlStable id (always a string) and links.
title, departments, team, location, allLocations, countryThe role and where it is.
remote, workplaceType, employmentType, compensationWork arrangement and pay, exactly as the platform publishes them (null when it does not).
publishedAt, updatedAtISO timestamps from the platform.
descriptionHtml, descriptionTextFull description when requested.
changeType, isDuplicate, firstSeenAt, lastSeenAt, contentHash, scrapedAtDedup and provenance. changeType is NEW, UPDATED or DUPLICATE.

Sample item from a live run:

{
"type": "job",
"provider": "greenhouse",
"company": "stripe",
"companyName": "Stripe",
"jobId": "7532733",
"title": "Account Executive, AI Sales",
"departments": [
"1175 Enterprise - Account Executives (NA)"
],
"team": null,
"location": "San Francisco, CA",
"allLocations": [
"San Francisco, CA",
"US"
],
"country": null,
"remote": null,
"workplaceType": null,
"employmentType": null,
"compensation": null,
"publishedAt": "2026-02-03T20:19:01.000Z",
"updatedAt": "2026-08-25T21:40:40.000Z",
"url": "https://stripe.com/jobs/search?gh_jid=7532733",
"applyUrl": "https://stripe.com/jobs/search?gh_jid=7532733",
"descriptionHtml": "<h2>Who we are</h2>\n<h3>About Stripe</h3>\n<p>Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest ente...",
"descriptionText": "Who we are\n\nAbout Stripe\n\nStripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to acce...",
"changeType": "NEW",
"isDuplicate": false,
"firstSeenAt": "2026-08-31T11:22:17.996Z",
"lastSeenAt": "2026-08-31T11:22:17.996Z",
"contentHash": "0b4c36c4ac25360e363e35ff63e43d476c69377c",
"scrapedAt": "2026-08-31T11:22:17.996Z"
}

The run-summary item reports per board: jobsOnBoard, matchedFilters, knownFromMemory, deliveredNew, deliveredDuplicates, withDescription, plus totals, filteredOut counts, charges, fillRates and a plain-language hint whenever a run delivers nothing.

How much does it cost to scrape Greenhouse jobs?

Pay per event, no Actor start fee, no monthly minimum. Two events are charged:

EventCharged whenPrice per 1,000Silver (10% off)Gold (20% off)
Job (job-result)A new, unique job is delivered to the dataset$1.20$1.08$0.96
Job description (job-description)That job also arrives with its full description$0.50$0.45$0.40

A complete record (job plus description) therefore costs $1.70 per 1,000 jobs at the base price. The Silver and Gold columns are the Apify Store subscription discounts, applied automatically to both result events; there is no start event to discount.

Worked examples at the base price: 200 new jobs across 3 companies = $0.24. The same 200 with descriptions = $0.34. A weekly re-run that finds 12 new jobs among 200 known ones = $0.014. A re-run that finds nothing new = $0.00.

Always free: duplicates across runs (under the same memory name), jobs removed by your filters, unknown companies, empty boards, status items and the run-summary.

Input reference

InputWhat it does
Companies / careers pagesOne entry per company: careers URL, provider:token, or bare token. Case, http/https, www. and trailing paths do not matter. EU-hosted Greenhouse and Lever boards are detected from eu. URLs. SmartRecruiters identifiers are case-sensitive.
ATS for bare company namesWhich platform a plain token belongs to. Preset to Greenhouse in this listing. URLs always win.
Include full job descriptionsAdds descriptionHtml and descriptionText. SmartRecruiters needs one extra request per job for this, which is why it is priced separately.
Must contain / Must not contain / Fields to checkCase-insensitive keyword filters over title, departments, team, location and (Greenhouse, Lever, Ashby, Workable) the description text.
LocationsSubstring match over the job's location, every listed office and the country, e.g. London, Remote, United States.
Remote jobs onlyKeeps jobs the ATS explicitly marks remote. Greenhouse has no remote flag, so Greenhouse jobs match only when the location text says "Remote".
Employment typesSubstring match over the platform's employment type, e.g. Full, Contract, Intern. Greenhouse does not publish it.
Posted within / Posted after / Posted beforeExact window on publishedAt (the date the platform reports the job was first published).
Memory nameThe key that makes repeat runs return only new jobs. Stored in a named key-value store in your account (ats-jobs-memory-v2-<name>).
Also return already-known jobsRe-sends known jobs flagged isDuplicate: true with changeType DUPLICATE or UPDATED (the posting changed since you last received it). Free.
Reset this memory firstForgets everything under the memory name before running.
Max new jobs per company / totalExact caps on new, charged jobs. Known duplicates do not count towards them.

Using the API

Same input as the form.

curl -X POST "https://api.apify.com/v2/acts/automation_craft~greenhouse-jobs-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"companies": ["https://boards.greenhouse.io/stripe"],
"includeDescription": true,
"postedWithinDays": "30",
"dedupMemoryName": "greenhouse-watch",
"maxJobsTotal": 500
}'
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation_craft/greenhouse-jobs-scraper').call({
companies: ['https://boards.greenhouse.io/stripe'],
includeKeywords: ['engineer'],
keywordFields: ['title'],
dedupMemoryName: 'engineering-watch',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
const jobs = items.filter((i) => i.type === 'job');
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("automation_craft/greenhouse-jobs-scraper").call(run_input={
"companies": ["https://boards.greenhouse.io/stripe"], "remoteOnly": True, "maxJobsTotal": 200})
jobs = [i for i in client.dataset(run["defaultDatasetId"]).iterate_items() if i.get("type") == "job"]

Scheduling: create a Schedule in Apify (daily or weekly) with a memory name set; each run then delivers only the jobs that appeared since the previous run, and a run that finds nothing new costs nothing.

What this Actor does NOT do

  • It does not scrape job boards such as LinkedIn, Indeed or Naukri. It reads company career boards hosted on Greenhouse, Lever, Ashby, Workable and SmartRecruiters.
  • It does not discover companies for you. You supply the careers pages; it gets every open job they list.
  • It does not return fields the platform does not publish (see the coverage table), applicant data, or anything behind a login.
  • It cannot see unlisted or internal-only postings.

FAQ

Do I need a Greenhouse API key to read a job board?

No. The Actor reads the public board endpoint that a company's own Greenhouse careers page uses, so there is no key, no login and no cookie.

Where do I find a company's Greenhouse board token?

It is the last part of the careers URL, in lower case: stripe in https://boards.greenhouse.io/stripe. You can paste the full URL instead, and job-boards.greenhouse.io/<company> and EU boards on boards.eu.greenhouse.io/<company> are recognised too.

Does Greenhouse publish salary or employment type?

No. Greenhouse publishes neither, so compensation and employmentType are always null on Greenhouse jobs, and there is no remote flag either: a Greenhouse job counts as remote only when its location text says Remote.

How do I get the full job description?

Tick Include full job descriptions. Every new job then arrives with descriptionHtml and descriptionText, charged as the separate Job description event only for jobs that actually got one.

How do I see only the Greenhouse jobs posted since last week?

Set Posted within to "Last 7 days", or give an exact Posted after date. The window is applied to publishedAt before billing, so filtered-out jobs cost nothing.

Why does this Actor run with limited permissions?

Least privilege: it only needs to read public careers pages and write its own storages, so it runs with limited permissions and cannot touch the rest of your account. The cross-run memory is a named key-value store (ats-jobs-memory-v2-<name>) that the Actor creates and owns itself.

Fair use

This Actor reads the same public, unauthenticated job-board endpoints the companies' own careers pages use, at a polite request rate. It collects no personal data and never logs in. You are responsible for complying with applicable laws and the target sites' terms in your jurisdiction and use case.

Support

Something missing or wrong? Open an issue on the Actor's Issues tab. Fixes usually ship within a day.

More data tools by Automation Craft