Greenhouse Jobs Scraper - No API Key
Pricing
from $3.00 / 1,000 job scrapeds
Greenhouse Jobs Scraper - No API Key
Pull every open role from any Greenhouse job board by company token. Titles, departments, offices, direct apply links. No API key, no login, no proxy.
Pricing
from $3.00 / 1,000 job scrapeds
Rating
0.0
(0)
Developer
Renzo Madueno
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Greenhouse Jobs Scraper — Company Career Pages, No API Key
Pull every open role from any Greenhouse job board by company token. No API key, no login, no cookies, no proxy, no browser. You give it stripe, it gives you all 572 of Stripe's open postings with titles, departments, offices, full descriptions and a direct apply link for each one.
Greenhouse powers the careers page of a very large share of funded startups and mid-market tech companies. Every one of those boards is served by a public JSON endpoint that Greenhouse documents and keeps open. This Actor reads that endpoint properly, normalises the payload and hands you clean rows.
What you get, with the fill rate actually measured
We ran this Actor across 9 real boards — 2,628 live postings (stripe, anthropic, databricks, airbnb, coinbase, discord, robinhood, figma, dropbox) and counted how often each field came back populated. These are the real numbers, not a wish list:
| Field | Fill rate | Notes |
|---|---|---|
title | 100% | |
companyName | 100% | As Greenhouse has it, e.g. "Stripe" |
jobId | 100% | Greenhouse's public posting id |
internalJobId | 100% | Greenhouse's internal id |
location | 100% | The single location string on the posting |
locations[] | 100% | Every office attached to the posting |
department / departments[] | 100% | Requires content=true, which we always send |
postedAt | 100% | first_published, ISO 8601 |
updatedAt | 100% | ISO 8601 |
jobUrl | 100% | Canonical posting URL |
applyUrl | 100% | Deep link straight into the application form |
boardUrl | 100% | |
descriptionText | 100% | Full description, HTML stripped and entities decoded |
descriptionHtml | on request | Off by default; roughly triples dataset size |
isRemote | 100% | Always present as true/false — see the honesty note below |
requisitionId | 78.2% | Placeholder junk is nulled, see below |
workplaceType | 25.1% | Only when the location text actually says so |
salaryMin / salaryMax / salaryCurrency / salaryText | 57.0% | Parsed from the posting text, see below |
educationRequirement | ~98% | Greenhouse's education field |
applicationDeadline | rare | Almost always null; most boards do not set it |
Three fields this Actor deliberately does not return
The Greenhouse board API simply has no such data, on any board we measured. Rather than ship you three columns that are null on 100% of rows, they are absent from the schema:
employmentType— Greenhouse's public board payload has no full-time/part-time/contract field.country— there is no ISO country on the posting. Uselocationandlocations[].team— Greenhouse exposes a flat department list, not a department/team hierarchy.
If you need employment type or a normalised country, the Ashby Jobs Scraper and Lever Jobs Scraper in this same fleet both return them at ~99–100% fill, because those platforms actually publish them.
The two fields you should read carefully
Salary: 57%, and it is parsed, not published
Greenhouse's public board API exposes no structured pay field. We checked pay_input_ranges across all 2,628 postings on 9 boards: populated on zero of them. Anyone claiming a structured Greenhouse salary field is not reading the same API.
What we do instead is parse the posting text, because a lot of companies now write the range into the description to satisfy pay-transparency laws. That gets a range on 57.0% of postings. We then sanity-check every hit: of 1,207 extractions in a 2,245-row audit, 1,193 (98.8%) landed in a plausible annual band, the rest being genuine hourly or part-time roles.
Every salary row carries salarySource: "description" so you always know it was parsed rather than published. salaryText keeps the original snippet so you can verify it yourself. European formatting is handled: €92.300 — €130.000 is read as 92300–130000, not 92.3–130.
Remote: honest nulls instead of confident guesses
isRemote is always present. remoteSource tells you where the signal came from:
inferred_location— 27% of rows. The location text says so:"US-Remote","Poland - Remote OR Romania - Remote","Remote in the US","SF, NYC, remote". These are reliable.unknown— 73% of rows, returned asisRemote: false.
An earlier build also scanned the description body for the word "remote". We killed that path after measuring it: on 2,245 Stripe/Databricks/Airbnb postings it flagged 95 roles as remote whose location was Singapore, Dublin or Bengaluru. It was matching company boilerplate ("we are a remote-friendly company"), which describes the employer, not the job. A confident wrong true is worse than an honest false, so that inference is gone.
Filter on remoteSource === "inferred_location" if you only want rows where the board really said it.
Requisition IDs: junk removed
requisition_id comes back populated on 100% of postings, but 572 of 2,245 (25%) contained the literal placeholder text "See Opening ID". Those are nulled out. Real fill after cleaning: 78.2%.
Input
The only required field is the list of companies.
{"companies": ["stripe", "anthropic", "databricks"],"maxItems": 1000,"titleKeywords": ["engineer", "data"],"locationKeywords": ["New York", "Remote"],"remoteOnly": false,"includeDescription": true}
Accepted company formats
All of these resolve to the same board:
stripehttps://boards.greenhouse.io/stripehttps://job-boards.greenhouse.io/stripehttps://job-boards.greenhouse.io/stripe/jobs/8130725https://boards-api.greenhouse.io/v1/boards/stripe/jobshttps://anything.com/careers?for=stripe
You can also paste a newline- or comma-separated blob into a single string; it is split for you.
Input aliases
Different tools name things differently, so the Actor accepts synonyms and you never have to guess:
- companies:
companies,company,boards,boardTokens,companyUrls,startUrls - limit:
maxItems,maxResults,limit,maxJobs - keywords:
titleKeywords,keywords,searchTitle - locations:
locationKeywords,locations,location - departments:
departmentKeywords,departments,department - descriptions:
includeDescription,includeContent,fullDescription
Filters
| Option | What it does |
|---|---|
maxItems | Hard ceiling on rows written, across all boards. This is your spend cap. |
maxJobsPerCompany | Stops one 800-posting board from eating the whole budget |
titleKeywords | Keep only titles containing one of these |
locationKeywords | Match against location and every entry in locations[] |
departmentKeywords | Match against the department list |
remoteOnly | Keep only isRemote: true |
postedAfter | ISO date; drops anything published earlier |
includeDescription | Off makes the dataset far smaller |
includeHtmlDescription | Adds raw HTML; off by default |
dedupe | Drops repeat ids, and repeats of company+title+location |
concurrency | Boards fetched in parallel, 1–15, default 5 |
Output sample
{"source": "greenhouse","companyToken": "stripe","companyName": "Stripe","jobId": "8130725","internalJobId": 3520748,"requisitionId": null,"title": "Account Executive, AI Startups (Hunter)","department": "1653 Startups - Account Executives (NA)","departments": ["1653 Startups - Account Executives (NA)"],"location": "San Francisco","locations": ["US", "San Francisco"],"isRemote": false,"workplaceType": null,"remoteSource": "unknown","salaryMin": null,"salaryMax": null,"salarySource": null,"postedAt": "2026-08-19T14:02:07-04:00","updatedAt": "2026-08-19T14:02:07-04:00","jobUrl": "https://stripe.com/jobs/search?gh_jid=8130725","applyUrl": "https://job-boards.greenhouse.io/stripe/jobs/8130725#app","boardUrl": "https://job-boards.greenhouse.io/stripe","descriptionText": "Who we are About Stripe ...","scrapedAt": "2026-08-22T04:11:02.884Z"}
Four saved dataset views ship with the Actor: Job overview, Compensation, Remote & locations and Apply links, so you can export a clean CSV without picking columns by hand.
How errors are handled
Errors never enter your dataset. Charging you for a row that says "this failed" is charging you for a message. Everything that did not work goes into the FAILURES record in the run's key-value store:
{"runFailed": false,"boardsRequested": 3,"boardsWithJobs": 2,"jobsDelivered": 815,"duplicatesDropped": 0,"byReason": { "board_not_found": 1 },"entries": [{ "target": "notarealboard", "reason": "board_not_found","detail": "no Greenhouse board with this token (HTTP 404)","at": "2026-08-22T04:11:02.884Z" }]}
The reasons are deliberately distinct, because they mean different things:
board_not_found— HTTP 404. That company is not on Greenhouse, or the token is wrong. This is the normal outcome when you feed a list of companies, and it is not a bug.board_empty— HTTP 200 with an empty list. The board is real; the company just has nothing open right now.unresolvable_input— we could not derive a token from what you passed.budget_exhausted—maxItemswas reached before this board was written. Nothing is ever skipped silently.fetch_failed— network or upstream error after retries.
If nothing was delivered, the run ends FAILED. A run that quietly succeeds with an empty dataset is how you find out a week later that your pipeline has been dead. You get a clear failure and a reason.
Pricing and the free tier
Pay per event, and only for results:
- Actor start — one small charge per gigabyte of memory
- Job scraped — charged after the row is written to the dataset
Rows are pushed in batches and the per-job event is only charged for rows that actually landed. Nothing is billed for a failure, a filtered-out posting or a deduplicated repeat.
The free tier returns real data. There is no proxy requirement, no API key, no credential of any kind. Everything this Actor touches is a public JSON endpoint. A free Apify account running {"companies": ["stripe"], "maxItems": 10} gets ten real jobs. Nothing in the code path throws because you lack a paid feature.
Use maxItems as your budget control: it is a hard ceiling enforced before rows are written.
Speed and cost efficiency
Greenhouse serves one board per request, so a board of 800 postings costs exactly one HTTP call. In our measured run, 2,628 postings from 9 boards took 5.0 seconds. Requests are gzip-compressed, retried with exponential backoff on 429/5xx, and never retried on 404.
Note that content=true is not optional and we always send it. Without it Greenhouse omits content, departments and offices from the response entirely — that is why some scrapers return no department data at all. The cost is a bigger payload (Stripe's board is 360 KB without descriptions, 4.4 MB with) which gzip absorbs.
If you are crawling 50+ large boards in one run, raise memory to 2 GB.
Common questions
Do I need a Greenhouse API key? No. The board endpoint is public and unauthenticated. The Harvest API needs a key; this is not that.
How do I find a company's board token? It is the path segment in their Greenhouse URL. If their careers page is job-boards.greenhouse.io/acme, the token is acme. Many companies embed the board on their own domain — view the page source and look for boards-api.greenhouse.io/v1/boards/<token> or a gh_jid parameter.
Why does jobUrl sometimes point to the company's own site? Greenhouse returns absolute_url, which large companies redirect to their branded careers page. applyUrl always points at the Greenhouse application form itself, so use that one for automation.
Can I get application questions or the form schema? Not in this Actor. It returns postings.
Is there an EU board host? Yes, and boards.eu.greenhouse.io/<token> URLs are parsed correctly.
How fresh is the data? Live. Every run hits Greenhouse directly; nothing is cached.
Related Actors in this fleet
- Ashby Jobs Scraper — structured salary ranges on ~72% of postings, plus a real remote flag
- Lever Jobs Scraper — team, commitment, workplace type and country at ~99–100%
- Workday Jobs Scraper — the enterprise side, by career-site URL
- Startup Jobs Aggregator — give it company handles, it finds each on Greenhouse, Ashby or Lever and returns one deduplicated schema
Legal
This Actor reads a public, unauthenticated JSON endpoint that Greenhouse publishes so job boards and aggregators can syndicate postings. It sends no credentials, solves no challenges and bypasses no access control. You are responsible for how you use the data, including any applicable data-protection rules.