LinkedIn, Workday, Greenhouse and ATS Jobs Scraper avatar

LinkedIn, Workday, Greenhouse and ATS Jobs Scraper

Pricing

from $1.40 / 1,000 job listeds

Go to Apify Store
LinkedIn, Workday, Greenhouse and ATS Jobs Scraper

LinkedIn, Workday, Greenhouse and ATS Jobs Scraper

Workday jobs scraper, Greenhouse jobs scraper, Lever jobs scraper and LinkedIn jobs in one careers page jobs scraper covering 15 sources. Schedule this company job listings scraper and it becomes a job listing feed: new, changed and closed postings only, with apply links.

Pricing

from $1.40 / 1,000 job listeds

Rating

0.0

(0)

Developer

Gerald Dobin

Gerald Dobin

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

2

Monthly active users

3 days ago

Last modified

Share

Workday, Greenhouse, Lever, Ashby and ATS Jobs Scraper

Give this Actor a list of companies and it returns every open job those companies are advertising, as clean rows in one schema. It reads each company's applicant tracking system through that system's own public job board endpoint, so the data is the same data the company publishes on its careers page: no rendering, no guessing, no stale copies. Fourteen systems are covered today: Workday, Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Personio, Teamtailor, Recruitee, Rippling, Jobvite, Breezy HR, BambooHR and iCIMS. You can point it at a board link such as https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite or https://jobs.lever.co/palantir, at a company's own careers page such as https://vercel.com/careers (the Actor loads it once and finds the board behind it), or at a short system:slug pair such as greenhouse:stripe. Titles, departments, locations, employment type, posting dates, salary ranges where the company publishes them, and the full job description in both HTML and plain text all come back in the same field names no matter which system the job came from.

LinkedIn is a fifteenth source, and it works the same way. Write linkedin:stripe in the same list and that company's LinkedIn job postings arrive as the same rows, with the same change tracking, alongside your Greenhouse and Workday boards. It is read through the public guest pages, one employer at a time. Read the LinkedIn section below before you use it: a LinkedIn slug is not a company, and the same job can arrive twice if you list a company's LinkedIn page and its ATS board together.

Who it is for

Job aggregators and job boards that need clean feeds. Point the Actor at a few thousand employers, schedule it daily, and load the rows straight into your index. Every row carries a stable jobId, so you can upsert on it and spot the jobs that disappeared since yesterday without diffing free text.

Recruiters and sourcers tracking target companies. Keep a watch list of the twenty companies your candidates care about and see the moment a role opens, with the apply link and the pay range already in the row instead of buried three clicks into a careers site.

Sales teams using hiring signals as intent data. A company that just posted four data engineering roles is buying a data stack. Filter on job title keywords, push the results into your CRM, and let your reps work accounts that are visibly investing rather than accounts that merely match a firmographic filter.

What you get

One dataset row per unique job. Here is a real row from a test run against Ramp's Ashby board, with the two long description fields shortened for readability:

{
"ats": "ashby",
"companySlug": "ramp",
"companyName": null,
"jobId": "ashby:ramp:34413f8d-26bf-4bbc-8ade-eb309a0e2245",
"nativeId": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
"title": "Security Engineer, Cloud",
"department": "Engineering",
"team": "Backend",
"location": "New York, NY (HQ)",
"city": "New York City",
"country": "United States",
"remote": false,
"workplaceType": "hybrid",
"employmentType": "Full-time",
"seniority": null,
"salaryMin": 211400,
"salaryMax": 290600,
"salaryCurrency": "USD",
"salaryPeriod": "year",
"descriptionHtml": "<div><p>Ramp is a financial operations platform ...</p></div>",
"descriptionText": "Ramp is a financial operations platform ...",
"postedAt": "2026-04-07T17:12:35.753Z",
"updatedAt": null,
"applyUrl": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245/application",
"sourceUrl": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
"scrapedAt": "2026-09-07T23:03:29.062Z",
"changeType": "salaryChanged",
"previous": {
"title": "Security Engineer, Cloud",
"department": "Engineering",
"location": "New York, NY (HQ)",
"salaryMin": 198000,
"salaryMax": 272000,
"salaryCurrency": "USD",
"salaryPeriod": "year",
"postedAt": "2026-04-07T17:12:35.753Z"
}
}

Every field is present on every row. A field is null when the applicant tracking system does not publish it, never missing and never an empty string. Dates are ISO 8601 in UTC. workplaceType is one of remote, hybrid, onsite or null, and remote always agrees with it. Salary fields are plain numbers plus a currency and a period, parsed out of whatever shape the board used.

How to use it

Inputs

InputWhat it does
companiesOne line per company: a board link, a careers page link, or a system:slug pair. Required. Workday boards are addressed by two names, so its pair carries both: workday:nvidia.wd5/NVIDIAExternalCareerSite.
keywordsKeep only jobs whose title contains one of these words. Case does not matter.
locationsKeep only jobs whose location, city or country contains one of these words.
remoteOnlyKeep only jobs the company marks as remote.
postedAfterKeep only jobs posted on or after this date, written as YYYY-MM-DD.
includeDescriptionOn by default. Turn it off for a faster, cheaper run when you only need titles and links; the two description columns stay in the output as nulls.
maxJobsPerCompanyStop after this many jobs per company. 0 means no limit. Worth setting on Workday boards, which often carry thousands of openings: the Actor stops paging as soon as it has this many. See the note below on how this interacts with filters.
trackChangesOn by default. Compare every job with the previous run on the same companies with the same settings, and fill in changeType and previous.
onlyChangesOff by default. Return only the jobs that moved, and pay only for those.
maxConcurrencyHow many companies to work on at once. Lower it if a board starts rate limiting.
proxyConfigurationWhich proxy to read LinkedIn through. Apify Proxy on automatic groups by default, which is what was measured. It is used for LinkedIn and nothing else: the fourteen applicant tracking systems are read directly, so a list with no LinkedIn source in it never spends a byte of proxy traffic.

How maxJobsPerCompany behaves when you also set a filter. You always get up to that many jobs that actually match. On the boards that publish the whole job in the listing, that costs nothing extra. On the boards that need a second request per job, some filters can only be judged once that second request has happened: a Workday listing carries no posting date at all, and an iCIMS job card may carry no location. When a filter of that kind is set, the Actor keeps reading further down the board until it has enough matching jobs, rather than taking the first few and then discarding most of them. That costs more requests than an unfiltered run of the same size, and it stops at 2000 jobs examined per company, so a filter that matches almost nothing on a very large board returns what it found by then rather than reading the whole board. Filtering on job title keywords never costs extra on any board, because every listing carries the title.

The job feed. With trackChanges on, which is the default, every row carries changeType and, where it applies, previous. Schedule the Actor against the same company list and you get a feed:

changeTypeWhat it means
newThis job was not on the board on the previous run. The first run labels every job new, because there is nothing yet to compare with.
closedThis job was on the board on the previous run and is not on it now. It has been filled or pulled. The row is built from what the last run stored, so it has no description, and its link will usually 404. That is the point of it.
salaryChangedThe pay range, currency or period moved. previous holds what it was. A pay move is always reported ahead of any other edit, so it is never hidden behind a reworded advert.
updatedThe title, department, team, location, workplace type, employment type, seniority, posting date, apply link or description text changed.
unchangedNothing a reader of the board would notice has moved.

Turn on onlyChanges and the unchanged jobs are dropped before you are charged for them, which is the whole point of it: a daily run over two thousand employers charges for the twenty roles that moved, not the forty thousand that did not. In that mode a company whose jobs did not move produces no rows at all, not even a note saying so, so a quiet morning can return an empty dataset. Companies whose boards could not be read are still reported, because silence from a broken board would otherwise look exactly like silence from a quiet one.

The comparison is kept in a key-value store called ats-jobs-scraper-state in your own account, one record set per company. Two things follow from that. Changing a filter, maxJobsPerCompany or includeDescription starts a fresh comparison for that combination of settings: two scheduled runs with different filters keep separate baselines, so the narrower one never reports everything outside its filter as closed. And closed is only ever reported for a company whose board this run read from end to end. If the board errored, if a listing on it could not be read, if maxJobsPerCompany was set, if a filter needed a second request per job, if the board's paging stopped for any reason other than reaching the end of the listing, or if the run hit its own limits, then nothing is reported closed for that company and the other companies in your list are unaffected. A job that is merely unseen is never called closed, and a job your own keyword or location filter removed is not unseen: every job the board listed counts as open, whether or not a row was delivered for it.

If that key-value store cannot be opened or read, change tracking gives way rather than the run. A run with trackChanges on still delivers every job, with no changeType on the rows and nothing reported closed, which is exactly what a run with tracking switched off delivers. A run with onlyChanges on does the opposite and stops without charging, because with no baseline every job looks new, and a feed asked for the handful that moved should never quietly return the whole list and bill for it. If only one company's baseline is unreadable, only that company is skipped, with a row saying so, and its stored baseline is left untouched for the next run.

If you would rather not use change tracking at all, turn trackChanges off and deduplicate on jobId on your side. The id is built from the ATS, the board slug and the system's own job id, so it stays the same across runs and across changes to the job title or location.

Calling it from the API

curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~ats-jobs-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"companies": ["greenhouse:stripe", "workday:ww.wd1/careers", "https://vercel.com/careers"],
"keywords": ["engineer"],
"includeDescription": false
}'

Add &format=csv to that URL for a spreadsheet instead of JSON. Results also stay in the run's dataset, so you can page through them later with the standard dataset endpoints.

MCP. Apify exposes its Actors over the Model Context Protocol, so an assistant such as Claude can call this scraper as a tool and answer questions like "which of these fifteen companies is hiring platform engineers in Berlin right now" without you writing any glue code.

Coverage and limits

Covered. Workday, Greenhouse (including EU boards and embedded job boards), Lever, Ashby, Workable, SmartRecruiters, Personio, Teamtailor, Recruitee, Rippling, Jobvite, Breezy HR, BambooHR and iCIMS. Between them these cover most startups and mid-market employers, a good share of large European ones, and with Workday and iCIMS a large part of the enterprise market too. LinkedIn is covered as a fifteenth source and has a section of its own below, because it behaves differently enough to be worth reading about before you use it.

Workday takes two names, not one. A Workday board lives at a tenant host and publishes one or more career sites on it, so both halves are needed. Paste the board URL in any shape it comes in, including the ones with /en-US/ or a /job/... path, and the Actor works out the pair. If you prefer the short form, write it as workday:tenant.wdN/SiteName, for example workday:nvidia.wd5/NVIDIAExternalCareerSite. Workday boards are often very large, so set maxJobsPerCompany when you only want a sample: the Actor stops paging as soon as it has enough instead of walking the whole board.

LinkedIn

How to address a company. Three forms work, and they all mean one employer:

What you writeWhat it is
linkedin:stripeThe slug in the company's LinkedIn page address.
linkedin:2135371The LinkedIn organization id, if you already know it.
https://www.linkedin.com/company/stripe/The company page itself.

A slug is not a company, and this is the one thing to get right. LinkedIn hands a slug to whoever registered it first, so the obvious spelling is often somebody else. Measured on 2026-09-25: linkedin:linear is organization 1353787, LINEAR GmbH, a German software company, while the issue tracker most people mean by that name is linkedin:linearapp, organization 29309454. Two things protect you. The Actor logs the organization id and the organization's real name for every LinkedIn entry before it delivers a single row, and every row carries that name in companyName and the organization id in companySlug. Check the first row. If the company name is not the company you meant, you have the wrong slug, and the fix is to open the company's LinkedIn page and read the slug out of the address bar. A company name on its own is never turned into a LinkedIn slug by guessing, because a guess here delivers a different company's jobs and bills you for them.

A slug that does not exist gets one row saying so, naming the slug you wrote. It is never a silent zero and it is never charged.

What a LinkedIn row carries. Title, company name, location, posting date and the apply link come off the listing itself. With includeDescription on, one further request per job adds descriptionHtml, descriptionText, seniority and employmentType. LinkedIn does not publish a department, a team or a salary range on a public posting, so department, team, salaryMin, salaryMax, salaryCurrency and salaryPeriod are null on every LinkedIn row. If you want pay ranges for a company, read its ATS board instead: that is where the range is published.

The same job can arrive twice, and that is deliberate. If you list both greenhouse:stripe and linkedin:stripe you get both sets of rows, and a role advertised in both places arrives twice with two different jobId values. Nothing is deduplicated across sources. The only way to match a job across two systems is to match on title and location, and doing that would silently drop a genuine second opening with the same title in the same city. Dropping a real job you are paying attention to is a worse mistake than handing you a duplicate you can see and filter, so the Actor does not guess. Pick one source per company, or dedupe on your own side where you know your data.

A thousand jobs per employer. LinkedIn's public listing stops at 1000 jobs per employer. A company advertising more than that is read down to 1000 and then marked as not fully read, which means nothing of theirs is ever reported closed. The same holds whenever a read is interrupted: a rate limit, a network failure, or an employer that lists nothing at all (a LinkedIn organization that does not exist answers exactly like one with no openings, so zero tells you nothing) all leave the read incomplete and suppress every closed row for that employer.

Every LinkedIn closed row is checked against the posting itself, and here is why. The public listing is not a stable set. One real employer was read four times over twenty minutes on 2026-09-25, each read walking the employer's whole range: the listing came back as 81, 91, 91 and 81 postings. Of the ten jobs that were in one read and not the next, every single one was still live on its own page. A scraper that trusted the comparison would have written ten closed rows and charged for all ten, and all ten would have been false.

So for LinkedIn, and only for LinkedIn, a job missing from a run's listing is treated as a question rather than an answer. Each one is looked up on its own posting page before anything is written. A page that is gone becomes a closed row. A page that still answers does not, and the job stays in the baseline so it is not billed to you again as new the next time the listing happens to show it. If hundreds go missing at once, that is a listing that was misread rather than an employer that closed its whole board, and no closures are reported for them at all.

The cost of that honesty is a slower run and a closure you may hear about a day late. You will not be billed for one that did not happen. On your ATS boards nothing changes: those endpoints are authoritative and stable, and closed there is reported straight from the comparison.

Rate and proxy. LinkedIn is read at about one request a second per company, through the proxy in proxyConfiguration, and at most four companies are read at once however high you set maxConcurrency. That is the rate that was measured as safe, and it is deliberately slower than the rest of the Actor. A company with 800 jobs takes a couple of minutes on its own. Apify Proxy datacenter groups, the default, were measured at zero blocks and roughly four times the speed of residential; switch to residential in that one input if you would rather.

Two things above cost extra requests, and both are spent only where they can change a row. Proving that a listing really ended costs up to twenty requests, and is skipped entirely on a first run, a capped run, or any run that could not report a closure anyway. Checking a missing job against its own posting page costs one request each, and only happens when there are missing jobs to check. A company of 90 jobs takes about 15 seconds to read on a first run and about 35 on a run that is comparing against a baseline.

Public pages only. Only what a logged out visitor is served is read: no login, no authentication bypass, no CAPTCHA solving, and nothing that harvests anybody's contact details out of a posting. There is no LinkedIn keyword search here either. This reads the jobs of employers you name, which is what makes it a feed rather than a scraper pointed at a search box.

Not covered. SuccessFactors, Taleo, Oracle Recruiting and JazzHR are not supported. Those systems either have no public job board endpoint or gate it behind per-tenant credentials.

A company must expose a public board. If an employer has switched its board to private, or publishes jobs only inside a rendered careers page with no underlying system, there is nothing to read and the Actor says so.

Descriptions. Greenhouse, Lever, Ashby, Teamtailor and Recruitee return the job body in the same call as the listing. Workday, SmartRecruiters, Workable, Rippling, BambooHR, Jobvite, iCIMS and Breezy need one extra request per job, so those companies take longer and cost more compute; turn includeDescription off when you do not need the body. LinkedIn needs one extra request per job too, and because LinkedIn is read at one request a second, that is where turning descriptions off makes the largest difference to how long a run takes. Personio returns descriptions only when the employer fills them in on the board.

Dates. postedAt and updatedAt are only filled in when the board publishes a date the Actor can read exactly, in ISO 8601 or that system's documented format. A date it cannot read becomes null rather than a guess. Workday is the clearest case: its board list dates a job only as "Posted Today" or "Posted 5 Days Ago", which become real dates, while "Posted 30+ Days Ago" is a floor rather than a date and becomes null instead of pretending to be exactly thirty days old. postedAfter is checked when the run starts, so a value the Actor cannot parse stops the run with a message instead of quietly disabling your filter.

Salary. Ashby, Lever, Recruitee and Breezy publish structured pay ranges and the Actor passes them through. BambooHR and Jobvite publish pay as a display string or a schema.org block, and the Actor reads it when the numbers and the currency are unambiguous; a range it cannot read confidently is left null rather than reported wrongly. Greenhouse, SmartRecruiters and Workday usually do not publish a structured range, so salaryMin and salaryMax are null there even when a range appears inside the description text.

Every entry in your list leaves one row, whichever way it goes. This matters most on a long list, where a company that quietly produced nothing is a company you would never notice. A board that is reachable and currently advertises no open jobs gets a row saying exactly that. So does a board whose jobs were all removed by your own keyword, location or date filters, and it tells you how many were removed. So does an entry naming a board that an earlier entry in the same list already named: it is read once, charged once, and the second entry's row points at the first. None of these rows are charged, because no job was delivered. A run of 84 employers returns at least 84 accounted outcomes, so you can reconcile the list you pasted against the rows you got by counting.

Every row carries the entry you pasted, in the input field, written exactly as you wrote it. If you give a careers page and the Actor finds the board behind it, the rows still carry your careers page URL alongside the board it resolved to, so a feed can be joined back to your own list without reverse engineering slugs.

Long lists and speed. 84 employers with full descriptions, about 9,400 jobs, take roughly eight minutes at the default concurrency of 10. Greenhouse and Lever serve every employer from one address each and neither rate limited 33 and 6 boards read at that concurrency. A single very large Workday board is what dominates a run of this shape: one 2,000 job board took four of those eight minutes on its own, and Workday is also the system most likely to answer with a rate limit while it is being paged. Set maxJobsPerCompany when a sample is enough, and lower maxConcurrency if a board starts refusing.

Companies that fail. A company that resolves to no known system produces a single row with error: "ATS not detected". A slug that no longer exists produces a row with error: "board not found". If a board answers with something that is not the shape it documents, you get error: "unexpected response from <system>" rather than a silent empty result, so a changed endpoint cannot look like a company that stopped hiring. These rows are free.

Addresses the Actor will not fetch. Only plain public http and https URLs on the default port are accepted. Links to an IP address, to localhost or a .internal name, or to any host that resolves onto a private, loopback, link local or cloud metadata address are refused, and that check is applied again on every redirect. Job descriptions are stripped down to ordinary formatting tags before they reach you, so no scripts, styles, embeds or javascript: links come through.

Run size. Up to 5000 companies and 200000 jobs in one run. Split larger lists across runs.

Pricing

Pay per unique job delivered. You are charged once for each job row the Actor writes into the dataset, and nothing else. There is no monthly subscription and no minimum.

  • Duplicate jobs inside a run are removed before they reach the dataset and are never charged.
  • Error rows are never charged.
  • A closed row is charged like any other row. It is a fact you asked for and it is the one a recruiter or a sales team acts on.
  • With onlyChanges on, the unchanged jobs are dropped before the charge, not after.
  • Jobs removed by your keywords, locations, remoteOnly or postedAfter filters never reach the dataset, so they are never charged. On the systems that need a second request per job, those filters are applied before that request, so a narrow filter is a cheap run.

A run that finds nothing costs nothing. If you set a maximum spend on the run, the Actor stops fetching once it hits that ceiling and finishes cleanly with whatever it has already delivered, rather than failing.

Support

Something wrong with a row, or a company whose board the Actor cannot see? Open an issue on the Actor's Issues tab with the exact companies entry you used and the run id. Requests for another applicant tracking system are welcome; say which one and name a company that uses it, so the endpoint can be checked against a real board.

Runs are stateless and store nothing beyond the dataset and the key value store record for the run itself. The Actor reads only public job board endpoints and sends no credentials anywhere.

What people search for

If you arrived looking for a Workday jobs scraper, this is it. The same Actor answers searches for: Workday jobs scraper, Greenhouse jobs scraper, Lever jobs scraper, Ashby jobs scraper, careers page jobs scraper, company job listings scraper, ATS jobs API, Rippling jobs scraper, BambooHR jobs scraper, iCIMS jobs scraper, Jobvite jobs scraper, SmartRecruiters jobs scraper, Workable jobs scraper, Personio jobs scraper, Teamtailor jobs scraper, Recruitee jobs scraper, Breezy jobs scraper.

Other searches this Actor answers: ATS jobs API, company job listings scraper, BambooHR jobs, Rippling jobs, iCIMS jobs scraper, Workday jobs API, Greenhouse job board API, Lever jobs API, Ashby jobs API, job postings API.