# Job Scraper - Workday, iCIMS, Greenhouse & Ashby Jobs API (`parseforge/career-site-jobs-scraper`) Actor

Reads job postings live from company career sites on Workday, iCIMS, Cornerstone, Dayforce, Greenhouse, Ashby, Lever, Comeet and 11 more ATS, plus 14 job feeds and any RSS URL. Geocoded location, salary, skills, and an incremental mode that returns only what changed.

- **URL**: https://apify.com/parseforge/career-site-jobs-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.68 / 1,000 job postings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Job Scraper - Workday, iCIMS, Greenhouse & Ashby Jobs API

**Read open roles straight off a company's own career site, live, across twenty-one applicant tracking systems, fourteen cross-company job feeds, 305 built-in company boards and an index of 168,000 postings.** Every row carries the title, employer, location, work arrangement, employment type, seniority, posting date and apply link. No login, no API key, no vendor database in the middle. Export to CSV, JSON, Excel, or XML.

Most job APIs resell a pre-indexed copy of the web and hand you what their crawler saw hours ago. This one calls Greenhouse, Lever, Ashby, Workday, iCIMS, Cornerstone, Dayforce, UKG Pro, Comeet, Pinpoint, JOIN, Oracle Cloud, Eightfold, SmartRecruiters, Recruitee, Workable, Rippling, Teamtailor, Breezy, Personio and BambooHR at the moment you run it, so a role posted five minutes ago is in your dataset and a role filled this morning is not.

| Who uses it | What they scrape job postings for |
|---|---|
| Recruiters and sourcers | Live openings at a named list of target companies, refreshed on a schedule |
| Sales and lead gen teams | Hiring signals: who is opening roles, in which team, in which city |
| Job board operators | A clean, deduplicated feed with salary, skills and full descriptions |
| Market and talent analysts | Pay bands, skill demand and headcount direction by company and function |
| Founders and investors | What a portfolio or a competitor is building next, read off their own board |

### What it does

This Actor collects job postings four ways, alone or together: paste career site URLs and it reads those boards directly, tick a job board (or paste any board's RSS URL) and it pulls that board's whole public feed, tick the built-in directory to sweep 305 verified company boards without naming anyone, or give it search terms and it queries a public index covering every Workable career site. All four paths return the same flat row, with a `sourceType` column saying whether it came from an applicant tracking system or a job board. Every posting carries:

- 🏢 **Core fields:** title, company, location, city, region, country, additional locations, apply URL and the source ATS.
- 🧭 **Structure:** work arrangement, employment type, department, team, seniority level, education level, requisition ID and language.
- 🗺️ **Geocoded location:** city, region, country, country code, latitude, longitude and timezone, resolved against 34,129 cities built into the Actor. No lookup service, no key, and it costs nothing extra.
- 📅 **Dates:** posted, updated and application deadline, plus a 300 character description snippet on every row for free.
- 💰 **Compensation:** salary minimum, maximum, currency and interval, with equity and bonus flags.
- 🧠 **Requirements:** skills named in the posting, years of experience, degree, visa sponsorship, security clearance, benefits, a job taxonomy, and the employer's own sentences summarising the responsibilities and the requirements.
- 🏛️ **Employer profile:** legal name, blurb, logo, headquarters, founding year, headcount and industry, plus the company's own LinkedIn, X and GitHub URLs.
- 📄 **Full description:** the complete posting as plain text and as HTML, plus the untouched source record if you want it.

Results export to CSV, JSON, Excel, or XML, or stream from the API.

### What you can do with career site data

**🎯 Track a list of target companies.**

Paste the career URLs of 50 companies and run it on a schedule. Every run returns their complete current openings, so a diff against the last run is your new-and-closed report.

**📡 Turn hiring into a sales signal.**

A company opening three roles in one team is a budget signal. Filter by department, seniority and posting date, and route the matches into your CRM.

**💵 Build a pay benchmark.**

Turn on compensation and keep only postings with a figure. Salary is read from the ATS where the employer published it and parsed out of the description where they did not, and every row with a figure says which of the two it was.

**🧩 Map skill demand.**

Turn on skills and count which technologies and certifications appear across a function, a city or an industry over time.

### Why choose this scraper

| | What you get |
|---|---|
| **Twenty-one ATS platforms, one schema** | Greenhouse, Lever, Ashby, Workday, iCIMS, Cornerstone OnDemand, Dayforce, UKG Pro (UltiPro), Comeet, Pinpoint, JOIN, Oracle Cloud HCM, Eightfold, SmartRecruiters, Recruitee, Workable, Rippling, Teamtailor, Breezy, Personio and BambooHR normalise into the same 37 columns. Eighteen of them are read board-wide; UKG Pro is read one posting at a time, because `recruiting.ultipro.com/robots.txt` allows its posting pages and disallows the only route that lists a board. |
| **305 boards built in** | Tick one box to search across verified company career sites without supplying a single URL. |
| **Fourteen job feeds on tap** | RemoteOK, Remotive, Arbeitnow, Jobicy, Himalayas, The Muse, Working Nomads, 4 Day Week, Landing.jobs, We Work Remotely, Jobspresso, Cryptocurrency Jobs EU Remote Jobs and SAP SuccessFactors, for breadth across thousands of employers at once. Scanning a board is free; you pay only for the rows you keep. |
| **A feed, not a re-scrape** | Turn on the incremental mode and a scheduled run returns only what is new since last time, and optionally one row per job that disappeared. |
| **Any other board, by feed** | Paste a board's RSS URL and it is read too. Most niche boards run WP Job Manager, whose feed carries the company, location and job type as structured fields, so a board nobody has mapped still comes back complete. |
| **Any company, on demand** | Paste a career URL and the platform is detected for you. No waiting for a vendor to add the company to their index. |
| **Live, not re-indexed** | The board is read when you press start, so what you get is what a candidate would see right now. |
| **Deterministic enrichment** | Salary, skills, seniority and sponsorship come from the source field first and a documented parser second, never from a one-shot LLM guess. |
| **Filters that cut noise** | Title, company, description, location, department, arrangement, type, language, date and salary. Filtering runs on the full posting, so you can filter on a block you did not buy. |
| **You pay for what you keep** | Only rows that pass your filters are written and billed, and each optional block bills only when it actually returned something. |

### How it compares

The best known actor in this category, `fantastic-jobs/career-site-job-listing-api`, sells a pre-indexed database that spans 54 ATS platforms with LinkedIn company data attached. It still covers more platforms than this Actor does, and if you need an untargeted firehose across the whole market it is the better tool. What it cannot do is read a specific company's board on demand, and its own documentation states a one hour delay plus a typical three hour indexing lag. This Actor covers twenty-one platforms, reads them live, and lets you name the company.

On company data the two differ in kind. They attach LinkedIn headcount, industry and follower counts. This Actor reads the employer's own website for the legal name, blurb, logo, headquarters and LinkedIn URL, and Wikidata for headcount and industry, and it accepts a match only when Wikidata's own official website is the same domain the posting resolved to. That is accurate where it fires and absent where it does not: measured on a live board run, 23% of rows carried a headcount. Follower counts are LinkedIn's alone and are not here. If you need a figure on every row, theirs is the one to buy.

On price, read your own plan rather than the headline. Their row is $0.012 on the free and Bronze plans and falls to $0.006 on Silver and $0.004 on Gold and above; the base row here is $0.004 falling to $0.00168, so it is cheaper on every plan, by 67% on the free plan and by 58% on Gold. Their row is all-inclusive, and so is that comparison: with every optional block ticked a row here is $0.0085 on the free plan and $0.0035 on Gold, still under their price on both.

| Feature | ParseForge | fantastic-jobs | jobo.world | webdata\_labs |
|---|---|---|---|---|
| Reads the company's board live | Yes | No, pre-indexed | Yes | Yes |
| Target a company by career URL | Yes, auto-detected | No | Partly | Yes |
| ATS platforms covered | 18 | 54 | Several | 3 |
| Cross-company search included | Yes, 14 job feeds + any RSS feed + 305 company boards + 168,000 postings | Yes, larger | No | No |
| Employer LinkedIn URL | Yes, from the company site | Yes, plus headcount | No | No |
| Salary with stated provenance | Yes | AI derived | No | No |
| Geocoded city, region, lat/lng, timezone | Yes | Yes | No | No |
| Incremental feed and expired jobs | Yes | Yes, plus a paid companion Actor | No | No |
| Company headcount and industry | Yes, from Wikidata | Yes, from LinkedIn | No | No |
| Screening questions | Yes, Recruitee | No | No | No |
| Price per job, free plan | $0.004 | $0.012 | $0.004 | $0.001 |
| Price per job, Gold and above | $0.00168 | $0.004 | $0.004 | $0.001 |

### What a job posting looks like

One real row from a verified run, unedited, with compensation, skills and screening questions turned on:

```json
{
  "jobId": "2699475",
  "title": "Customer Success Manager - Benelux",
  "company": "Tellent",
  "companyWebsite": "https://careers.tellent.com",
  "jobUrl": "https://careers.tellent.com/o/customer-success-manager-benelux-3",
  "applyUrl": "https://careers.tellent.com/o/customer-success-manager-benelux-3/c/new",
  "ats": "recruitee",
  "atsBoard": "careers.tellent.com",
  "boardUrl": "https://careers.tellent.com",
  "location": "Amsterdam, Noord-Holland, Netherlands",
  "city": "Amsterdam",
  "region": "Noord-Holland",
  "country": "NL",
  "additionalLocations": [],
  "isRemote": "No",
  "workArrangement": "Hybrid",
  "employmentType": "Full-time",
  "department": "1. GTM",
  "team": "sales",
  "seniorityLevel": "Mid level",
  "educationLevel": "Bachelor's degree",
  "datePosted": "2026-08-04T12:35:17.000Z",
  "dateUpdated": "2026-08-04T18:17:24.000Z",
  "applicationDeadline": "Not Disclosed",
  "requisitionId": "customer-success-manager-benelux-3",
  "language": "Not Disclosed",
  "hasSalary": "Yes",
  "salarySummary": "EUR 55000 - 65000 per year",
  "descriptionSnippet": "As a Customer Success Manager in the Benelux market, you will manage your own portfolio. As part of the Customer Experience team, you are responsible for customer retention and for helping our clients achieve their HR business goals through the use of our products. We act as trusted advisors, stayi…",
  "descriptionLength": 4761,
  "salaryMin": 55000,
  "salaryMax": 65000,
  "salaryCurrency": "EUR",
  "salaryInterval": "year",
  "salarySource": "ats-structured",
  "salaryText": "EUR 55000 - 65000 per year",
  "equityOffered": "No",
  "bonusOffered": "No",
  "skills": [
    "Roadmap",
    "Account Management",
    "Business Development",
    "English",
    "German",
    "French",
    "Dutch"
  ],
  "skillCount": 7,
  "yearsExperienceMin": "Not Disclosed",
  "experienceLevel": "Senior",
  "educationRequired": "Not Disclosed",
  "visaSponsorship": "Not Disclosed",
  "securityClearance": "No",
  "screeningQuestions": [
    {
      "question": "Do you have legal right to work in the country where this role is based? (e.g., EU citizenship, HSM, orientation visa, valid permit)?",
      "kind": "boolean",
      "required": "Yes"
    },
    {
      "question": "Please state your desired annual salary (gross)",
      "kind": "string",
      "required": "Yes"
    },
    {
      "question": "Are you proficient in both written and spoken Dutch at least at C1 level? (please note: this is a must-have requirement)",
      "kind": "boolean",
      "required": "Yes"
    },
    {
      "question": "Are you proficient in both written and spoken English at least at B2 level? (please note: this is a must-have requirement)",
      "kind": "boolean",
      "required": "Yes"
    },
    {
      "question": "Do you currently reside in the Netherlands?",
      "kind": "boolean",
      "required": "Yes"
    },
    {
      "question": "<p><em>As English is our official company language, we kindly request that all application materials are submitted in <strong>English</strong>. Please note that applications submitted in other languages <strong>cannot be reviewed</strong>. Thank you for your understanding!</em></p>",
      "kind": "infobox",
      "required": "No"
    }
  ],
  "screeningQuestionCount": 6,
  "scrapedAt": "2026-08-31T21:19:22.197Z"
}
```

Fields the source did not publish read `Not Disclosed` rather than `null`, so a CSV never has a hole in it.

### Configure the run

Two source types, alone or together: `careerSiteUrls` for named companies and `searchTerms` for the cross-company index. Filters run as each posting is read, so only matches reach your dataset. The Input tab lists every parameter.

Read named career sites in full, whatever the platform:

```json
{ "careerSiteUrls": ["https://job-boards.greenhouse.io/stripe", "https://jobs.lever.co/spotify", "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"], "maxItems": 500 }
```

Pull three job boards at once, keeping only the roles that pay:

```json
{ "jobBoards": ["remoteok", "himalayas", "weworkremotely"], "jobFeedUrls": ["https://example.com/?feed=job_feed"], "includeCompensation": true, "maxItems": 400 }
```

Run it hourly and get only what changed, plus the jobs that closed:

```json
{ "jobBoards": ["remoteok", "weworkremotely"], "onlyNewJobs": true, "includeExpiredJobs": true, "incrementalKey": "my-feed", "maxItems": 1000 }
```

Search across the built-in directory of 305 company boards without naming anyone:

```json
{ "searchTerms": ["staff software engineer"], "scanKnownCompanies": true, "includeCompensation": true, "maxItems": 300 }
```

Search every Workable career site for remote engineering roles posted this week, with pay:

```json
{ "searchTerms": ["software engineer"], "workArrangement": ["Remote"], "postedWithinDays": 7, "includeCompensation": true, "hasSalaryOnly": true, "maxItems": 300 }
```

Watch a target list for senior openings and pull the full description and skills:

```json
{ "careerSiteUrls": ["greenhouse:stripe", "lever:spotify", "https://careers.tellent.com"], "titleIncludes": ["senior", "staff", "principal"], "includeDescription": true, "includeSkills": true, "maxItems": 200 }
```

### Pricing

Pay-per-event: **$0.004 per job posting**, dropping to $0.00168 on Gold and above, plus a run-start fee of $0.002 on the free plan and $0.0002 on any paid one. Optional blocks are billed separately and only when they actually returned data: full description $0.001, compensation $0.001, skills $0.0008, company profile $0.0006, screening questions $0.0006, raw source record $0.0004. Discovery adds $0.004 per career site scanned and $0.002 per page of 20 search results, and neither is billed when nothing was delivered.

| Jobs collected | Base row only, free plan | Base row only, Gold | Every block ticked, Gold ceiling |
|---|---|---|---|
| 100 | $0.41 | $0.17 | $0.36 |
| 1,000 | $4.10 | $1.70 | $3.55 |
| 10,000 | $41.00 | $16.82 | $35.30 |

The right hand column is the ceiling: it assumes every optional block returned data on every row. Blocks that find nothing are not billed, so a real run with all six ticked lands below it.

For comparison, 1,000 jobs on `fantastic-jobs/career-site-job-listing-api` cost $12.01 on the free plan, $6.00 on Silver and $4.00 on Gold, against $4.10, $2.48 and $1.68 for the base row here. Even with every optional block ticked, a Gold row here is $0.0035 against their $0.004. New Apify accounts start with $5 in free credit.

### Free users

Free-plan runs return up to 10 jobs as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 jobs per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [Career Site Jobs Scraper](https://apify.com/parseforge/career-site-jobs-scraper?fpr=vmoqkp).
3. Paste `careerSiteUrls`, or type `searchTerms`, add any filters, tick the optional blocks you want, and click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to company career sites through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/career-site-jobs-scraper"
```

Then prompt it in plain language:

- *"Pull every open engineering role at Stripe and Spotify and group them by location."*
- *"Find remote data analyst jobs posted in the last week that publish a salary, and list the pay bands."*
- *"Read these ten career pages and tell me which companies are hiring for sales right now."*

Copy this into ChatGPT, Claude, or Cursor to start:

```
Use the Apify Actor "parseforge/career-site-jobs-scraper" to collect job postings from company career sites. Input: { "careerSiteUrls": ["<career page URL>"], "searchTerms": ["<keyword>"], "maxItems": <n>, "includeDescription": true, "includeCompensation": true }. It returns title, company, location, workArrangement, employmentType, department, seniorityLevel, datePosted, jobUrl, applyUrl, ats and salary fields per job. Call it with the ApifyClient and my APIFY_TOKEN.
```

### Troubleshooting

**Why am I getting no results?**

The career URL may point at a board with no open roles, or at a platform this Actor does not read. Check that the URL is a Greenhouse, Lever, Ashby, SmartRecruiters, Recruitee or Workable page. If you only used `searchTerms`, widen the term: the index matches words in the title and description, not fuzzy synonyms.

**Why fewer jobs than I asked for?**

The boards you named have that many open roles, or your filters removed the rest. A `location` or `postedWithinDays` filter cuts hard. Add more career sites or relax the filters.

**Why is a field empty?**

Not every ATS publishes every field. Greenhouse sends education and requisition ID but no salary; Ashby and Recruitee send structured pay; Lever puts pay in the description text. `Not Disclosed` means the source withheld it, and turning on compensation will still recover a figure from the description where the employer wrote one there.

**Why did the search stop early?**

The cross-company index rate limits by IP, and Apify's outbound addresses are shared with other runs. The Actor waits and retries the same page, and if the throttle persists it keeps everything collected so far and stops the search rather than burning your run. Career site URLs are read directly and are not affected.

**Why is the run slow?**

Search pages hold 20 results, so a large `maxItems` on `searchTerms` fetches many pages. Career site URLs are far faster, usually one request per company. The slowest combination is the built-in directory with the company profile turned on, because every new employer costs one extra request to their website: expect several minutes over the full directory. Turn the profile off, or narrow with filters, and the same sweep runs in well under a minute.

### FAQ

| Question | Answer |
|---|---|
| Do I need an API key for any of these platforms? | No. Every endpoint this Actor reads is public and anonymous, so there is nothing to register and no token to manage. |
| Which ATS platforms are supported? | Twenty-one: Greenhouse, Lever, Ashby, Workday, iCIMS, Cornerstone OnDemand, Dayforce, UKG Pro (UltiPro), Comeet, Pinpoint, JOIN, Oracle Cloud HCM, Eightfold, SmartRecruiters, Recruitee, Workable, Rippling, Teamtailor, Breezy, Personio and BambooHR. Paste any board or posting URL from those and the platform is detected automatically. UKG Pro takes a posting URL only: their robots.txt disallows the route that lists a board, so a UKG board URL is refused with that reason instead of being read. |
| Can I search without naming companies? | Yes, three ways. Tick one or more of the 14 `jobBoards`, or paste any board's RSS into `jobFeedUrls`, or use `searchTerms` for the cross-company index of Workable career sites, or `scanKnownCompanies` to sweep 305 verified company boards. |
| Do I get LinkedIn company data? | You get the company's own LinkedIn URL, plus legal name, logo, HQ and founding year from the employer's website. Headcount, industry and follower counts are LinkedIn's own data and are not included. |
| Can I scrape a company that is not in any index? | Yes. That is the point of `careerSiteUrls`. If the company runs one of the six platforms, its board is readable the moment it goes live. |
| How fresh is the data? | It is read at run time. There is no cached copy and no indexing lag between the employer posting and you seeing it. |
| Does it return the full job description? | Every row carries a 300 character snippet for free. Tick `includeDescription` for the complete text and HTML. |
| Where does the salary come from? | From the ATS field when the employer published one, otherwise parsed from the description. `salarySource` says which of the two it was, on every row that carries a figure. |
| Are the skills and seniority AI generated? | No. They come from the source field when there is one, and from a documented dictionary and regex pass otherwise, so the same posting always yields the same values. |
| How many jobs per run? | Free plan: 10. Paid: up to 1,000,000, bounded by how many open roles your sources actually have. |
| Does it deduplicate? | Yes, by job URL across every source in the run. One Workable board returned the same posting once per country; those collapse into one row with every location kept. |
| Can I run it as an incremental feed? | Yes. `onlyNewJobs` skips anything an earlier run under the same `incrementalKey` already delivered, and `includeExpiredJobs` adds a row per job that disappeared, marked `status=expired`. The first run seeds the state and returns everything. |
| Are the locations geocoded? | Yes, on every row and at no extra cost. 34,129 cities are built into the Actor, so `city`, `region`, `country`, `latitude`, `longitude` and `timezone` are filled from the location text without any lookup service. |
| Where do headcount and industry come from? | Wikidata, and only when its own record for that company points at the same website domain the posting resolved to. Where that check fails the columns are empty rather than guessed. |
| Is this an official product of any of these platforms? | No. It is unofficial and reads only public job board data. |

### Related actors

- [LinkedIn Jobs Scraper](https://apify.com/parseforge/linkedin-jobs-scraper?fpr=vmoqkp): public LinkedIn job postings by keyword and location.
- [Google Jobs Scraper](https://apify.com/parseforge/google-jobs-scraper?fpr=vmoqkp): aggregated listings from Google's job search results.
- [Greenhouse Jobs Scraper](https://apify.com/parseforge/greenhouse-jobs-scraper?fpr=vmoqkp): a single-platform reader for Greenhouse boards.
- [VC Portfolio Jobs Scraper](https://apify.com/parseforge/vc-portfolio-jobs-aggregator-scraper?fpr=vmoqkp): open roles across a venture firm's portfolio companies.
- [Hacker News Who's Hiring Jobs Scraper](https://apify.com/parseforge/hn-whoishiring-scraper?fpr=vmoqkp): structured jobs from the monthly hiring threads.

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Greenhouse, Lever, Ashby, SmartRecruiters, Recruitee, Workable, or any employer whose postings it reads. It collects only publicly available job board data. You are responsible for using the data in compliance with each platform's terms and applicable laws, including GDPR, CCPA, and PIPL. Do not use it to identify, profile, or target individuals.

# Actor input Schema

## `searchTerms` (type: `array`):

Job titles or keywords. They query the cross-company index of 168,000+ live postings, and they also filter the career sites you name below, keeping only jobs whose title or description matches one of them. Leave empty to get every open role from the career sites you list.

## `careerSiteUrls` (type: `array`):

Company career pages to read directly, one per line. Paste any Greenhouse, Lever, Ashby, Workday, iCIMS, Cornerstone (CSOD), Dayforce, UKG Pro, Comeet, Pinpoint, JOIN, Oracle Cloud, Eightfold, SmartRecruiters, Recruitee, Workable, Rippling, Teamtailor, Breezy, Personio or BambooHR URL (a board or a single posting) and the platform is detected for you, or use an explicit handle such as greenhouse:stripe, icims:careers-rambus, comeet:team8|61.003 or dayforce:cathaybank. UKG Pro (UltiPro) takes a posting URL rather than a board URL: their robots.txt disallows the route that lists a board.

## `scanKnownCompanies` (type: `boolean`):

Read every board in the bundled directory of verified company career sites as well, so you get cross-company results without naming any company. Combine with search terms and filters to keep it focused; a bare run over the whole directory returns a lot of rows.

## `jobBoards` (type: `array`):

Also pull from these cross-company job boards. Each one is a public feed of openings from thousands of employers, so this is the fastest way to get breadth without naming companies. Scanning a board costs nothing on its own; you pay only for the jobs you keep.

## `jobFeedUrls` (type: `array`):

Any other job board's RSS feed, one per line. Most niche boards run WP Job Manager, whose feed is at /?feed=job\_feed and carries the company, location and job type as structured fields, so a board nobody has mapped still comes back complete.

## `maxItems` (type: `integer`):

Free users: limited to 10 items (preview). Paid users: up to 1,000,000.

## `onlyNewJobs` (type: `boolean`):

Skip jobs already delivered by an earlier run with the same key. Ideal for a scheduled feed: you get today's new postings, not the whole board again.

## `includeExpiredJobs` (type: `boolean`):

Add one row per job that was in the previous run and is gone from this one, marked status=expired, so you can close them on your side. Only meaningful when this run covers the same sources as the last.

## `incrementalKey` (type: `string`):

Name for the remembered state, so you can keep several independent feeds side by side. Anything short and stable, for example eu-engineering.

## `locations` (type: `array`):

Keep only jobs whose location, city, region or country matches one of these. In search mode this also narrows the query at the source. Use the place name as the site writes it, for example Berlin, Germany, or New York.

## `locationExcludes` (type: `array`):

Drop jobs whose location matches any of these.

## `titleIncludes` (type: `array`):

Keep only jobs whose title contains one of these words. Case insensitive.

## `titleExcludes` (type: `array`):

Drop jobs whose title contains one of these words, for example intern or senior.

## `companyIncludes` (type: `array`):

Keep only jobs from companies whose name contains one of these words.

## `companyExcludes` (type: `array`):

Drop jobs from companies whose name contains one of these words. Useful for filtering out staffing agencies.

## `descriptionIncludes` (type: `array`):

Keep only jobs whose title or description contains one of these words, for example Kubernetes or visa sponsorship.

## `descriptionExcludes` (type: `array`):

Drop jobs whose title or description contains one of these words.

## `departmentIncludes` (type: `array`):

Keep only jobs whose department or team contains one of these words, for example Engineering or Sales.

## `workArrangement` (type: `array`):

Keep only jobs with one of these arrangements. Pick exactly one and the search query is narrowed at the source as well.

## `employmentTypes` (type: `array`):

Keep only jobs of these employment types. Pick exactly one and the search query is narrowed at the source as well.

## `languages` (type: `array`):

Keep only jobs whose posting language matches one of these two letter codes, for example en, de, fr. Only Greenhouse, SmartRecruiters and the search index publish this field.

## `postedWithinDays` (type: `integer`):

Keep only jobs published in the last N days. In search mode this also narrows the query at the source, which is much faster than filtering afterwards.

## `hasSalaryOnly` (type: `boolean`):

Keep only jobs where a pay figure was found, either published by the ATS or read out of the description.

## `minSalary` (type: `number`):

Keep only jobs paying at least this much per year. Hourly, daily, weekly and monthly figures are converted before the comparison. Leave empty to ignore pay.

## `includeDescription` (type: `boolean`):

Add the complete job description as plain text and as HTML.

## `includeCompensation` (type: `boolean`):

Add salary minimum, maximum, currency and interval, plus equity and bonus flags. Read from the ATS when it publishes pay, otherwise parsed out of the description, and always stamped with which of the two it was.

## `includeSkills` (type: `boolean`):

Add the skills named in the posting, the minimum years of experience, the experience band, the degree required, visa sponsorship and security clearance.

## `includeCompanyProfile` (type: `boolean`):

Add the employer logo, the company blurb and the company domain, where the source publishes them.

## `includeScreeningQuestions` (type: `boolean`):

Add the application's screening questions with their type and whether they are required. Recruitee career sites only.

## `includeRawSource` (type: `boolean`):

Add the untouched JSON record exactly as the ATS returned it, for auditing or for fields this actor does not map.

## Actor input object example

```json
{
  "searchTerms": [
    "software engineer"
  ],
  "careerSiteUrls": [],
  "scanKnownCompanies": false,
  "jobBoards": [],
  "jobFeedUrls": [],
  "maxItems": 10,
  "onlyNewJobs": false,
  "includeExpiredJobs": false,
  "incrementalKey": "default",
  "locations": [],
  "locationExcludes": [],
  "titleIncludes": [],
  "titleExcludes": [],
  "companyIncludes": [],
  "companyExcludes": [],
  "descriptionIncludes": [],
  "descriptionExcludes": [],
  "departmentIncludes": [],
  "workArrangement": [],
  "employmentTypes": [],
  "languages": [],
  "hasSalaryOnly": false,
  "includeDescription": false,
  "includeCompensation": false,
  "includeSkills": false,
  "includeCompanyProfile": false,
  "includeScreeningQuestions": false,
  "includeRawSource": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `csv` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "software engineer"
    ],
    "careerSiteUrls": [],
    "jobBoards": [],
    "jobFeedUrls": [],
    "maxItems": 10,
    "locations": [],
    "locationExcludes": [],
    "titleIncludes": [],
    "titleExcludes": [],
    "companyIncludes": [],
    "companyExcludes": [],
    "descriptionIncludes": [],
    "descriptionExcludes": [],
    "departmentIncludes": [],
    "workArrangement": [],
    "employmentTypes": [],
    "languages": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/career-site-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["software engineer"],
    "careerSiteUrls": [],
    "jobBoards": [],
    "jobFeedUrls": [],
    "maxItems": 10,
    "locations": [],
    "locationExcludes": [],
    "titleIncludes": [],
    "titleExcludes": [],
    "companyIncludes": [],
    "companyExcludes": [],
    "descriptionIncludes": [],
    "descriptionExcludes": [],
    "departmentIncludes": [],
    "workArrangement": [],
    "employmentTypes": [],
    "languages": [],
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/career-site-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "software engineer"
  ],
  "careerSiteUrls": [],
  "jobBoards": [],
  "jobFeedUrls": [],
  "maxItems": 10,
  "locations": [],
  "locationExcludes": [],
  "titleIncludes": [],
  "titleExcludes": [],
  "companyIncludes": [],
  "companyExcludes": [],
  "descriptionIncludes": [],
  "descriptionExcludes": [],
  "departmentIncludes": [],
  "workArrangement": [],
  "employmentTypes": [],
  "languages": []
}' |
apify call parseforge/career-site-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/career-site-jobs-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/M0cXf7wYCgkwZ9ioo/builds/9EFODRlkLxq5yBv9D/openapi.json
