# Workday Jobs Scraper (`scrapyx/workday-jobs-scraper`) Actor

Jobs from any Workday career site (<tenant>.myworkdayjobs.com) via its public CXS API: title, req id, locations, country, time and remote type, posting date, full description, apply URL. Any tenant by URL - NVIDIA, Salesforce, Adobe - with keyword search and facets; robots.txt honoured per host.

- **URL**: https://apify.com/scrapyx/workday-jobs-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Workday Jobs Scraper

Jobs from **any Workday career site** — `<tenant>.wd<N>.myworkdayjobs.com`,
the platform behind the careers pages of NVIDIA, Salesforce, Adobe and
thousands of other employers — through its public, keyless CXS API: title,
req id, location and additional locations, country, time type, remote type,
the exact posting date, the full description (text and HTML) and the apply
URL.

HTTP only, no login, no key, no browser. One request per 20 jobs, plus one
per job for the detail.

### What it is for

- **Employer watchlists** — the same input on a schedule, diff by `jobReqId`.
- **Cross-company search** — several sites, one keyword, one dataset.
- **ATS research** — every field the site's own careers page shows, as data.

### Input

| field | what it does |
| --- | --- |
| `careerSiteUrls` | One or more Workday site URLs. Any URL on the site works. |
| `searchTerms` | The site's own keyword search. Empty = every job. |
| `facetFilters` | `{facetParameter: [valueIds]}` — ids come from `availableFacets` in a summary row. |
| `includeDetails` | On by default: one request per job for description, exact date, locations. |
| `maxItems`, `maxConcurrency`, `minRequestInterval`, `proxyConfiguration` | Limits. |

Targets are `careerSiteUrls` × `searchTerms`; each gets a `SEARCH_SUMMARY`
with the host's robots verdict and its facet vocabulary.

### Four things about this API worth knowing before you trust a run

#### 1. Each tenant is its own site, and its robots.txt is checked at run time

A Workday tenant is a separate host with its own `robots.txt`. The Actor
fetches it before the first job request and refuses any host whose rules
disallow the career-site path or the `/wday/cxs/` API — for everyone or for
Claude by name. The verdict is on every summary row (`robotsVerdict`). The
tenants measured while building this allow their sites and name no AI bot.

#### 2. The API pages to 2,000 rows and then serves page 1 again

Page size is 20 (anything larger is an HTTP 400). Offset 1,980 serves the
last new page; offset 2,000, 2,020, 5,000 all serve **the same 20 jobs as
offset 0** — no error, no empty page. NVIDIA's own facet counts add to
\~2,300 jobs, so the site holds more than the API pages to. The Actor never
asks past offset 1,980, ends on a page with no new ids, and reports
`wallReached`. Narrow with `searchTerms` or `facetFilters` to get under
2,000.

#### 3. `total` is right on page 1 and zero on every page after

A paginator that re-reads the total per page stops after page 1 with "0
results". It is read once, from page 1.

#### 4. The list's date is relative text

The list says "Posted Today" / "Posted 3 Days Ago" / "Posted 30+ Days Ago";
the detail carries the ISO date (`postedDate`). That is why `includeDetails`
is on by default — the list alone is title, location text, relative date and
req id.

### Other things measured

- An unknown career site is a 404 and an unknown tenant a 422 —
  `site_not_found` with the API's message. A facet id the site does not
  know is a 400 — `bad_request`, with a pointer to `availableFacets`.
  Nonsense search text is a clean 0.
- Facet values can be nested (locations inside regions); `availableFacets`
  flattens them to `descriptor`, `id`, `count`, up to 60 per parameter.
- No throttling on any tenant tried (101 unpaced requests).

### Output

- **`JOB`** — `jobReqId`, `jobPostingId`, `title`, `url`, `externalUrl`,
  `location`, `locationsText`, `additionalLocations`, `country`,
  `timeType`, `remoteType`, `postedDate`, `postedOnText`, `canApply`,
  `description`, `descriptionHtml`, `tenant`, `careerSite`, `host`, `query`,
  `resultPosition`.
- **`SEARCH_SUMMARY`** — `careerSiteUrl`, `robotsVerdict`, `totalResults`,
  `jobsReturned`, `pagesFetched`, `stoppedReason`, `wallReached`,
  `detailsFetched`, `detailsFailed`, `availableFacets`.
- **`ERROR`** — `robots_disallowed`, `site_not_found`, `bad_request`,
  `payload_shape_changed`, `fetch_failed`, with detail.

### Known limits

- Salary is not in the CXS payloads for the tenants measured; it appears
  only inside descriptions when the employer writes it there.
- Internal (employee-only) career sites need a login and are out of scope.
- Sites that disallow crawling in `robots.txt` return no rows, by design.

# Actor input Schema

## `careerSiteUrls` (type: `array`):

One or more Workday career sites: https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite, https://salesforce.wd12.myworkdayjobs.com/External\_Career\_Site, https://adobe.wd5.myworkdayjobs.com/external\_experienced ... Any URL on the site works (an /en-US/ segment or a job path is fine). Each host's robots.txt is checked before anything is read; a site that disallows crawling is reported as robots\_disallowed.

## `searchTerms` (type: `array`):

Keywords for the site's own search (title, description, req id). Each term is its own target per site. Empty = every job on the site.

## `facetFilters` (type: `object`):

Workday facet ids, as {"facetParameter": \["valueId", ...]} - e.g. {"timeType": \["5509c0b5959810ac0029943377d47364"]}. Ids differ per tenant: run once and copy them from availableFacets in the summary row (each facet lists its values with descriptor, id and count).

## `includeDetails` (type: `boolean`):

On by default: one extra request per job for the description, exact posting date, location list, country, time type, remote type and apply URL. The list alone carries title, location text, a relative date and the req id.

## `maxItems` (type: `integer`):

Overall cap on JOB rows across every target. The API pages 20 at a time and stops at 2,000 rows per search - past that it silently serves page 1 again; a target that gets there reports wallReached. Narrow with searchTerms or facetFilters.

## `maxConcurrency` (type: `integer`):

Parallel requests across sites and job details.

## `minRequestInterval` (type: `integer`):

Politeness delay between request starts. The API tolerated 101 unpaced requests.

## `proxyConfiguration` (type: `object`):

Optional. No anti-bot layer on any tenant tried. Enable Apify's free datacenter proxy only if a cloud run reports fetch\_failed.

## Actor input object example

```json
{
  "careerSiteUrls": [
    "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"
  ],
  "searchTerms": [
    "engineer"
  ],
  "facetFilters": {},
  "includeDetails": true,
  "maxItems": 100,
  "maxConcurrency": 4,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "careerSiteUrls": [
        "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"
    ],
    "searchTerms": [
        "engineer"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/workday-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "careerSiteUrls": ["https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"],
    "searchTerms": ["engineer"],
}

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/workday-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "careerSiteUrls": [
    "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"
  ],
  "searchTerms": [
    "engineer"
  ]
}' |
apify call scrapyx/workday-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/workday-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/h1VJqPlQqr4vJSPkS/builds/9Mg6bpyOye7ifTURD/openapi.json
