# Workday Jobs Scraper - Real Locations, No API Key (`maydit/workday-jobs-scraper`) Actor

Scrape live job postings from any Workday career site. Expands "12 Locations" into the real city list, adds descriptions, exports CSV/JSON.

- **URL**: https://apify.com/maydit/workday-jobs-scraper.md
- **Developed by:** [Brandt May](https://apify.com/maydit) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.60 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Workday Jobs Scraper - real locations, no API key

Scrape live job postings from any company's **Workday career site** (`*.myworkdayjobs.com`) and export them to CSV, JSON or Excel. Give the Actor the career-site URL you see in your browser; it reads the same public JSON the career page itself calls, so there is no API key, no login and no proxy to configure.

The part that makes this different: Workday's job list collapses multi-city roles into a label like **"12 Locations"**. Most exports stop there. This Actor opens each of those postings and writes the **real list of twelve cities** into a `locations` array, plus `locationCount`, alongside the raw label.

This Actor reads Workday career sites and nothing else. If an employer you want is not on Workday, this is not the Actor for them - an entry that is not a `*.myworkdayjobs.com` career site is skipped with a warning rather than quietly fetched from somewhere else.

Typical uses: job-board aggregation, recruiting and talent-market research, competitor hiring signals, tracking where a company is opening headcount, and feeding an "new openings since yesterday" alert.

***

### What you get

One row per posting, same shape for every employer.

| Field | Type | Notes |
|---|---|---|
| `company` | string | Workday tenant (`nvidia`) |
| `source` | string | Always `"workday"` - the system the row came from. Constant, and also part of the change-detection key |
| `tenant` | string | Workday tenant |
| `site` | string | Workday career-site name |
| `jobId` | string | Requisition id (`JR336652`) |
| `title` | string | Job title |
| `locationsText` | string | Exactly what the job list shows, including `"12 Locations"` |
| `locations` | array | The real locations. Filled from the detail call when `fetchDetails` is on |
| `locationCount` | number | Length of `locations` (falls back to the number in the label) |
| `locationsExpanded` | boolean | `true` when a collapsed `"N Locations"` label was opened into real cities |
| `postedOn` | string | Workday's human label (`Posted Today`, `Posted 30+ Days Ago`) |
| `postedDate` | string | `YYYY-MM-DD`. Exact when details are fetched; derived from the label otherwise. `"30+ Days Ago"` is left `null` rather than guessed |
| `timeType` | string | `Full time` / `Part time` |
| `remoteType` | string | Tenant-dependent: Salesforce publishes `Office - Flexible`, NVIDIA publishes nothing |
| `url` | string | Public posting URL |
| `descriptionText` | string | Plain text, HTML stripped, capped at 5,000 characters. Present only when `fetchDetails` is on |
| `scrapedAt` | string | ISO timestamp of the fetch |

A `SUMMARY` record is written to the run's key-value store with per-company counts, the reason any company failed, and whether the run stopped early on its time budget.

***

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `companies` | array | NVIDIA + Salesforce career sites | One entry per employer. Paste the career-site URL as-is (`https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite`); a `/en-US/` segment or a link to a single job is fine. A `tenant|wdN|site` triple also works |
| `searchText` | string | *(empty)* | Passed to Workday's own search, which matches title and description |
| `maxResultsPerCompany` | integer | `40` | Postings per employer. Workday serves 20 rows per request, so 40 is two list calls |
| `fetchDetails` | boolean | `true` | Opens each posting for the real locations, exact posting date, time type, remote type and description |
| `onlyNewSinceLastRun` | boolean | `false` | Emit only postings that were not in the previous run's baseline |
| `stateStoreName` | string | `workday-jobs-scraper-state` | Named key-value store holding that baseline |
| `maxRunSeconds` | integer | `240` | Wall-clock budget. The run stops cleanly at this point and keeps what it has |

#### Finding your company's career-site URL

Open the employer's "Search jobs" page. If the address bar reads

```
https://salesforce.wd12.myworkdayjobs.com/en-US/External_Career_Site
```

then `salesforce` is the tenant, `wd12` is the data centre and `External_Career_Site` is the site name. Paste the whole URL - the Actor pulls the three pieces out of it. Both the data centre number and the site name are case- and spelling-sensitive, and the Actor tells you which of the two is wrong if a company fails.

***

### Example output

A real row from a default run (description truncated here for readability):

```json
{
  "company": "salesforce",
  "source": "workday",
  "tenant": "salesforce",
  "site": "External_Career_Site",
  "jobId": "JR336652",
  "title": "Success Architect - Agentforce (U.S. Citizen)",
  "locationsText": "12 Locations",
  "locations": [
    "Illinois - Chicago",
    "Massachusetts - Boston",
    "Washington - Seattle",
    "Massachusetts - Burlington",
    "Washington - Bellevue",
    "Texas - Dallas",
    "California - Irvine",
    "Colorado - Denver",
    "Georgia - Atlanta",
    "Indiana - Indianapolis",
    "Virginia - Mclean",
    "Texas - Austin"
  ],
  "locationCount": 12,
  "locationsExpanded": true,
  "postedOn": "Posted Today",
  "postedDate": "2026-09-23",
  "timeType": "Full time",
  "remoteType": "Office - Flexible",
  "url": "https://salesforce.wd12.myworkdayjobs.com/External_Career_Site/job/Illinois---Chicago/Agentforce-Success-Architect_JR336652",
  "descriptionText": "To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months...",
  "scrapedAt": "2026-09-24T01:25:44.304Z"
}
```

The run that produced this row wrote **80 postings from two employers in 9 seconds**, 27 of which had a collapsed `"N Locations"` label that was expanded into real cities.

***

### FAQ

**Why does Workday show "12 Locations" instead of the cities, and what does this Actor do about it?**
Workday's list endpoint returns a single `locationsText` string per posting, and for multi-city requisitions that string is just a count. The city list only exists on the individual posting. When `fetchDetails` is on, this Actor opens each posting and writes every location into the `locations` array, keeping the original label in `locationsText` so nothing is lost.

**Do I need an API key, a Workday account or a proxy?**
No. These are the employers' own public career-site endpoints and they answer anonymous requests, with no key and no account. Some employers do restrict their career-site JSON to non-datacentre traffic; if one of yours does, the run reports it as an HTTP 403 for that company, tells you a proxy is the only thing that would change it, and carries on with the rest. If you hit that, switch the run to a residential proxy in the Actor's run settings.

**The employer I want is not on Workday - can this Actor still get them?**
No. This Actor speaks Workday's career-site API only, and it will not silently fall back to a different job-posting system just to return rows. A non-Workday entry is skipped with a warning naming what a valid career-site URL looks like. Check the employer's "Search jobs" page: if the address bar is not `*.myworkdayjobs.com`, they are on another platform and you want a scraper built for that platform.

**What happens on the first run with "Only postings new since the last run" turned on?**
There is no baseline yet, so **everything is returned** and saved as the baseline; the log says so plainly. From the second run on you get only postings that were not there before, and an empty dataset means nothing new was posted. The baseline lives in a named key-value store (`stateStoreName`), so give separate scheduled jobs separate names.

**Can I get salary or compensation data?**
There is no separate pay field in Workday's public career-site JSON, so this Actor does not invent one. Where an employer states a pay range, it is inside the posting text and therefore inside `descriptionText`.

**Why does a big employer report exactly 2,000 results?**
Workday's `total` saturates at 2000 on large boards (NVIDIA reports 2000 with or without a search term), so it is an estimate of board size, not a precise count. Use `searchText` to narrow a large board instead of trying to page through all of it.

***

### Data source, access and limits

- **Where the data comes from.** Every row comes from the employer's own career-site API, `https://{tenant}.wd{N}.myworkdayjobs.com/wday/cxs/{tenant}/{site}/jobs` and the matching per-job endpoint. These are the public, keyless endpoints the career pages themselves use. Nothing here is behind a login. Every row is an employer-published job advertisement, not a person: the Actor collects no candidate, applicant or employee data, and it never touches an employer's applicant-tracking data. Note that `descriptionText` reproduces the employer's own advertisement text verbatim, and some employers print a recruiting or accommodation-request mailbox in it. Those addresses are the employer's own published contact details. Turn `fetchDetails` off if you would rather not store the ad text at all.
- **Page size.** Workday caps its list endpoint at **20 rows per request** (21 returns HTTP 400), so a 1,000-posting board is 50 sequential list calls plus one detail call per posting. Budget time accordingly with `maxRunSeconds`.
- **Politeness.** Requests are sequential per employer with a short pause between list pages; detail lookups run at most five at a time.
- **Time budget.** The run stops cleanly when it reaches `maxRunSeconds` (or the platform run timeout), keeps everything collected so far, and logs exactly how much was left and what to change. A partial run is a successful run with a warning, not a failure.
- **Failures are loud.** If one employer's URL is wrong, that company is reported with the specific reason (wrong site name vs wrong `wd` data centre) and the others still run. If **every** employer fails, the run fails rather than handing you an empty file that looks like "no jobs".
- **Freshness.** Every row is fetched live at run time; nothing is served from a cache.
- **Billing.** This Actor is a paid Actor: the price and pricing model in force are the ones shown on its Apify Store page and on the run screen before you start a run. Apify platform usage (compute units, storage) is charged on top under your own plan. Use `maxResultsPerCompany` and `maxRunSeconds` to cap the size of a run.

# Actor input Schema

## `companies` (type: `array`):

One entry per employer. Paste the career-site URL exactly as it appears in the browser, e.g. https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite - the tenant, the wd number and the site name are all read out of it (a /en-US/ segment or a link to a single job is fine). A 'tenant|wdN|site' triple also works. Workday career sites only; an employer that publishes on another system is skipped with a warning.

## `searchText` (type: `string`):

Keyword passed to Workday's own job search (it matches title and description). Leave empty for the whole board.

## `maxResultsPerCompany` (type: `integer`):

Cap on postings written for each employer. Workday serves 20 rows per request (its hard limit), so 40 means two list calls per company. Raise it for a full board - a 1,000-posting board is 50 list calls plus one detail call each, so raise 'Max run seconds' with it.

## `fetchDetails` (type: `boolean`):

Calls each posting's detail endpoint. This is what turns Workday's collapsed "12 Locations" label into a real list of twelve cities, and it also adds the posting date, time type, remote type and the description text (capped at 5,000 characters). Turn it off for a faster, shallower run.

## `onlyNewSinceLastRun` (type: `boolean`):

Compares this run against the job ids stored by the previous run in a named key-value store and writes only postings that were not there. The FIRST run has no baseline, so it returns everything and saves it - from the second run on you get just the new openings. An empty dataset then means nothing new was posted.

## `stateStoreName` (type: `string`):

Name of the named key-value store that holds the seen-job-id baseline. Only used when 'Only postings new since the last run' is on. Give each scheduled job its own name if you track different company lists separately.

## `maxRunSeconds` (type: `integer`):

Wall-clock budget. The Actor stops cleanly at this point and keeps everything collected so far, instead of being killed mid-write. Raise it for large boards; the platform run timeout still applies on top.

## Actor input object example

```json
{
  "companies": [
    "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
    "https://salesforce.wd12.myworkdayjobs.com/External_Career_Site"
  ],
  "searchText": "engineer",
  "maxResultsPerCompany": 40,
  "fetchDetails": true,
  "onlyNewSinceLastRun": false,
  "stateStoreName": "workday-jobs-scraper-state",
  "maxRunSeconds": 240
}
```

# Actor output Schema

## `results` (type: `string`):

One row per job posting, with collapsed "N Locations" labels expanded into the real location list.

## `summary` (type: `string`):

Totals for the run, including how many records were requested versus returned and whether the run stopped early on its time budget.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
        "https://salesforce.wd12.myworkdayjobs.com/External_Career_Site"
    ],
    "maxResultsPerCompany": 40
};

// Run the Actor and wait for it to finish
const run = await client.actor("maydit/workday-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
        "https://salesforce.wd12.myworkdayjobs.com/External_Career_Site",
    ],
    "maxResultsPerCompany": 40,
}

# Run the Actor and wait for it to finish
run = client.actor("maydit/workday-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
    "https://salesforce.wd12.myworkdayjobs.com/External_Career_Site"
  ],
  "maxResultsPerCompany": 40
}' |
apify call maydit/workday-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,maydit/workday-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/PKbrbcPx5GPuLgLSc/builds/ckrHjYNUlvdVjbj3N/openapi.json
