Workday Jobs Scraper avatar

Workday Jobs Scraper

Pricing

from $0.84 / 1,000 results

Go to Apify Store
Workday Jobs Scraper

Workday Jobs Scraper

Jobs from any Workday career site (<tenant>.myworkdayjobs.com) via its public CXS API: title, req id, locations, country, time and remote type, posting date, full description, apply URL. Any tenant by URL - NVIDIA, Salesforce, Adobe - with keyword search and facets; robots.txt honoured per host.

Pricing

from $0.84 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Jobs from any Workday career site — <tenant>.wd<N>.myworkdayjobs.com, the platform behind the careers pages of NVIDIA, Salesforce, Adobe and thousands of other employers — through its public, keyless CXS API: title, req id, location and additional locations, country, time type, remote type, the exact posting date, the full description (text and HTML) and the apply URL.

HTTP only, no login, no key, no browser. One request per 20 jobs, plus one per job for the detail.

What it is for

  • Employer watchlists — the same input on a schedule, diff by jobReqId.
  • Cross-company search — several sites, one keyword, one dataset.
  • ATS research — every field the site's own careers page shows, as data.

Input

fieldwhat it does
careerSiteUrlsOne or more Workday site URLs. Any URL on the site works.
searchTermsThe site's own keyword search. Empty = every job.
facetFilters{facetParameter: [valueIds]} — ids come from availableFacets in a summary row.
includeDetailsOn by default: one request per job for description, exact date, locations.
maxItems, maxConcurrency, minRequestInterval, proxyConfigurationLimits.

Targets are careerSiteUrls × searchTerms; each gets a SEARCH_SUMMARY with the host's robots verdict and its facet vocabulary.

Four things about this API worth knowing before you trust a run

1. Each tenant is its own site, and its robots.txt is checked at run time

A Workday tenant is a separate host with its own robots.txt. The Actor fetches it before the first job request and refuses any host whose rules disallow the career-site path or the /wday/cxs/ API — for everyone or for Claude by name. The verdict is on every summary row (robotsVerdict). The tenants measured while building this allow their sites and name no AI bot.

2. The API pages to 2,000 rows and then serves page 1 again

Page size is 20 (anything larger is an HTTP 400). Offset 1,980 serves the last new page; offset 2,000, 2,020, 5,000 all serve the same 20 jobs as offset 0 — no error, no empty page. NVIDIA's own facet counts add to ~2,300 jobs, so the site holds more than the API pages to. The Actor never asks past offset 1,980, ends on a page with no new ids, and reports wallReached. Narrow with searchTerms or facetFilters to get under 2,000.

3. total is right on page 1 and zero on every page after

A paginator that re-reads the total per page stops after page 1 with "0 results". It is read once, from page 1.

4. The list's date is relative text

The list says "Posted Today" / "Posted 3 Days Ago" / "Posted 30+ Days Ago"; the detail carries the ISO date (postedDate). That is why includeDetails is on by default — the list alone is title, location text, relative date and req id.

Other things measured

  • An unknown career site is a 404 and an unknown tenant a 422 — site_not_found with the API's message. A facet id the site does not know is a 400 — bad_request, with a pointer to availableFacets. Nonsense search text is a clean 0.
  • Facet values can be nested (locations inside regions); availableFacets flattens them to descriptor, id, count, up to 60 per parameter.
  • No throttling on any tenant tried (101 unpaced requests).

Output

  • JOB — jobReqId, jobPostingId, title, url, externalUrl, location, locationsText, additionalLocations, country, timeType, remoteType, postedDate, postedOnText, canApply, description, descriptionHtml, tenant, careerSite, host, query, resultPosition.
  • SEARCH_SUMMARY — careerSiteUrl, robotsVerdict, totalResults, jobsReturned, pagesFetched, stoppedReason, wallReached, detailsFetched, detailsFailed, availableFacets.
  • ERROR — robots_disallowed, site_not_found, bad_request, payload_shape_changed, fetch_failed, with detail.

Known limits

  • Salary is not in the CXS payloads for the tenants measured; it appears only inside descriptions when the employer writes it there.
  • Internal (employee-only) career sites need a login and are out of scope.
  • Sites that disallow crawling in robots.txt return no rows, by design.