Workday Jobs Scraper
Pricing
from $0.84 / 1,000 results
Workday Jobs Scraper
Jobs from any Workday career site (<tenant>.myworkdayjobs.com) via its public CXS API: title, req id, locations, country, time and remote type, posting date, full description, apply URL. Any tenant by URL - NVIDIA, Salesforce, Adobe - with keyword search and facets; robots.txt honoured per host.
Pricing
from $0.84 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Jobs from any Workday career site — <tenant>.wd<N>.myworkdayjobs.com,
the platform behind the careers pages of NVIDIA, Salesforce, Adobe and
thousands of other employers — through its public, keyless CXS API: title,
req id, location and additional locations, country, time type, remote type,
the exact posting date, the full description (text and HTML) and the apply
URL.
HTTP only, no login, no key, no browser. One request per 20 jobs, plus one per job for the detail.
What it is for
- Employer watchlists — the same input on a schedule, diff by
jobReqId. - Cross-company search — several sites, one keyword, one dataset.
- ATS research — every field the site's own careers page shows, as data.
Input
| field | what it does |
|---|---|
careerSiteUrls | One or more Workday site URLs. Any URL on the site works. |
searchTerms | The site's own keyword search. Empty = every job. |
facetFilters | {facetParameter: [valueIds]} — ids come from availableFacets in a summary row. |
includeDetails | On by default: one request per job for description, exact date, locations. |
maxItems, maxConcurrency, minRequestInterval, proxyConfiguration | Limits. |
Targets are careerSiteUrls × searchTerms; each gets a SEARCH_SUMMARY
with the host's robots verdict and its facet vocabulary.
Four things about this API worth knowing before you trust a run
1. Each tenant is its own site, and its robots.txt is checked at run time
A Workday tenant is a separate host with its own robots.txt. The Actor
fetches it before the first job request and refuses any host whose rules
disallow the career-site path or the /wday/cxs/ API — for everyone or for
Claude by name. The verdict is on every summary row (robotsVerdict). The
tenants measured while building this allow their sites and name no AI bot.
2. The API pages to 2,000 rows and then serves page 1 again
Page size is 20 (anything larger is an HTTP 400). Offset 1,980 serves the
last new page; offset 2,000, 2,020, 5,000 all serve the same 20 jobs as
offset 0 — no error, no empty page. NVIDIA's own facet counts add to
~2,300 jobs, so the site holds more than the API pages to. The Actor never
asks past offset 1,980, ends on a page with no new ids, and reports
wallReached. Narrow with searchTerms or facetFilters to get under
2,000.
3. total is right on page 1 and zero on every page after
A paginator that re-reads the total per page stops after page 1 with "0 results". It is read once, from page 1.
4. The list's date is relative text
The list says "Posted Today" / "Posted 3 Days Ago" / "Posted 30+ Days Ago";
the detail carries the ISO date (postedDate). That is why includeDetails
is on by default — the list alone is title, location text, relative date and
req id.
Other things measured
- An unknown career site is a 404 and an unknown tenant a 422 —
site_not_foundwith the API's message. A facet id the site does not know is a 400 —bad_request, with a pointer toavailableFacets. Nonsense search text is a clean 0. - Facet values can be nested (locations inside regions);
availableFacetsflattens them todescriptor,id,count, up to 60 per parameter. - No throttling on any tenant tried (101 unpaced requests).
Output
JOB—jobReqId,jobPostingId,title,url,externalUrl,location,locationsText,additionalLocations,country,timeType,remoteType,postedDate,postedOnText,canApply,description,descriptionHtml,tenant,careerSite,host,query,resultPosition.SEARCH_SUMMARY—careerSiteUrl,robotsVerdict,totalResults,jobsReturned,pagesFetched,stoppedReason,wallReached,detailsFetched,detailsFailed,availableFacets.ERROR—robots_disallowed,site_not_found,bad_request,payload_shape_changed,fetch_failed, with detail.
Known limits
- Salary is not in the CXS payloads for the tenants measured; it appears only inside descriptions when the employer writes it there.
- Internal (employee-only) career sites need a login and are out of scope.
- Sites that disallow crawling in
robots.txtreturn no rows, by design.