Jobs.ac.uk Search Scraper
Pricing
from $2.99 / 1,000 jobs.ac.uk job records
Jobs.ac.uk Search Scraper
Scrape job listings from Jobs.ac.uk, the leading UK job board for academic, research, and higher education positions. Extract job titles, institutions, locations, salary ranges, contract types, and descriptions for academic recruitment.
Pricing
from $2.99 / 1,000 jobs.ac.uk job records
Rating
0.0
(0)
Developer
Jobs API
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
11 days ago
Last modified
Categories
Share
Jobs.ac.uk Jobs Search Scraper
This Actor performs bounded extraction of public Jobs.ac.uk academic job pages using ordinary direct HTTPS requests and Cheerio parsing. Requests are sequential, have no retry, custom user-agent, proxy, or fingerprint controls, and stop on a public access barrier. It does not sign in, submit applications, or fetch external application destinations.
Modes
searchdiscovers public Jobs.ac.uk vacancies through the official search page, then fetches each detail page.singlefetches one URL fromjobUrl,directUrl, orstartUrl.multiplefetches URLs fromjobUrls,directUrls, orstartUrls.
Every emitted row is buffered until its detail page, canonical URL, job identity, description, and required fields are verified. Incomplete or mismatched jobs are skipped in search mode; transport, access, and source-integrity failures write structured diagnostics to the key-value store and emit no misleading row.
Output
Rows preserve the public academic vacancy data available on the page: title, institution, department, public advert categories (role, subject areas, and location areas), full text and HTML description, sectioned responsibilities/qualifications/benefits, locations, salary, hours, contract type, posting and closing dates, job reference, structured JobPosting data, description links, contact emails, explicit external application links, source listings, request receipts, field coverage, and execution-quality evidence. The dataset accepts a record only when at least 21 distinct meaningful source facts are populated; the count excludes identifiers, repeated aliases, URLs used only for provenance, and runtime metadata.
applicationUrl is emitted only when Jobs.ac.uk publishes an explicit application action. The canonical Jobs.ac.uk vacancy URL is never presented as an application link.
Local run
npm install$env:APIFY_LOCAL_STORAGE_DIR = 'storage-validation-20260927-01'apify run --resurrect --input-file INPUT.jsonnpm run validate -- $env:APIFY_LOCAL_STORAGE_DIR
The reproducible fixtures INPUT.json, INPUT-single.json, and INPUT-multiple.json exercise all three modes. The dataset is stored in storage/datasets/default; run state and diagnostics are stored under storage/key_value_stores/default.
This actor has local validation evidence only; it has not been pushed or run in Apify Cloud. OUTPUT, OUTPUT_SUMMARY, RUN_HEALTH, RUN_DIAGNOSTICS, RUN_SKIPS, and SOURCE_RECEIPTS preserve run results, status, and request evidence.
Run status is only the platform outcome (SUCCEEDED or FAILED); resultStatus separately records COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED. Only explicit source-authored denial/challenge evidence is SKIPPED; generic HTTP/transport failures are deferred. Dataset row counts advance only after successful Dataset.pushData. A later non-block failure preserves already stored rows as SUCCEEDED/LIMITED; fatal zero-row errors call Actor.fail() and report FAILED/FAILED. Per-request receipts and run diagnostics remain in KVS, not job rows.
The validator reconciles dataset rows with OUTPUT, checks the schema and KVS source receipts, and rejects nulls, blank strings, placeholders, empty arrays/objects, duplicate IDs/URLs, non-Jobs.ac.uk URLs, inconsistent URL identity, records below the 21-fact threshold, and unverifiable records. Empty output is reconciled against separate platform and result statuses; prior legacy SKIPPED receipts remain readable without rewriting storage.
Bounded settings
maxItems is limited to 10, maxPages to 3, request timeout to 30 seconds, and sequential detail requests are separated by at least 250 ms. HTTP 401, 403, 429, 451, or a visible source-authored challenge stops requests and yields result SKIPPED while platform status remains SUCCEEDED; HTTP 407, other HTTP errors, and unverified transport failures yield DEFERRED. The actor does not retry or switch routes after a barrier.