Jobs.ac.uk Search Scraper avatar

Jobs.ac.uk Search Scraper

Pricing

from $2.99 / 1,000 jobs.ac.uk job records

Go to Apify Store
Jobs.ac.uk Search Scraper

Jobs.ac.uk Search Scraper

Scrape job listings from Jobs.ac.uk, the leading UK job board for academic, research, and higher education positions. Extract job titles, institutions, locations, salary ranges, contract types, and descriptions for academic recruitment.

Pricing

from $2.99 / 1,000 jobs.ac.uk job records

Rating

0.0

(0)

Developer

Jobs API

Jobs API

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

11 days ago

Last modified

Categories

Share

Jobs.ac.uk Jobs Search Scraper

This Actor performs bounded extraction of public Jobs.ac.uk academic job pages using ordinary direct HTTPS requests and Cheerio parsing. Requests are sequential, have no retry, custom user-agent, proxy, or fingerprint controls, and stop on a public access barrier. It does not sign in, submit applications, or fetch external application destinations.

Modes

  • search discovers public Jobs.ac.uk vacancies through the official search page, then fetches each detail page.
  • single fetches one URL from jobUrl, directUrl, or startUrl.
  • multiple fetches URLs from jobUrls, directUrls, or startUrls.

Every emitted row is buffered until its detail page, canonical URL, job identity, description, and required fields are verified. Incomplete or mismatched jobs are skipped in search mode; transport, access, and source-integrity failures write structured diagnostics to the key-value store and emit no misleading row.

Output

Rows preserve the public academic vacancy data available on the page: title, institution, department, public advert categories (role, subject areas, and location areas), full text and HTML description, sectioned responsibilities/qualifications/benefits, locations, salary, hours, contract type, posting and closing dates, job reference, structured JobPosting data, description links, contact emails, explicit external application links, source listings, request receipts, field coverage, and execution-quality evidence. The dataset accepts a record only when at least 21 distinct meaningful source facts are populated; the count excludes identifiers, repeated aliases, URLs used only for provenance, and runtime metadata.

applicationUrl is emitted only when Jobs.ac.uk publishes an explicit application action. The canonical Jobs.ac.uk vacancy URL is never presented as an application link.

Local run

npm install
$env:APIFY_LOCAL_STORAGE_DIR = 'storage-validation-20260927-01'
apify run --resurrect --input-file INPUT.json
npm run validate -- $env:APIFY_LOCAL_STORAGE_DIR

The reproducible fixtures INPUT.json, INPUT-single.json, and INPUT-multiple.json exercise all three modes. The dataset is stored in storage/datasets/default; run state and diagnostics are stored under storage/key_value_stores/default.

This actor has local validation evidence only; it has not been pushed or run in Apify Cloud. OUTPUT, OUTPUT_SUMMARY, RUN_HEALTH, RUN_DIAGNOSTICS, RUN_SKIPS, and SOURCE_RECEIPTS preserve run results, status, and request evidence.

Run status is only the platform outcome (SUCCEEDED or FAILED); resultStatus separately records COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED. Only explicit source-authored denial/challenge evidence is SKIPPED; generic HTTP/transport failures are deferred. Dataset row counts advance only after successful Dataset.pushData. A later non-block failure preserves already stored rows as SUCCEEDED/LIMITED; fatal zero-row errors call Actor.fail() and report FAILED/FAILED. Per-request receipts and run diagnostics remain in KVS, not job rows.

The validator reconciles dataset rows with OUTPUT, checks the schema and KVS source receipts, and rejects nulls, blank strings, placeholders, empty arrays/objects, duplicate IDs/URLs, non-Jobs.ac.uk URLs, inconsistent URL identity, records below the 21-fact threshold, and unverifiable records. Empty output is reconciled against separate platform and result statuses; prior legacy SKIPPED receipts remain readable without rewriting storage.

Bounded settings

maxItems is limited to 10, maxPages to 3, request timeout to 30 seconds, and sequential detail requests are separated by at least 250 ms. HTTP 401, 403, 429, 451, or a visible source-authored challenge stops requests and yields result SKIPPED while platform status remains SUCCEEDED; HTTP 407, other HTTP errors, and unverified transport failures yield DEFERRED. The actor does not retry or switch routes after a barrier.