Higherin Scraper avatar

Higherin Scraper

Under maintenance

Pricing

from $5.00 / 1,000 job scrapeds

Go to Apify Store
Higherin Scraper

Higherin Scraper

Under maintenance

Scrape Higherin graduate jobs, internships, placements and apprenticeships. Search by keyword or collect all jobs, with optional UK visa sponsor data.

Pricing

from $5.00 / 1,000 job scrapeds

Rating

0.0

(0)

Developer

Shaheer Sarfaraz

Shaheer Sarfaraz

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

15 days ago

Last modified

Categories

Share

Higherin Job Scraper: UK Graduate Jobs, Internships & Apprenticeships

An Apify Actor and TypeScript scraper for collecting structured job listings from Higherin. Scrape UK graduate schemes, internships, placements, apprenticeships, and entry-level roles into an Apify dataset with job-detail data and optional employer visa-sponsorship enrichment.

Features

  • Search Higherin with a job keyword, or leave the query blank to crawl all available listings.
  • Follow Higherin pagination automatically and de-duplicate job URLs.
  • Visit every job detail page and extract descriptions, deadlines, job types, roles, locations, company information, salary, and structured JobPosting metadata.
  • Apply fuzzy matching across listing and detail-page text after enrichment.
  • Keep a listing when its detail page fails, and report failed detail URLs in the run summary.
  • Optionally match employers against the GB sponsor registry using the @dakheera47/visa-sponsor-data package.
  • Export normalized records through the Apify dataset and run statistics through the OUTPUT key-value record.

Input

The Actor input is a JSON object. Both controls are available in the Apify input UI:

{
"searchQuery": "software engineer",
"enrichSponsorship": true
}
FieldTypeDefaultDescription
searchQuerystring""Optional fuzzy search term. An empty string returns all available Higherin jobs.
enrichSponsorshipbooleantrueAdd sponsorRegistry employer matches. Set to false to skip sponsor matching and its paid event.

Scraped job data

Every dataset item includes the original Higherin fields plus normalized fields for downstream job search, filtering, and visa-sponsorship analysis:

  • Core fields: source, sourceJobId, title, employer, and jobUrl
  • Listing fields: jobTypeName, deadline, employmentStartDate, salary, salaryNotes, jobLocationNames, relevantFor, company IDs, logos, and employer badges
  • Detail-page fields: jobDescription, jobRoles, locations, location, companyDescription, companyProfileUrl, datePosted, validThrough, and employmentType
  • Optional sponsorship data: sponsorRegistry, including match status, confidence, match method, registry records, and source update time

The dataset schema is defined in .actor/dataset_schema.json. The Apify input and output definitions are in .actor/input_schema.json and .actor/output_schema.json.

How the scraper works

  1. Fetch Higherin’s job-search JSON endpoint.
  2. Follow every nextPageUrl returned by Higherin pagination.
  3. Normalize and de-duplicate the listings.
  4. Fetch each canonical job detail page and parse server-rendered HTML plus JobPosting JSON-LD metadata.
  5. Apply searchQuery matching across the completed records.
  6. Add optional sponsor-registry matches and push the final records to Apify.

Detail-page failures do not discard the original listing. The OUTPUT record reports the result count, failed detail count, sponsorship status, and any sponsorship setup error.

Sponsor enrichment is enabled by default and is visible as the Include sponsor enrichment control in the Apify input UI. When enabled, the Actor initializes the GB sponsor registry matcher and pushes the enriched batch with the sponsor-enrichment pay-per-event name.

If the pricing event is not configured or the registry cannot be initialized, the scrape continues without sponsor data and records the reason as sponsorshipError in OUTPUT. Set enrichSponsorship to false to avoid matcher setup and paid enrichment entirely.

Local development

npm ci
npm run check
CRAWLEE_STORAGE_DIR=/tmp/higherin-job-actor npm run start:dev

The production image is built with the included Dockerfile and runs dist/actor.js. To deploy, connect this repository to an Apify Actor and use the generated input, dataset, and output schemas from the .actor directory.

This Higherin scraper is useful for building graduate job boards, internship alerts, UK placement search tools, apprenticeship datasets, employer research pipelines, and visa-sponsorship job discovery workflows.