Higherin Scraper
Under maintenancePricing
from $5.00 / 1,000 job scrapeds
Higherin Scraper
Under maintenanceScrape Higherin graduate jobs, internships, placements and apprenticeships. Search by keyword or collect all jobs, with optional UK visa sponsor data.
Pricing
from $5.00 / 1,000 job scrapeds
Rating
0.0
(0)
Developer
Shaheer Sarfaraz
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
15 days ago
Last modified
Categories
Share
Higherin Job Scraper: UK Graduate Jobs, Internships & Apprenticeships
An Apify Actor and TypeScript scraper for collecting structured job listings from Higherin. Scrape UK graduate schemes, internships, placements, apprenticeships, and entry-level roles into an Apify dataset with job-detail data and optional employer visa-sponsorship enrichment.
Features
- Search Higherin with a job keyword, or leave the query blank to crawl all available listings.
- Follow Higherin pagination automatically and de-duplicate job URLs.
- Visit every job detail page and extract descriptions, deadlines, job types, roles, locations, company information, salary, and structured JobPosting metadata.
- Apply fuzzy matching across listing and detail-page text after enrichment.
- Keep a listing when its detail page fails, and report failed detail URLs in the run summary.
- Optionally match employers against the GB sponsor registry using the
@dakheera47/visa-sponsor-datapackage. - Export normalized records through the Apify dataset and run statistics through
the
OUTPUTkey-value record.
Input
The Actor input is a JSON object. Both controls are available in the Apify input UI:
{"searchQuery": "software engineer","enrichSponsorship": true}
| Field | Type | Default | Description |
|---|---|---|---|
searchQuery | string | "" | Optional fuzzy search term. An empty string returns all available Higherin jobs. |
enrichSponsorship | boolean | true | Add sponsorRegistry employer matches. Set to false to skip sponsor matching and its paid event. |
Scraped job data
Every dataset item includes the original Higherin fields plus normalized fields for downstream job search, filtering, and visa-sponsorship analysis:
- Core fields:
source,sourceJobId,title,employer, andjobUrl - Listing fields:
jobTypeName,deadline,employmentStartDate,salary,salaryNotes,jobLocationNames,relevantFor, company IDs, logos, and employer badges - Detail-page fields:
jobDescription,jobRoles,locations,location,companyDescription,companyProfileUrl,datePosted,validThrough, andemploymentType - Optional sponsorship data:
sponsorRegistry, including match status, confidence, match method, registry records, and source update time
The dataset schema is defined in .actor/dataset_schema.json. The Apify input and output definitions are in .actor/input_schema.json and .actor/output_schema.json.
How the scraper works
- Fetch Higherin’s job-search JSON endpoint.
- Follow every
nextPageUrlreturned by Higherin pagination. - Normalize and de-duplicate the listings.
- Fetch each canonical job detail page and parse server-rendered HTML plus
JobPostingJSON-LD metadata. - Apply
searchQuerymatching across the completed records. - Add optional sponsor-registry matches and push the final records to Apify.
Detail-page failures do not discard the original listing. The OUTPUT record
reports the result count, failed detail count, sponsorship status, and any
sponsorship setup error.
Sponsor enrichment and pricing
Sponsor enrichment is enabled by default and is visible as the
Include sponsor enrichment control in the Apify input UI. When enabled, the
Actor initializes the GB sponsor registry matcher and pushes the enriched batch
with the sponsor-enrichment pay-per-event name.
If the pricing event is not configured or the registry cannot be initialized,
the scrape continues without sponsor data and records the reason as
sponsorshipError in OUTPUT. Set enrichSponsorship to false to avoid
matcher setup and paid enrichment entirely.
Local development
npm cinpm run checkCRAWLEE_STORAGE_DIR=/tmp/higherin-job-actor npm run start:dev
The production image is built with the included Dockerfile and runs
dist/actor.js. To deploy, connect this repository to an Apify Actor and use
the generated input, dataset, and output schemas from the .actor directory.
Related use cases
This Higherin scraper is useful for building graduate job boards, internship alerts, UK placement search tools, apprenticeship datasets, employer research pipelines, and visa-sponsorship job discovery workflows.