CV-Library Jobs Scraper
Pricing
from $0.35 / 1,000 results
CV-Library Jobs Scraper
Scrapes job listings from CV-Library, one of the UK's largest job boards. Search any keyword and location; returns title, company, salary range, category, contract type and full description from a single search call.
Pricing
from $0.35 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
CV-Library Jobs Scraper (UK)
Scrapes job listings from CV-Library —
one of the UK's largest job boards. A second UK job-search actor in this
portfolio alongside reed-jobs-scraper (general listings) and
efinancialcareers-jobs-scraper (finance-only) — distinct platform,
distinct customer base.
Public data only. No login, no cookies, no browser.
The one thing you need to know before using this
CV-Library's robots.txt disallows every query string except two
explicit carve-outs (?jobId= and ?page_number=) — the site's own
human-facing search URL (?q=...&geo=...) is not robots-permitted for a
generic client. This actor uses the equivalent path-based search URL
instead (/{keyword}-jobs-in-{location}), which the site itself also
serves and which carries no matching disallow rule. See
CRAWLING_METHOD.md §3 for the full policy
reasoning.
Separately, 3 of 7 TLS profiles tried in recon get a plain HTTP 403 (not a disguised-200 challenge) — this actor's client pool keeps only the 4 confirmed-clean profiles.
What you get
Three record types share one dataset, told apart by recordType.
JOB — one row per listing
Search rows (listing) already carry title, employer, salary range,
location, category, contract type and a highlighted description snippet.
Turn on Fetch job detail pages (on by default) to also attach
jobDetail, the Schema.org JobPosting block with the full, untruncated
description.
SEARCH_SUMMARY — one row per (keyword, location) query
CV-Library's own real match count (totalMatches — a genuine structured
field, not a heuristic), pages fetched, and locationApplied — whether
the requested location actually narrowed the search upstream.
ERROR — one row per input that failed
So every entry in Searches maps to at least one output row.
Input
| Field | What it does |
|---|---|
| Searches | list of {keyword, location} — location is optional |
| Fetch job detail pages | adds the full untruncated description (on by default) |
| Max jobs / max pages per search | pagination caps — this actor computes the exact page count from CV-Library's own total, these are extra safety valves |
| Max concurrent requests / Min seconds between requests | standard pacing controls |
Example
{"queries": [{"keyword": "developer", "location": "manchester"},{"keyword": "nurse"}],"includeJobDetails": true,"maxItems": 100}
Notes on reliability
- Location can silently widen upstream, keyword cannot. An
unrecognised
locationdoes not error — CV-Library quietly falls back to the keyword-only national result set. This actor detects that (the location metadata upstream would otherwise return is simply absent) and reportslocationApplied: falsewith zero rows rather than shipping mislabeled data. An unrecognisedkeyword, in contrast, answers a clean HTTP 404 and is reported as anot_founderror — the two behave oppositely, so both are handled explicitly. totalMatchesis a real structured field, not a page-copy estimate — confirmed live to change correctly per keyword and per location.- A de-listed/expired job answers a clean HTTP 404 — the search row is
still emitted, with
jobDetail: null.
Output envelope
Every record carries _input, _source and _scrapedAt. Upstream field
names pass through verbatim under listing (and jobDetail when
requested) — no renaming, and the <mark> highlight tags CV-Library wraps
around matched keywords in the search snippet are left intact rather than
stripped.
See CRAWLING_METHOD.md for the full reverse-engineering trail, including the robots.txt policy reasoning and what was NOT verified this session (category/salary/contract-type filtering).