jobs.cz Jobs Scraper
Pricing
from $0.56 / 1,000 results
jobs.cz Jobs Scraper
Czech job listings from jobs.cz: search by keyword and locality, with full job details on request — description, required education and languages, employment type, contract length and contact person.
Pricing
from $0.56 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Scrapes job listings from jobs.cz, one of the Czech Republic's largest job boards. Search by keyword and locality, then optionally pull the full job detail: description, required education and languages, employment type, contract length and contact person.
What you get
One JOB row per posting, plus a SEARCH_SUMMARY row per query.
From the search page (always)
jobId · title · url · companyName · locality
From the detail page (includeJobDetails, on by default — read the limit below first)
description (full job ad text) · infoItems (all the label/value pairs
jobs.cz itself shows — required education, languages, employment type,
contract length, contract form, and whatever else the posting includes) ·
locationText (full street address) · contactPerson · breadcrumbs
(category path)
Input
{"queries": [{ "term": "programátor", "locality": "praha" }],"includeJobDetails": true,"maxItemsPerQuery": 50}
Each entry needs a term; locality is optional (omit it for a
nationwide search).
One term per query, not a list — here's why
jobs.cz's own multi-keyword parameter combines terms as OR, not AND:
measured in Prague, programátor alone → 494 results, senior alone →
686, both together → 1,099 (close to the sum, not an intersection). A
list input here would silently mean "any of these keywords" when "all of
these" is the natural reading — refused by design. Run separate entries in
queries for separate searches instead, each with its own honest total.
locality is free text, and that's safe here
Real values are Czech city/region URL slugs (praha, brno, ostrava,
plzen, and others not listed). There's no lookup endpoint to validate
against — but jobs.cz is honest about a miss: an unrecognised slug answers
a clean HTTP 404, not a silent fallback to a nationwide baseline.
Read this before turning on includeJobDetails
A large share of postings from bigger employers cannot be enriched at
all — by design of the site, not a bug here. Some employers' job ads
redirect from the standard jobs.cz/rpd/{id}/ detail URL to a
company-branded subdomain ({company}.jobs.cz) — a completely different
page template with zero markup overlap with the standard one (no title
heading, no structured fields of any kind). Measured on one small sample:
4 of 5 detail fetches for a Prague "programátor" search hit this
(ČSOB, Česká spořitelna, and others). Those rows keep every field the
search page already gave them (title, company, locality, url) and carry a
specific _detailError naming the branded subdomain they redirected to —
never a silent blank. includeJobDetails is still worth turning on: it's
not an all-or-nothing switch, just don't expect 100% enrichment coverage
on a batch that includes large employers.
Read this before trusting a large maxItemsPerQuery
Pagination clamps to the LAST real page past the end — the opposite
direction from a related actor in this same portfolio
(jobup-jobs-scraper, which clamps to page 1). Past the true final page,
jobs.cz re-serves that last page's exact result list rather than an empty
page. The actor's stop condition is "this page contributed zero new job
ids", never "the page was empty", so this is handled correctly regardless
of which direction the clamp goes — and the summary row reports
paginationClampHit: true when it happens.
Known limits
1. Company-branded detail pages are not parsed (see above) — the single biggest limitation on this target.
2. The total-count text uses a non-breaking space as its thousands
separator ("Našli jsme 1 099 nabídek"), not a comma — handled internally,
mentioned here only because it's the kind of thing that silently breaks a
naive \d+ regex if you're parsing this site yourself.
Anti-bot posture
None observed. 8/8 TLS profiles clean on both the search and detail surfaces during recon. The actor still ships a rotating profile pool, a retry ladder and challenge-marker detection, and paces requests by default — a target being clean in local recon has not always meant the same from Apify's cloud egress elsewhere in this portfolio.
Policy
robots.txt (User-agent: *) names no ClaudeBot/anthropic-ai/GPTBot
group. It disallows /api/, /iapi/, /muj/ (account area), /asmt/,
/status/, /translations/, /js/, /nabidky-podle-cv/,
/session-log/, and two tokenised query strings on unrelated paths — none
of which overlap /prace/ (search) or /rpd/ (detail), the two surfaces
this actor fetches.