jobs.cz Jobs Scraper avatar

jobs.cz Jobs Scraper

Pricing

from $0.56 / 1,000 results

Go to Apify Store
jobs.cz Jobs Scraper

jobs.cz Jobs Scraper

Czech job listings from jobs.cz: search by keyword and locality, with full job details on request — description, required education and languages, employment type, contract length and contact person.

Pricing

from $0.56 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Scrapes job listings from jobs.cz, one of the Czech Republic's largest job boards. Search by keyword and locality, then optionally pull the full job detail: description, required education and languages, employment type, contract length and contact person.


What you get

One JOB row per posting, plus a SEARCH_SUMMARY row per query.

From the search page (always)

jobId · title · url · companyName · locality

From the detail page (includeJobDetails, on by default — read the limit below first)

description (full job ad text) · infoItems (all the label/value pairs jobs.cz itself shows — required education, languages, employment type, contract length, contract form, and whatever else the posting includes) · locationText (full street address) · contactPerson · breadcrumbs (category path)


Input

{
"queries": [
{ "term": "programátor", "locality": "praha" }
],
"includeJobDetails": true,
"maxItemsPerQuery": 50
}

Each entry needs a term; locality is optional (omit it for a nationwide search).

One term per query, not a list — here's why

jobs.cz's own multi-keyword parameter combines terms as OR, not AND: measured in Prague, programátor alone → 494 results, senior alone → 686, both together → 1,099 (close to the sum, not an intersection). A list input here would silently mean "any of these keywords" when "all of these" is the natural reading — refused by design. Run separate entries in queries for separate searches instead, each with its own honest total.

locality is free text, and that's safe here

Real values are Czech city/region URL slugs (praha, brno, ostrava, plzen, and others not listed). There's no lookup endpoint to validate against — but jobs.cz is honest about a miss: an unrecognised slug answers a clean HTTP 404, not a silent fallback to a nationwide baseline.


Read this before turning on includeJobDetails

A large share of postings from bigger employers cannot be enriched at all — by design of the site, not a bug here. Some employers' job ads redirect from the standard jobs.cz/rpd/{id}/ detail URL to a company-branded subdomain ({company}.jobs.cz) — a completely different page template with zero markup overlap with the standard one (no title heading, no structured fields of any kind). Measured on one small sample: 4 of 5 detail fetches for a Prague "programátor" search hit this (ČSOB, Česká spořitelna, and others). Those rows keep every field the search page already gave them (title, company, locality, url) and carry a specific _detailError naming the branded subdomain they redirected to — never a silent blank. includeJobDetails is still worth turning on: it's not an all-or-nothing switch, just don't expect 100% enrichment coverage on a batch that includes large employers.


Read this before trusting a large maxItemsPerQuery

Pagination clamps to the LAST real page past the end — the opposite direction from a related actor in this same portfolio (jobup-jobs-scraper, which clamps to page 1). Past the true final page, jobs.cz re-serves that last page's exact result list rather than an empty page. The actor's stop condition is "this page contributed zero new job ids", never "the page was empty", so this is handled correctly regardless of which direction the clamp goes — and the summary row reports paginationClampHit: true when it happens.


Known limits

1. Company-branded detail pages are not parsed (see above) — the single biggest limitation on this target.

2. The total-count text uses a non-breaking space as its thousands separator ("Našli jsme 1 099 nabídek"), not a comma — handled internally, mentioned here only because it's the kind of thing that silently breaks a naive \d+ regex if you're parsing this site yourself.


Anti-bot posture

None observed. 8/8 TLS profiles clean on both the search and detail surfaces during recon. The actor still ships a rotating profile pool, a retry ladder and challenge-marker detection, and paces requests by default — a target being clean in local recon has not always meant the same from Apify's cloud egress elsewhere in this portfolio.

Policy

robots.txt (User-agent: *) names no ClaudeBot/anthropic-ai/GPTBot group. It disallows /api/, /iapi/, /muj/ (account area), /asmt/, /status/, /translations/, /js/, /nabidky-podle-cv/, /session-log/, and two tokenised query strings on unrelated paths — none of which overlap /prace/ (search) or /rpd/ (detail), the two surfaces this actor fetches.