EuropeEnglishJobs
Pricing
from $0.01 / 1,000 results
EuropeEnglishJobs
This scraper will help in finding jobs in Europe, specially in Germany and Switzerland
Pricing
from $0.01 / 1,000 results
Rating
0.0
(0)
Developer
akaji Banpu
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 hours ago
Last modified
Categories
Share
English-speaking Jobs — Germany & Switzerland
An Apify Actor that aggregates English-speaking job listings focused on
Germany and Switzerland, plus remote/work-from-anywhere jobs regardless of
country (Netherlands is still supported via countries, just no longer the
primary focus).
Sources
Free, direct API calls (no key needed, no separate Actor cost, no known legal/ToS issue with this kind of use):
- Arbeitnow
- Remotive
- Jobicy — supports Germany, Switzerland, Netherlands via
geo=. Does not use Jobicy's owntagparam for keyword filtering — that param only accepts values from Jobicy's fixed taxonomy (e.g. "python"), and an arbitrary keyword like "QA" causes a 400 Bad Request. Filtering is done entirely client-side against the title instead, like the other sources. - RemoteOK — 100%-remote listings. Its own API response embeds an explicit terms clause: "please link back... and mention Remote OK as a source, so we get traffic back from your site. If you do not we'll have to suspend API access." If you build anything user-facing on top of this Actor's output, credit RemoteOK and link back to the job's URL on remoteok.com.
- We Work Remotely — 100%-remote listings via the RSS feed the site itself publishes for external consumption.
- Adzuna (optional — needs a free
adzunaAppId/adzunaAppKey) — officially supports Germany (de), Switzerland (ch), and Netherlands (nl) - Himalayas — free public JSON API, no key, documented, supports a
countryfilter - The Muse — free without a key (500 req/hr); its own
locationquery param didn't reliably narrow results in testing, so this fetches broadly and filters client-side like Arbeitnow - Jooble (optional — needs a free
joobleApiKeyfrom a signup form) — legitimate free API meant for exactly this use, but the free tier is a lifetime cap of 500 requests total (not per-day), so it's opt-in and should be used sparingly
RemoteOK, We Work Remotely, Himalayas, and The Muse are all global, English-by-default remote/job boards, so — unlike Arbeitnow, a German board mixing German/English postings — their results are not run through the English-language filter; nearly everything on them already qualifies, and the filter would mostly produce false negatives.
Keyword matching is title-only, not title+description. Found via
testing: Remotive (via Lemon.io postings) and other sources reuse the same
generic marketing boilerplate paragraph across every job listing
regardless of actual role — a "QA" search matched "Senior Golang
Developer" and other unrelated roles because that shared boilerplate text
happened to mention "qa automation" once, deep in the description. None of
the upstream APIs' own keyword/search/tag params were found to reliably
filter either (Himalayas' keyword and Remotive's search silently
returned everything; Jobicy's tag outright 400s on non-taxonomy values) —
every free source's keyword filter is applied client-side, against the
title only, as the actual source of truth.
One official option investigated but not added: EURES, the EU/EFTA's own job mobility network (Switzerland is a full member via a bilateral agreement, confirmed). It has no sanctioned public API — only a reverse-engineered, community-documented endpoint exists, which could change or break without notice since it isn't actually meant for third-party use. Worth revisiting if you want the reach of an official government jobs database and are OK with that fragility.
Composed from existing Apify Store Actors (each is a separate, pay-per-result Actor run):
- LinkedIn — default
curious_coder/linkedin-jobs-scraper - Indeed — default
borderline/indeed-scraper - Xing — default
fatihtahta/xing-jobs-scraper - StepStone — default
trev0n/stepstone-scraper— covers Germany, Austria, Netherlands, and Belgium (only DE/NL are in ourcountrieslist; Switzerland is skipped, StepStone doesn't operate there). Filters for English-language ads directly via its ownadLanguageparam, so this source's results are more precisely English than the others in this section. Verified against a real run —title/location/url/companyNameall confirmed correct. - Robert Half — default
studio-amba/roberthalf-scraper— a staffing agency covering Germany, Netherlands, and Switzerland (via its English-languageCH-enlocale). Unverified — every test attempt (from two different Apify accounts) was rejected pre-flight with "exceeds remaining usage," seemingly due to this Actor's own Store pricing configuration rather than the caller's resource request. Verify before relying on it. - Relocate.me (visa-sponsored jobs) — default
khadinakbar/visa-sponsored-jobs-scraper— international tech roles with explicit visa sponsorship and a full relocation package (visa services, flights, housing, language courses); covers Germany and Netherlands, not Switzerland. Every listing from this source is inherently visa-sponsored, so it's tagged as"Relocate.me (Visa Sponsored)"in thesourcefield rather than adding a dedicated schema field for one source. Verified against a real run — note there's no singlelocationfield, it's split into separatecity/countryfields (handled in code).
Why compose existing Actors instead of scraping these sites directly
LinkedIn, Indeed, Xing, StepStone, Robert Half, and Relocate.me all prohibit automated scraping in their Terms of Service and run active anti-bot defenses (LinkedIn in particular has pursued legal action against scrapers). Rather than writing bespoke scraping logic against those defenses in this repo,
src/sources/{linkedin,indeed,xing,stepstone, robert_half,visa_sponsored}.pyActor.call(...)) and read
its output dataset. That plumbing (proxies, anti-bot handling,
markup-change maintenance) already lives in those Actors. You are still
bound by Apify's own platform terms and by whatever terms the upstream
Actor's author has agreed to — this doesn't remove the underlying legal
risk, it just avoids re-implementing scraping logic that the target sites
explicitly forbid. These six sources are excluded from the default
sources list precisely because of that residual legal exposure — the
free/direct-API sources above are the default, ToS-clean set.
You can swap in a different Actor for any of these via the
linkedinActorId / indeedActorId / xingActorId / stepstoneActorId /
robertHalfActorId / visaSponsoredActorId input fields.
Robert Half and Relocate.me both have small user counts on Apify (2 and 81 respectively) — lower track record than the others in this list, so treat their output field names as even less certain until you've run a real test.
Every composed-Actor call passes explicit memory_mbytes=512/
timeout_secs=180 (see src/sources/apify_actor.py) so cost stays bounded
regardless of the called Actor's own resource defaults — found necessary
during testing when one Actor's default allocation alone made Apify's
pre-flight cost estimate exceed a small account's remaining monthly free
credit, even before any real per-result charges applied. Each source is
also isolated in main.py (_safe() wrapper) so one failing composed
Actor — insufficient balance, the Actor being down, a bad input field —
can never discard the other sources' already-collected results; a failure
is logged and that source just comes back empty.
Two boards considered but not added: Glassdoor and Monster both have
usable Apify Store Actors, but weren't wired in this round — Monster in
particular describes itself as US-focused, a weak fit for Germany/Switzerland.
Add them the same way (a new src/sources/<site>.py following the
stepstone.py pattern) if you want them later.
Four expat/English-specific job sites were also investigated (iamexpat.nl, undutchables.nl, thelocal.de/jobs, berlinstartupjobs.com) as candidates for a custom scraper. Only undutchables.nl came back clear (no scraping clause in its ToS, permissive robots.txt); the other three either explicitly prohibit scraping in their terms or have no real job-listing surface left to scrape. None are implemented yet — see git history for the full per-site findings if you want to revisit undutchables.nl.
English-language filtering for the Actor-composed sources: LinkedIn,
Indeed, Xing, StepStone, and Robert Half are all run through
looks_english_language (src/models.py) — real language detection via
langdetect against each listing's description, not a keyword-presence
check. This distinction matters: an earlier version reused
looks_english_speaking (which only checks whether text explicitly says
"English", the right check for Arbeitnow where most postings are German
and we want the subset that calls out English-speaking teams) for these
sources too — but a normal English LinkedIn posting essentially never
contains the literal word "English", so that check would have wrongly
rejected good English listings, not just filtered German ones. Found via a
real test: a LinkedIn run with no keywords returned plainly German-only
postings ("Dachdecker (m/w/d) in Vollzeit gesucht!") completely unfiltered,
since LinkedIn has no server-side language param at all (unlike StepStone's
adLanguage: "en", which is now double-checked with the same client-side
detector rather than trusted blindly — consistent with this project's
"verify server-side filters, don't trust them" pattern). Relocate.me is the
one exception, left unfiltered as an English-by-default international
board, same reasoning as RemoteOK/We Work Remotely/Himalayas/The Muse.
Detection needs a reasonably long description to be reliable (job titles
alone are too short — verified langdetect misclassifying "DevOps
Engineer" as Dutch and "QA" as Vietnamese), so text under 40 characters is
never filtered out rather than risk a false rejection.
CV-based job scoring
Both keywords and CV scoring are entirely optional and independent — you
can search with neither (broad fetch), either one alone, or both together
(CV scoring applies on top of whatever keywords/countries/sources
already narrowed down to).
If keywords is empty and a CV is provided, the CV's own top skills are
now used to drive the actual search (cvAutoKeywordCount, default 5) —
not just to re-rank results afterward. This was a real, confirmed gap:
running with a CV but no keywords used to fetch each source's generic
default feed (LinkedIn/StepStone with no keyword returned a roofer, a
florist, "Product Owner" — nothing security/engineering-related) and rely
on match_score to sort that essentially-random set, producing near-zero
scores and irrelevant top results no matter how good the scoring math was.
The CV was never actually driving what got searched for. Now it is.
Provide a CV via cvText (paste plain text) or cvFile (upload a PDF,
DOCX, or plain-text file) and every job gets a match_score (0-100),
matched_keywords, and cv_suggestions, and the results are sorted
best-match-first.
Scoring is free deterministic matching, not AI/semantic matching — a
deliberate trade-off (see src/cv_matching.py):
- Extract the ~40 most frequent non-stopword words from the CV text (a large, CV/job-posting-specific stopword list filters out both normal grammar words and generic resume boilerplate like "experience", "team", "senior", "led" — otherwise scores are dominated by words that appear in almost every job posting regardless of actual role).
- For each job, extract its OWN significant terms the same way — title terms weighted higher than description terms (more central to what the role actually is). Score = what fraction of the JOB's own terms the CV covers — framed this way deliberately, not "what fraction of the CV the job uses": a CV naturally lists far more skills than any single job description mentions, so using the CV's full keyword set as the denominator was tried first and produced misleadingly low scores even for excellent matches (the best real match in testing scored 24/100). Scoring against the job's own requirements instead gives a number that means what people actually expect "match %" to mean.
- A seniority-fit adjustment: years-of-experience extracted from the CV (regex for "N years... experience") is mapped to a 0-5 level (intern through director) and compared against the same scale detected from the job's title/description — a close fit adds a bonus, a big mismatch (e.g. a senior candidate vs. an internship) subtracts a penalty.
cv_suggestions: the job's own significant terms that don't appear anywhere in the CV, listed as "consider adding this if you genuinely have that experience" — an honest, computable stand-in for AI-generated ATS advice, not a claim of understanding what "ATS-friendly formatting" means beyond keyword coverage.
This won't catch synonyms ("k8s" vs "kubernetes") or truly reason about experience the way an LLM-based matcher would — an explicit choice over adding an AI API dependency/cost (asked directly, chose the free deterministic version). Verified locally and live on the deployed Actor against a sample security-analyst CV: "Senior Penetration Tester" scored 73, "Junior QA Tester" scored 14, "Florist" scored 0.
Jobs requiring German language are always excluded entirely (not just
scored lower), regardless of whether a CV is provided —
models.requires_german matches phrasings like "fluent in German",
"verhandlungssicheres Deutsch", "German C1", "Deutschkenntnisse
erforderlich" against title+description, deliberately excluding "nice to
have" framings ("German is a plus") so only genuinely mandatory
requirements are filtered. This is distinct from the English-language
detection elsewhere (looks_english_language) — a posting can be written
entirely in English and still require German as a stated prerequisite,
which written-language detection alone can't catch. Verified against 13
realistic phrasings (7 that should filter, 6 that shouldn't) before
deploying.
Verification status of the two CV input paths:
cvText— fully verified end-to-end, live on the deployed Actor.cvFile— the extraction logic (PDF/DOCX/plain-text parsing, and reading an Apify key-value-store URL with the run's own authenticated credentials) is verified in isolation, but Apify's own input-schema docs don't actually specify what URL format thefileuploadeditor hands to the Actor at runtime ("up to the Actor developer to interpret"). A test against a manually-created key-value-store record correctly failed with a permissions error (the run is sandboxed to its own storages underLIMITED_PERMISSIONS) — consistent with the real Console upload widget likely storing the file in the run's own default store, which this code should then read correctly, but that specific path needs one real test: upload an actual file via the Console's Input tab and run it.
Note: the description field (used as scoring context) is only populated
for sources whose API/Actor actually exposes one — empty for others, so
scoring for those falls back to title-only.
Security & Privacy
Three legitimate concerns worth addressing directly, with evidence rather than just reassurance — check the referenced line numbers yourself.
Does the CV (cvText/cvFile) get sent to LinkedIn/Indeed/Xing/StepStone/
Robert Half/Relocate.me, or anywhere else? No, with one precise nuance.
cv_matching.py (regex + word counting, no network calls with the CV
content) is the only code that ever touches the CV text or file — none of
the six composed-Actor fetch() functions accept a CV parameter at all,
check their signatures in
src/sources/{linkedin,indeed,xing,stepstone, robert_half,visa_sponsored}.py(actor, actor_id, keyword, countries, max_items, log)keywords is empty and a CV is
provided, individual extracted keywords (single common words like
"python", "security" — never the CV text itself) are used as the search
keyword for those same calls, functionally identical to a user typing
those words into keywords by hand. The full CV text/file never reaches
any composed Actor under any configuration. Separately, the extracted
keywords (occasionally a name fragment) are written to the
run's own log via Actor.log.info in main.py — visible only to you as
the account owner in Apify Console, never sent anywhere else. Remove that
log line if you want it tighter.
Do the Adzuna/Jooble API keys go anywhere besides Adzuna/Jooble? No —
app_id/app_key are sent only as query params to api.adzuna.com
(fetch_adzuna in free_apis.py); api_key is used only in the URL path
for jooble.org/api/{key} (fetch_jooble). Never logged, never passed to
any composed Actor. Both are marked isSecret: true in the input schema,
so Apify encrypts and masks them.
Can the *ActorId override fields be used to run arbitrary code with my
credentials? This Actor declares LIMITED_PERMISSIONS (visible in every
run's log: ACTOR: Running under "LIMITED_PERMISSIONS"). Per
Apify's own permissions docs:
a limited-permission Actor "can only call other Actors that also have
limited permissions," each called Actor "receives a restricted token" (not
a copy of your account's real credentials), and it "can't access any other
data in your Apify account" beyond its own storages and whatever's
explicitly passed to it. What we explicitly pass to a composed Actor is
just keyword/country/maxItems (see the fetch() signatures above) —
never your CV, never your Adzuna/Jooble keys. The real residual risk from
pointing an override at an untrusted Actor ID is narrower than credential
theft: it would run under your account's billing/budget (confirmed
directly during testing — a composed-Actor call failed with "you will
exceed your remaining usage" when the account's balance was too low), and
it would receive that same plain keyword/country/maxItems input. It cannot
reach your CV, your other API keys, or anything else in your account.
Input
See .actor/input_schema.json. Key fields:
| Field | Description |
|---|---|
keywords | Array of job title/keyword filters, e.g. ["QA", "security analyst"]. Optional — leave empty to fetch broadly with no keyword filter, unless a CV is provided (see below). Each keyword is searched separately across every enabled source and results are merged/deduplicated. Matches against job title only (not description) — see below for why. linkedin/indeed/xing/stepstone/roberthalf/visasponsored each run once per keyword, so their cost multiplies with the number of keywords |
cvText / cvFile | Optional. See "CV-based job scoring" above |
cvAutoKeywordCount | Default 5 — when keywords is empty and a CV is provided, how many of the CV's top skills to actually search for instead of fetching each source's generic feed |
countries | ["germany", "switzerland"] by default; "netherlands" also supported |
remoteAnywhere | Default true — also include jobs tagged Worldwide/Anywhere/Europe/EMEA/Flexible regardless of countries, from arbeitnow/remotive/remoteok/weworkremotely/himalayas/themuse |
sources | Which sources to query. Default is the free, ToS-clean set: arbeitnow, remotive, jobicy, remoteok, weworkremotely, himalayas, themuse. jooble is free but opt-in (signup + 500-request lifetime cap). linkedin/indeed/xing/stepstone are opt-in and cost money per run — see the legal note above before enabling |
maxItemsPerSource | Cap on results per paid Actor — keep this low while testing |
adzunaAppId / adzunaAppKey | Only needed if adzuna is in sources |
joobleApiKey | Only needed if jooble is in sources |
linkedinActorId / indeedActorId / xingActorId / stepstoneActorId / robertHalfActorId / visaSponsoredActorId | Override which Store Actor to call |
Local development
pip install -r requirements.txtapify run
apify run needs the Apify CLI installed and
an APIFY_TOKEN configured (apify login) so it can call the LinkedIn/
Indeed/Xing/StepStone/Robert Half/Relocate.me sub-Actors if those sources
are enabled. The free sources (the default) work without a token.
Before enabling linkedin/indeed/xing/stepstone/roberthalf/
visasponsored: verify the output field names assumed in the corresponding
src/sources/*.py module (pick(item, "title", ...) etc.) against a small
test run of each upstream Actor in the Apify Console — Store search doesn't
expose output schemas, only input schemas, so those were written
defensively but unverified against live output.
Deploy to Apify
apify loginapify push