StepStone.de Jobs Scraper — German Job Listings & Leads avatar

StepStone.de Jobs Scraper — German Job Listings & Leads

Pricing

$0.50 / 1,000 per job listing returneds

Go to Apify Store
StepStone.de Jobs Scraper — German Job Listings & Leads

StepStone.de Jobs Scraper — German Job Listings & Leads

From $0.50 per 1,000 listings. Scrape StepStone.de by keyword, city, Bundesland, radius, sector, discipline, contract type, career level, home-office & date: employer name, company page, location, posted date, snippet, apply route. Incremental 'new mandates only' mode with Slack/webhook alerts.

Pricing

$0.50 / 1,000 per job listing returneds

Rating

0.0

(0)

Developer

Scrapers Delight

Scrapers Delight

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

🇩🇪 StepStone.de Jobs Scraper — German Job Listings, Employers & Hiring Leads

Scrape StepStone.de — Germany's largest job board — by keyword, city, Bundesland, radius, sector, discipline, contract type, career level, home-office and date posted. Every row is one live vacancy: the employer, their StepStone company page, the location, the posting date and a description snippet. 20+ server-side filters, an incremental "new mandates only" monitor, and one flat price of $0.50 per 1,000 listings.

Why this one?

This actorTypical StepStone scraper
Price per 1,000 listings$0.50$1.00 – $3.00
Actor-start feenone$0.005 – $0.09 per run
Server-side filters20+ (sector, discipline, career level, application method, advert language, skills, region, city…)keyword + location
Proves the filters bound✅ reads StepStone's own categorization echo back on every search and every route (keyword, company, pasted URL) and fails the run rather than bill you for rows that don't match your search; on a zero-row page it names the unbound filter instead of blaming StepStone
Duplicates✅ collapsed on StepStone's listing id before billing — 250 raw → 100 delivered, 150 collapsed on an overlapping 4-query run (run JcpJGdYucxuFcbd6g, 2026-09-06)often billed twice
StepStone's "similar jobs" padding✅ dropped by default, so a narrow search can't bill you 24 non-matching rowsbilled as results
Staffing-agency noise (Arbeitnehmerüberlassung)✅ removed server-side, so you don't pay for competitors reposting mandates
Incremental "new listings only" mode✅ named key-value store that survives between scheduled runs
Employer-level sweep✅ every open role at one company, by id or company URL
Salary claimsand we say so — see Honest limitssome claim it; the field is empty on this surface

No login, no API key, no browser. The actor reads the structured JSON that StepStone's own public result pages are rendered from — 25 listings per request; a healthy page lands in about a second, and the median run on 2026-09-06 took 21 s.


What does StepStone.de Jobs Scraper do?

It extracts live German job postings from stepstone.de into clean rows you can export as JSON, CSV, Excel or pull through the Apify API.

  • 🔎 Search the way the site does — keyword (Was), city / postal code / Bundesland (Wo) + radius, and StepStone's full facet set. All applied server-side, so you only pay for rows that already match.
  • 🏢 Employer-first output — company name, numeric company id, and the employer's StepStone profile URL on every row. That's the lead.
  • 🗺️ Real German geography — all 16 Bundesländer by name, exact-city ids that ignore radius spill, and sub-state region ids (Ruhrgebiet, Rhein-Main…).
  • 🏭 Segment like an agency — sector (Branche), discipline (Berufsfeld), contract type (Vertragsart), career level (Karrierestufe), working hours, advert language, extracted skills.
  • 📮 Application routeEXTERNAL means the employer runs its own ATS (a direct mandate you can pitch); INTERNAL is StepStone quick-apply.
  • 🔔 Incremental monitor — schedule it and get only the listings that are new since the last run, with Slack / webhook alerts.
  • 🎯 Noise control that worksexcludeStaffingAgencies removes Arbeitnehmerüberlassung postings (staffing agencies reposting other people's mandates) server-side, before you are charged. Toggles for sponsored and anonymous rows exist too, but read Honest limits first — those flags are currently always false.

What data does it extract?

Every listing is one dataset row. Fill rates below are measured on 430 real rows across seven runs on 2026-09-06, not estimated (runs 88I8ff2ju9nnwzgkk · vhizaSgmf5DnI2j9e · JcpJGdYucxuFcbd6g · FabH8m3esnd22xaQM · 3QlZWTy2rpUx1Eq7t · zs29FBmZPZzoRgTUM · Unqkj9aV8fbhbSavU):

FieldFillWhat it is
jobId100%StepStone's numeric listing id — the dedupe key
title100%Job title
jobUrl100%Absolute listing URL, click-tracking suffix stripped
companyName100%The employer — the lead
companyId100%Numeric employer id (feed it back in to sweep all their roles)
companyUrl430/430 on keyword searches; 10/10 on a Company page URL run; 0/10 on a companyId runThe employer's StepStone profile / jobs page. StepStone omits it when the result list is already scoped to one employer by id — use companyUrl (the page route) if you need that column (runs 88I8ff2ju9nnwzgkk · UtizOckyVXZN5JDgE · 64wRh1OSbv8CF59vZ)
companyLogoUrl100%Employer logo
location100%City, or a multi-city string for nationwide roles
datePosted100%ISO-8601 with timezone
textSnippet100%~300-character description teaser, HTML entities decoded — StepStone ships this field escaped, so German umlauts arrive as ä/ö/ü, never as ä
harmonisedId100%StepStone's cross-brand UUID
workFromHome / workFromHomeCode100%Whether home office is offered (read Honest limits)
isSponsored, isTopJob, isHighlighted, isAnonymouspresent on 100%, but false on 430/430Placement + anonymity flags — see Honest limits
isPartnershipJob, isBackfilled, isCrossPostedpresent on 100%, false on 430/430Syndication flags
partnerSourceSite0/430Always null on this surface
labels, labelTypes174/430 overall (40.5%); 80/200 (40.0%) on a single 200-row vertrieb run (run UP3f6WyjxARTBWmJD)Badges such as QUICK_APPLY, NO_COVER_LETTER. Genuinely sparse and clustered by page — inside that one run the per-page fill ran from 4/25 to 16/25, and two 50-row vertrieb runs the same day read 7/50 and 8/50. So ~40% is the fill; any 50-row slice is noise, not a per-keyword rate
isQuickApplytrue/false on the ~40% of rows that carry labels, null on the restMeasured 2026-09-06 on 100 maschinenbau rows (run Unqkj9aV8fbhbSavU): 30 true, 16 false, 54 null; and on 200 vertrieb rows (run UP3f6WyjxARTBWmJD): 76 true, 4 false, 120 null. null means "this row carries no badges at all", not "no quick apply"
section100%main = a real match (padding rows are dropped by default)
searchQuery, sourceUrl, resultPage, scrapedAt100%Provenance for every row
positionOnPage, positionAbsolute100%Ad-ranking position (turn off with compactOutput)
isNewmonitor modeSet on listings unseen since the previous run
rawopt-inStepStone's untouched source object, via includeRawJson

Who is it for?

  • 🤝 Recruitment & staffing agencies — a live map of who in Germany is hiring what, where, and through whose ATS. Filter to applicationMethod: "EXTERNAL" and experienceLevels: ["manager"] for direct mandates with a decision-maker attached.
  • 🎯 B2B / GTM teams — hiring is the strongest public buying signal there is. A company posting five DevOps roles is buying tooling; one posting warehouse staff is expanding logistics.
  • 📊 Labour-market and HR analysts — corpus sizes measured 2026-09-06 run from 8,731 listings for maschinenbau to 47,962 for pflege, sliceable by sector, region and date.
  • 🏗️ Job boards & aggregators — a clean, deduplicated German feed with stable ids.
  • 💼 Sales teams doing account-based selling — pass a companyId and get every open role at one employer.

Quick start

{ "searchQueries": ["vertrieb"], "maxItems": 50 }

Fresh mandates in Bavaria, direct employers only

{
"searchQueries": ["vertriebsingenieur", "key account manager"],
"bundesland": ["Bayern"],
"postedWithin": "7d",
"applicationMethod": "EXTERNAL",
"excludeStaffingAgencies": true,
"sortBy": "date",
"maxItems": 500
}

Daily "new listings only" monitor with Slack alerts

{
"searchQueries": ["pflegefachkraft"],
"location": "Hamburg",
"radius": 30,
"postedWithin": "24h",
"sortBy": "date",
"incrementalMode": true,
"slackWebhookUrl": "https://hooks.slack.com/services/...",
"maxItems": 0
}

Attach an Apify Schedule and each run emits only what is new since the last one.

Every open role at one employer

{ "companyUrl": "https://www.stepstone.de/cmp/de/gi-group-deutschland-gmbh-213579/jobs", "maxItems": 0 }

Paste a search you built in your browser

{ "startUrls": ["https://www.stepstone.de/jobs/pflege/in-hamburg?radius=30&wfh=2"], "maxItems": 100 }

The URL's filters are parsed back out and the whole result set is re-paged. A pasted URL is re-paged exactly as pasted, so the filter fields on the input form (location, radius, Bundesland, contract type…) do not apply to it — set any of them alongside a start URL and the run names each one it is ignoring, in the log and in the status message (measured run tgfpr06C8zrQTQsUD, where location: "Berlin" and radius: 50 were both named as not applied).


How much does it cost?

Pay-per-event, one event, no start fee, no subscription.

EventWhat it coversPrice
job-listing-scrapedeach job listing returned and saved$0.0005

$0.50 per 1,000 listings. Duplicates, rows you excluded with the filters, StepStone's non-matching "similar jobs" padding, and rows your outputFields list would leave empty are all dropped before delivery, so you are never charged for them. Delivery and billing happen in the same call, so a charge cap can never leave you paying for rows you didn't receive.

Measured 2026-09-06 on the shipped build: outputFields: ["salary","jobTitle"] — two names this actor does not produce — returns 0 rows and charges $0.00 (run YYL1cHyQAWvbXzTh8), with the reason in the status message. It used to deliver and bill five empty {} rows.


Honest limits

Things this actor deliberately does not claim:

  • There are no salary figures on this surface. StepStone's result pages carry a unifiedSalary object on about half the rows, but every field inside it — min, max, currency, period — was null on 25 of 25 rows measured. Real salary numbers live on the individual listing page, which StepStone's robots.txt disallows. Some competing actors advertise StepStone salary data anyway; this one does not return a salary field at all rather than ship an empty column.
  • workFromHome tells you whether, not which. Measured across 450 rows: the row-level code is effectively binary (none / offered). Code offered covers 97–100% of both the "Nur Home-Office" and the "Teilweise Home-Office" result sets, so it cannot separate fully-remote from hybrid. To actually select fully-remote roles, use the server-side homeOffice: "only" filter — that one provably narrows the corpus (22,136 → 112 listings for vertrieb, runs OHj6iaW9j9X3bJmbC / kZ34fh3c6bKmDmYb4, 2026-09-06).
  • The sponsored / anonymous / syndication flags are always false here. Re-measured 2026-09-06 across 430 listings on 6 keywords, isSponsored, isTopJob, isHighlighted, isAnonymous, isTrafficFromPartner, hasFuturePosting, isPartnershipJob, isBackfilled and isCrossPosted came back false on 430 of 430 rows on StepStone's logged-out result pages. The fields are returned because they are part of the source record and may start populating, and the matching excludeSponsored / excludeAnonymous toggles are wired correctly — but as of today they have nothing to remove. The exclusion that does bite is excludeStaffingAgencies, which filters server-side on contract type.
  • No commute-time (Pendelzeit) filter. StepStone's own commute facet has no URL form — it is a client-side widget that calls an endpoint robots.txt disallows. Rather than ship a parameter that silently does nothing, it is not offered.
  • No detail-page enrichment. Full descriptions and requirements live under /listing/, which robots.txt disallows. This actor scrapes result pages only — one site, one function.
  • StepStone is a commercial job board and can tighten its defences. Transport is measured, not assumed (below). If first-attempt rates ever slide, switch proxyConfiguration to RESIDENTIAL with country DE.

Reliability — measured, not asserted

The actor runs plain HTTP through the Apify datacenter proxy, forced to HTTP/1.1.

That detail is the whole ballgame. Over HTTP/2 the datacenter pool draws NGHTTP2_INTERNAL_ERROR stream resets that poison an entire proxy session — one session failed 10 of 10 consecutive calls while another was clean 10 of 10. Fresh-session HTTP/2 measured 60.0% (12/20) and 71.4% (25/35) first-attempt. Forcing HTTP/1.1 and retiring the session on every transport error removes that failure class.

Measured on the Apify platform itself (not on a dev machine, which flatters the numbers), on the shipped transport, across 59 result pages in 34 runs on 2026-09-06:

Runs completed34 of 34 SUCCEEDED — 0 FAILED, 0 TIMED-OUT
First-attempt success64.4% (38/59 pages)
Pages lost after exhausting retries0 of 59
Rows delivered vs rows charged987 / 987 — exact match on every run
Residential rescues needed1 page of 59 (run 88I8ff2ju9nnwzgkk)
Run timemedian 21 s, worst single-search run 145 s (the one that needed the residential rescue), against a 300 s health-check threshold

Be aware that this number moves. The same test on 2026-09-04 measured 86.8% first-attempt over 114 pages; on 2026-09-06 the Apify datacenter pool was materially worse and the same code measured 64.4%. What did not move is the outcome: every run still succeeded and no page was lost, because the retry budget (default 8, floor 5 in code) and the residential rescue absorb it. That is the honest shape of this lane — the first-attempt rate is a weather report, the delivered/charged match and the zero lost pages are the guarantee.

Retries are what turn that first-attempt rate into complete results, which is why the retry floor is enforced in code and the default is 8. Retries cost time, not money: you are only ever charged for rows you receive.

And if the datacenter proxy burns every one of those attempts on a page — measured: it happens during bad patches, and it happened once in these 34 runs — the actor automatically retries that page through RESIDENTIAL / country DE rather than leaving a hole in your results. Tested by cutting the datacenter path entirely: 5 out of 5 dead pages were rescued, full 25 rows each. It only ever fires after the cheap path is exhausted, so a healthy run never pays for it, and it steps aside if you pin your own proxy groups.

If a page is still unreachable after all of that, the run says so explicitly in its status message and log — you are told about the coverage gap rather than quietly handed a short dataset, and you are not charged for the missing rows.

Pagination is ?page=N, 25 rows per page. (?of=<offset> is silently ignored by StepStone — it echoes offset 0 and re-serves page 1.) Crawled maschinenbau contiguously through page 14 plus deep probes at 20/30/40/60/100/…/347: every page served 25 rows, page 347 served 17, and pages beyond it returned nothing with pageCount pinned at 347.

If first-attempt rates ever fall meaningfully below this, switch proxyConfiguration to RESIDENTIAL with country DE — measured working, just slower.


The actor reads publicly available job advertisements — no login, no paywall, and no job-seeker data. Employer names and company pages are business information the postings publish deliberately.

StepStone's robots.txt, fetched 2026-09-04, reads in relevant part:

User-agent: *
Disallow: /*?*
Disallow: /jobs/*?*
Allow: /jobs/*?q=*
Disallow: /jobs/*?q*&*
Disallow: /jobs/vollzeit/
Disallow: /jobs/teilzeit/
Disallow: /public-api/
Disallow: /listing
Disallow: /listing/*

This actor deliberately stays off /public-api/ and /listing/ — which is where most competing scrapers get their salary and description data — and routes the working-hours filter through the query parameter wt= rather than the disallowed /jobs/vollzeit/ path.

Scraping may still conflict with StepStone's Terms of Service, and you are responsible for how you use the data, including your obligations under the GDPR for any personal data a listing happens to contain (a named contact person, for instance). Review StepStone's current terms before commercial use.


FAQ

What is StepStone.de? StepStone is Germany's largest job board, carrying hundreds of thousands of live vacancies across every sector. Broad keywords are deep — counts read off StepStone's own result pages on 2026-09-06: pflege 47,962 listings, einkauf 23,284, vertrieb 22,138, logistik 14,583, it 12,361, maschinenbau 8,731. (These move by a few rows an hour; the run ids are in SIGNOFF.md.)

Do I need an account, login, or API key? No. The actor reads public result pages.

Does it return salaries? No — and that is deliberate. See Honest limits: the field exists on StepStone's result pages but is empty on every row.

How do I get only fully-remote jobs? Set homeOffice: "only". Do not filter on the output workFromHome column for this — it cannot separate remote from hybrid.

How do I know the filters actually applied? The actor reads StepStone's own categorization object back off the page, which echoes every parameter the server bound, and compares it to what you asked for. If they disagree it fails the run instead of delivering and billing rows that don't match your search. The check runs on every search in the run and on every route, including a pasted start URL — verified 2026-09-06: a 3-task run (two keywords + a company page) logged three separate binding verifications (run IrjPnRC8EgptlVO58), and a start URL carrying wfh=2 logged its own (run fLEmCBMsJgTIVAeGW). Two honest caveats: filters you set on the form do not apply to a pasted URL at all (the run says which ones it ignored), and StepStone echoes an id-type filter (regionIds, cityIds, sectorIds, skillIds) straight back even when the id matches nothing — so on a zero-row run the message tells you the echo does not vouch for those values and names them (run pPcdlgljvzmCB9Px7).

Will I be charged twice for the same job across overlapping searches? No. Deduplication is on by default and collapses on StepStone's numeric listing id across every page and every keyword in the run, before delivery. Measured 2026-09-06: the same search fed four times read 250 raw rows and delivered 100, all unique — 150 repeats collapsed before billing (run JcpJGdYucxuFcbd6g). The control with deduplicate:false delivered 50 rows containing only 25 unique ids (run FabH8m3esnd22xaQM).

What are "similar jobs"? When a search is narrow StepStone pads the page with recommended/regional listings that do not match your filters — a one-result query came back as 1 real match plus 24 padded rows. Those are dropped (and never billed) unless you set includeSimilarJobs: true.

Can I scrape every open role at one employer? Yes — pass companyId or companyUrl. On both routes the numeric employer id is checked against every row on page 1 and the run fails rather than deliver another employer's leads (companyUrl's id is read out of the URL slug). Verified 2026-09-06: companyUrl → 10/10 rows owned by employer 213579 (run UtizOckyVXZN5JDgE); companyId → 5/5 (run W0QQmkjLapgK4j2cy).

Can I combine an employer with a keyword? Yes, with companyId. Give it a companyId and searchQueries and the two are intersected into one server-side search per keyword — measured 2026-09-06: companyID 213579 + "lager" = 83 matching jobs, 10/10 delivered rows owned by that employer (run 64wRh1OSbv8CF59vZ). A company page URL cannot be narrowed this way (StepStone's /cmp/ page takes no keyword); the run says so in the log and sweeps the page in full.

Can I resume or shard a large crawl? Yes — startPage shifts the window (page N covers results (N−1)×25 … N×25) and maxPagesPerQuery caps the depth. If you start past the last page, the run tells you so with the real number instead of blaming StepStone: {"searchQueries":["vertrieb"],"startPage":2000} returns 0 rows, charges nothing, and says "this search only has 889 page(s) … page 2000 does not exist … start at a page between 1 and 889" (run pg9BuwQhY9AH6B4bC, 2026-09-06). The page count is read off StepStone's own pagination before the empty answer is explained — and if the out-of-range page carries no pagination at all, page 1 is fetched to get it (that probe delivers nothing and is never billed).

In incremental mode, does a maxItems cap mean I get fewer rows each run? No — and it's worth understanding why. Incremental mode suppresses listings it has already seen, then keeps paging to try to fill your cap with unseen ones. So on a huge keyword, run #2 can still return a full 50 rows (just deeper, older ones). To get a true "only what's new today" feed, bound the corpus rather than the row count: pair incrementalMode with postedWithin: "24h" and sortBy: "date", and set maxItems: 0.

How do I keep costs down on a daily monitor? sortBy: "date" + stopAfterDays: the crawl stops paging as soon as a page is entirely older than your cutoff. Add incrementalMode so you only get (and only pay for) listings you haven't seen.

What happens if a search matches nothing? The run exits cleanly with zero rows, charges nothing, and the status message names which of the causes it actually was — it never falls back to blaming StepStone. Measured 2026-09-06, one run per case:

CauseWhat the run says
You started past the last page"this search only has 889 page(s)… page 2000 does not exist… StepStone is NOT empty" (run pg9BuwQhY9AH6B4bC)
Your outputFields names no real column"names no field this run produces… dropped instead of delivered and nothing was charged — StepStone had results, this is a column-name problem" (run YYL1cHyQAWvbXzTh8)
A disciplines value isn't one StepStone offers"None of the Discipline (Berufsfeld) value(s) you set is one StepStone offers… nothing was scraped and nothing was charged" (run 1SrziGfztnbPB5OtW)
Start URLs that aren't stepstone.dethe URLs are named and nothing else is scraped in their place
Your Max total charge cap was hit firstsays so, and tells you to raise the cap
StepStone blocked every attemptsays it is a transient block, not an empty search
StepStone genuinely has nothing"it served a valid result page with an empty list", plus whether the filters were verified as bound (run tgfpr06C8zrQTQsUD)
Incremental mode had already delivered everything on the page"All 25 listing(s) on this search were already delivered by a previous run of this monitor… nothing is new since last time. That is the normal quiet day of a scheduled monitor, not a misconfiguration", $0.00 (run KicEWtqZDAUo7JvMd, 2026-09-07)
Everything on the page was on your skipJobIds list"25 listing(s) were on the 'Skip job IDs' list you supplied", $0.00 (run ZHXTb8VOEV1bA1jTu, 2026-09-07)
A filter you set removed the restnames the count and the toggle(s) you actually set — it never suggests one you left off

If a filter you set never bound, that is named too, rather than reported as an empty StepStone. And if the page carried no result-list block at all, the zero is reported as unconfirmed instead of confirmed (run ILdkl9AiL7mSFXhPm).

The actor's own memory is never reported as your filters. Incremental mode and skipJobIds are counted separately from the filters you set, in both the status message and the run log (0 excluded by the filters you set | 25 already delivered by an earlier run (incremental mode) — run KicEWtqZDAUo7JvMd). A monitor's quiet day says so in those words; it does not tell you to widen a search that is working correctly.

Can I choose my own columns? Yes — outputFields takes a list and returns exactly those, in that order. A name this actor does not produce is ignored and named in a log warning alongside the full field list, so a typo cannot silently cost you a column (run gdy0Q8yfZGmWgXZgU; a mixed list still works — ["companyName","title","salary"] returned 25 two-column rows on run fLEmCBMsJgTIVAeGW). If none of your names is real, the rows would be empty — those rows are dropped rather than delivered, and you are not charged for them: ["salary","jobTitle"] returns 0 rows and $0.00 with the reason in the status message (run YYL1cHyQAWvbXzTh8, 2026-09-06). compactOutput drops the ad-tracking positions for a clean CSV.

Can I integrate with Make, Zapier, n8n or my CRM? Yes — webhookUrl and slackWebhookUrl push new listings in monitor mode, or pull the dataset through the Apify API.

Which proxy should I use? Leave it on the default. Apify datacenter is the cheap path and is measured sufficient — 64.4% of pages first try on 2026-09-06, 0 pages lost, 34/34 runs SUCCEEDED — and when a page does fail every datacenter attempt the actor escalates that page to RESIDENTIAL / country DE by itself. You only need to pin RESIDENTIAL manually if you want every request on it from the start (fewer retries, so a faster run, at a higher proxy cost).


Feedback

Missing a field or a filter? Open an issue on the actor — fixes and feature requests welcome.