Arbeitsagentur Jobs Scraper - German Federal Jobbörse
Pricing
from $0.80 / 1,000 per job returneds
Arbeitsagentur Jobs Scraper - German Federal Jobbörse
Scrape the German Federal Employment Agency Jobbörse: ~929,000 live vacancies with employer, PLZ/city, pay, contract type, start date and a per-job Zeitarbeit (Arbeitnehmerüberlassung) flag no rival exposes. Auto-splits past the site's 10,000-row query ceiling. No login, no API key.
Pricing
from $0.80 / 1,000 per job returneds
Rating
0.0
(0)
Developer
Scrapers Delight
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
🇩🇪 Arbeitsagentur Jobs Scraper — the German Federal Jobbörse, as structured rows
Turn arbeitsagentur.de/jobsuche — the German Federal Employment Agency's Jobbörse, the single largest vacancy corpus in the German-speaking market — into clean JSON/CSV/Excel rows: employer, PLZ and city, occupation, contract type, working-time pattern, start date, pay where stated, and the Zeitarbeit (Arbeitnehmerüberlassung) classification that tells you whether a posting comes from a direct employer or from a temp-staffing agency.
No login. No API key. No CAPTCHA. Datacenter proxies are enough.
📊 What you are scraping (measured live 2026-09-04)
Offer type (angebotsart) | Live vacancies |
|---|---|
Arbeit — regular jobs (1) | 746,533 |
Ausbildung & duales Studium (4) | 170,391 |
Praktikum & Trainee (34) | 14,516 |
Selbständigkeit (2) | 3,226 |
| Total | ~934,700 |
Churn is roughly 9% a day — on a sample occupation, 276 of 3,109 postings were less than 24 hours old. Set Published within → Today only and this actor becomes a daily new-vacancy feed.
🚀 What does Arbeitsagentur Jobs Scraper do?
- 🔎 Searches the way the site does. Keyword (
was), location + radius (wo+umkreis), offer type, publication window, working time, contract type, Berufsfeld, Beruf, employer name, private-placement opt-in, disability-friendly, sort order — every filter is wired to real behaviour and every one was verified to genuinely narrow the result set, not just relabel it. - 🏢 Classifies Zeitarbeit per job.
isTemporaryStaffing/istArbeitnehmerUeberlassungis the qualifier a staffing firm, an HR-SaaS SDR or a Personalberatung actually buys on: is this posting a competitor agency, or a direct employer you can sell to? No rival in this lane exposes it. - 🧩 Breaks the site's 10,000-row ceiling. arbeitsagentur.de refuses to page past 10,000 rows for
any single query (page 400 is the last one that returns data — an Elasticsearch
max_result_window). Auto-split cuts a bigger query by federal state → publication window → Berufsfeld → Beruf → contract type → working time, then merges and de-duplicates. That is how "all 746,533 German vacancies" is actually deliverable. - 🧾 Never double-bills. Rows are de-duplicated on
referenznummeracross every search term, location and auto-split slice before anything is charged. - 🌍 Emits English aliases.
jobTitle,employer,city,postalCode,publishedAt,isTemporaryStaffing… alongside the original German keys, so rows drop straight into a CRM. - 🔗 Paste any Jobbörse URL. Build the search in the site's own UI, paste the URL, and every
query parameter is honoured verbatim — including filters this schema does not expose. A
/jobsuche/jobdetail/<referenznummer>link is fetched as a single job. - 📝 Optional detail enrichment adds the full job-ad text (median 2,386 characters), the employer's street address, the Zeitarbeit and private-placement flags, the required education level, and the partner job board that syndicated the posting.
📦 Output — measured fill rates
List fields, measured on a 350-row national sample across 7 varied searches (Pflegefachkraft · Softwareentwickler · LKW-Fahrer · Elektroniker · Buchhalter · Erzieher · whole corpus), 350 of 350 unique:
| Field (German → English alias) | Fill |
|---|---|
referenznummer → jobId (stable id, dedupe key) | 100% |
stellenangebotsTitel → jobTitle | 100% |
firma → employer | 100% |
ort → city · land · region · breite/laenge → latitude/longitude | 100% |
plz → postalCode | 97.4% |
hauptberuf → occupation · alleBerufe | 100% |
stellenangebotsart → offerType | 100% |
veroeffentlichtAm · datumErsteVeroeffentlichung · aenderungsdatum | 100% |
eintrittsdatum → startDate | 100% |
vertragsdauer → contractDuration | 100% |
verguetungsangabe → salaryBasis | 100% |
arbeitszeitVollzeit → isFullTime | 95.7% |
quereinstiegGeeignet → isCareerChangerFriendly | 84.9% |
istGeringfuegigeBeschaeftigung → isMiniJob | 84.6% |
homeofficemoeglich → isHomeOfficePossible | 58.9% |
arbeitgeberKundennummerHash → employerKey (groups all postings by one company) | 51.7% |
externeURL → externalUrl (partner-board deep link) | 48.3% |
strasse → street | 41.4% |
artDerVerguetung → salaryKind | 36.9% |
gehaltsspanneVon/Bis → salaryFrom/salaryTo | 21.1% |
festgehalt → salaryFixed | 15.7% |
chiffrenummer (anonymised listing) | 14.6% |
Detail fields (fetchJobDetails: true), measured on 60 enriched rows:
| Field | Fill |
|---|---|
stellenangebotsBeschreibung → description (394–4,079 chars, median 2,386) | 100% |
allianzpartnerName / allianzpartnerUrl → partnerJobBoard | 100% |
istBetreut | 100% |
istArbeitnehmerUeberlassung → isTemporaryStaffing | 85.0% |
istPrivateArbeitsvermittlung → isPrivatePlacement | 76.7% |
istBehinderungGefordert | 75.0% |
strasse → street | 66.7% |
hausnummer → houseNumber | 48.3% |
geforderterBildungsabschluss → requiredEducation | 0–50%, occupation-dependent (see limits) |
ausbildungsverguetungJahr1/2/3 → apprenticePayYear1/2/3 appear only on Ausbildung postings
(measured 4 of 40 on angebotsart=4).
Turn on Include the raw source objects to keep the untouched ergebnisliste / jobdetail
payload under raw for any field this actor does not map.
💵 Pricing
Pay per event. No monthly rental, no platform-usage surcharge, no charge for starting a run.
| Event | Price | When it fires |
|---|---|---|
Per job returned (job-scraped) | $0.0008 / job — $0.80 per 1,000 | Each unique job delivered to your dataset |
Per detail fetch (job-detail-enriched) | $0.0012 / job — $1.20 per 1,000 | Only when Fetch job details is on, and only for enriched rows that actually shipped |
Rows ship through Apify's budget-aware push, so delivered always equals billed: if you set a max-total-charge cap, the run stops at the cap instead of handing you rows you were never charged for. Duplicates across overlapping searches are removed before billing.
A full sweep of the 746,533-vacancy Arbeit corpus costs ≈ $597, refreshed as often as you like. A daily "everything posted in the last 24 hours" feed is ~9% of that per day.
⚙️ Example inputs
Daily new-vacancy feed for a staffing firm — direct employers only, Berlin + 50 km:
{"searchTerms": ["Pflegefachkraft", "Altenpfleger"],"locations": ["Berlin"],"radiusKm": 50,"publishedWithinDays": "1","temporaryStaffing": "exclude","fetchJobDetails": true,"maxItems": 2000}
Map the competition — every Zeitarbeit posting in Bayern:
{"locations": ["Bayern (Bundesland)"],"radiusKm": 0,"temporaryStaffing": "only","fetchJobDetails": true,"maxItems": 1000}
Whole-country apprenticeship sweep past the 10,000-row ceiling:
{"offerTypes": ["4"],"autoSlice": true,"maxItems": 0}
Paste a search you built on the site:
{"startUrls": [{ "url": "https://www.arbeitsagentur.de/jobsuche/suche?angebotsart=1&was=Elektroniker&wo=K%C3%B6ln&umkreis=25&veroeffentlichtseit=7" }],"maxItems": 500}
🧠 How it works
The Jobbörse is a server-side-rendered Angular app: every search page and every job-detail page
embeds its complete state as JSON in a <script id="ng-state"> tag. This actor reads that blob —
no DOM scraping, no brittle CSS selectors, and no API key.
- Search:
GET /jobsuche/suche?<filters>→#ng-state→suchergebnis.ergebnisliste[](25 rows per page), withsuchergebnis.maxErgebnisseas the structurally-read hit count. - Detail:
GET /jobsuche/jobdetail/<referenznummer>→#ng-state→jobdetail. - The documented
rest.arbeitsagentur.de/jobboersev4 API returns 403 without an app key; this actor never touches it, because the SSR blob is the identical payload with no key at all.
Transport: plain HTTP over Apify datacenter proxies, HTTP/2, de-DE accept-language,
responses decoded explicitly as UTF-8 (the default decoding mangles umlauts — "Universität" becomes
"Universit?t"). 512 MB, no browser. Verified over 200+ calls with no CAPTCHA, no Cloudflare, no
DataDome, no rate limiting. RESIDENTIAL + country DE also works and is exposed as a fallback.
⚠️ Honest limits — read these before you buy
- The 10,000-row ceiling is the site's, not ours. Any single query stops at 10,000 rows: page 400 returns 25 rows, page 401 returns an empty page state. Auto-split works around it by cutting a big query into smaller ones, but if you pin every axis yourself (a keyword AND a city AND a date window AND a contract type…) and the remainder is still over 10,000, the actor logs a warning and returns the first 10,000 for that query. Narrow it further to reach the rest.
- About 5% of responses come back HTTP 200 with an empty page state. These are retried, never reported as "no results" — a naive parser reads them as zero rows and finishes SUCCEEDED with nothing. Measured: 19/20 clean on the first attempt, 18/18 clean with one retry. Keep Max retries per request at 3 or higher.
- Page size is fixed at 25. The site ignores
size=. Budget one request per 25 rows. geforderterBildungsabschlussis patchy — 10 of 20 across mixed offer types, but 0 of 60 on a nursing/retail/mechatronics sample. Do not build a workflow that depends on it.- Salary is mostly absent.
gehaltsspanneVonis present on ~21% of rows andfestgehalton ~16%. German postings usually state a collective-agreement reference (verguetungsangabe, 100%) instead of a number. - Street address is a detail-page field and present on ~2 of 3 enriched rows — 41% at list level. If you need it on every row, turn on Fetch job details.
employerNameis exact-match, not fuzzy.Vivantes Netzwerk für Gesundheit GmbHreturns ~320 rows;Charitéreturns 0.- Six federal-state names collide with same-named towns.
wo=Hessenreturns a town in Sachsen-Anhalt. Use the disambiguating form —Hessen (Bundesland)— when you mean the state. The 16 states in that form sum to 98.3% of the national Arbeit corpus. publishedWithinDaysonly honours 0, 1, 7, 14 and 28. Any other number is silently ignored by the site and returns the unfiltered set; this actor warns and drops the filter rather than pretending it applied.- Zeitarbeit "only" costs roughly twice the requests. The site has no positive Zeitarbeit filter — only a negative one — so "only" crawls both sets and takes the difference. Turn on Fetch job details with it to have every returned row verified against the official flag.
- Some postings are outside Germany. The corpus includes Austrian and other cross-border
vacancies. The
landfield tells you which; filter on it if you only want DE.
❓ FAQ
Is scraping arbeitsagentur.de allowed?
The site's robots.txt is fully open, verbatim:
User-agent: * / Disallow: / Allow: / — with no crawl-delay. This actor reads only public,
unauthenticated pages that the Bundesagentur für Arbeit publishes for exactly this purpose. You are
responsible for complying with the site's terms of use and with GDPR for any personal data the
postings contain.
Do I need an API key or a login? No. Neither. The documented v4 API needs an app key; this actor does not use it.
How many jobs can I get? All of them. ~934,700 across the four offer types. Auto-split handles the site's per-query ceiling.
How fresh is the data? Live — every run reads the site at that moment. Roughly 9% of the corpus turns over daily.
Can I get only new postings since yesterday? Yes. Set Published within to Today only or Last 1 day.
Can I tell a temp agency from a direct employer?
Yes — that is this actor's wedge. Use Zeitarbeit → Exclude for direct employers only, or
Only for the agencies themselves. Add Fetch job details to have every row carry the official
istArbeitnehmerUeberlassung flag. Verified end-to-end: an "exclude" run returned 0/20 agency
postings, an "only" run returned 20/20.
Can I group every vacancy from one company?
Yes — arbeitgeberKundennummerHash / employerKey is a stable per-employer key (present on ~52%
of rows). employerName also filters by exact registered name.
Does it get the full job-ad text? Yes, with Fetch job details on: 100% fill, median 2,386 characters.
Are the German umlauts correct? Yes. Responses are decoded explicitly as UTF-8 — "Fachverkäuferin", "Backstube Wünsche", "München".
How do I search a specific city, PLZ or federal state?
locations accepts a city, a postal code, a federal state or Deutschland. Use the
<Name> (Bundesland) form for the six state names that collide with towns.
What happens if my run hits its charge cap? It stops cleanly and tells you so. Every row you received was billed and every row billed was delivered — the cap never leaks unpaid rows in either direction.
What happens if there are genuinely no matches? The run exits cleanly with 0 rows and a status message explaining why (no matches vs. blocked), and charges nothing. It does not fail.
Can I run this on a schedule? Yes — Apify Schedules. A "last 1 day, my region, direct employers only" run is the standard setup.
⚖️ Legal & fair use
This actor collects public, unauthenticated job postings published by the Bundesagentur für
Arbeit, a German federal agency, on a site whose robots.txt explicitly permits crawling of all
paths. It sends no credentials, solves no challenges, and forges no signatures.
Job postings can contain personal data (named contacts). How you store, process and use that data is your responsibility, including under the GDPR/DSGVO and the site's own terms of use. Concurrency is deliberately conservative by default (5) because this is a public service funded by German taxpayers — please leave it there unless you have a reason not to.
Not affiliated with, endorsed by, or connected to the Bundesagentur für Arbeit.