# Germany Job Finder & Data Normalizer (BA API + ATS Feeds) (`wonderful_beluga/germany-job-finder`) Actor

Multi-source job search actor for Germany. Extracts structured listings from the Bundesagentur für Arbeit API, Personio, Greenhouse, Lever, and SmartRecruiters. Normalizes salaries, seniority, skills, and locations.

- **URL**: https://apify.com/wonderful\_beluga/germany-job-finder.md
- **Developed by:** [Zaher el siddik](https://apify.com/wonderful_beluga) (community)
- **Categories:** Automation, Jobs, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Germany Job Finder & Data Normalizer

![Germany Job Finder](apify.png)

[![Apify Actor](https://img.shields.io/badge/Apify-Actor-orange.svg)](https://apify.com)
[![TypeScript](https://img.shields.io/badge/TypeScript-ES2022-blue.svg)](https://www.typescriptlang.org/)
[![License](https://img.shields.io/badge/License-Apache_2.0-green.svg)](LICENSE)

Search the entire German job market from a single API. This Actor aggregates live job listings from the official Bundesagentur für Arbeit REST API and the public feeds of four major applicant tracking systems (Personio, Greenhouse, Lever, and SmartRecruiters), then normalizes every listing into one consistent JSON schema with enriched fields for salary, seniority, skills, contract type, and location.

Because it consumes structured API and XML streams rather than scraping fragile HTML pages, it runs without headless browsers, CAPTCHAs, or anti-bot blocking, and it does not break when career sites change their layout.

### Contents

- [What it does](#what-it-does)
- [Data sources](#data-sources)
- [Key features](#key-features)
- [Use cases](#use-cases)
- [Input parameters](#input-parameters)
- [Output schema](#output-schema)
- [Performance and cost](#performance-and-cost)
- [Running locally](#running-locally)
- [Integrations and API access](#integrations-and-api-access)
- [FAQ](#faq)
- [Compliance](#compliance)

### What it does

1. Queries all selected sources concurrently with your keywords, location, and filters.
2. Normalizes each raw listing into a single schema: parsed salary, resolved Bundesland and postal code, classified contract type, seniority level, workplace arrangement, and extracted tech skills.
3. Filters out staffing agencies (Zeitarbeit), enforces recency and contract filters, and removes duplicates across sources.
4. Pushes clean, deduplicated job objects to the dataset, ready for export as JSON, CSV, or Excel.

### Data sources

| Source | Type | Coverage |
| --- | --- | --- |
| `arbeitsagentur` | Official REST API of the German Federal Employment Agency (Bundesagentur für Arbeit) | Nationwide, all industries; the largest job index in Germany |
| `personio` | Personio ATS XML feeds | German SMEs, startups, and scaleups |
| `greenhouse` | Greenhouse Job Board API | German tech scaleups and unicorns (HelloFresh, N26, SumUp, Celonis, Flix, and others) |
| `lever` | Lever Postings API | Companies with German offices; results are restricted to postings located in Germany |
| `smartrecruiters` | SmartRecruiters Postings API | Large German enterprises (Bosch, Delivery Hero, Sixt, Vattenfall); queried with `country=de` |

All company boards are verified as live. Boards that go offline are skipped gracefully without failing the run. All sources run concurrently for fast execution.

### Key features

- **Official API ingestion.** No headless browsers, no CAPTCHAs, no anti-bot rate limiting. Direct JSON and XML streams keep runs fast, cheap, and reliable.
- **Salary normalization.** German salary expressions such as `50.000 EUR - 75.000 EUR pro Jahr`, `4.500 EUR / Monat`, or `60k - 80k EUR` are parsed into structured `{ min, max, currency, interval }` objects. Extraction is currency-anchored and context-validated, so arbitrary numbers or benefit amounts in long descriptions are never misread as salaries.
- **Seniority classification.** Each listing is classified as `INTERN`, `JUNIOR`, `MID`, `SENIOR`, `LEAD`, or `UNKNOWN` based on German and English title conventions (Werkstudent, Praktikum, Berufseinsteiger, Senior, Head of, Leiter, and more).
- **Staffing agency filter.** Listings from recruiters, headhunters, and Zeitarbeit agencies (Randstad, Adecco, Hays, Ferchau, Akkodis, GULP, and dozens more) are identified and excluded on request. For the Arbeitsagentur source this filter is additionally applied server-side.
- **Skill and tech stack extraction.** Over 60 technologies, frameworks, certifications, and enterprise tools (TypeScript, AWS, Kubernetes, SAP, DATEV, ISO 27001, GDPR/DSGVO, and others) are extracted into a clean array.
- **Location resolution.** German cities are mapped to their federal states (Bundesländer), 5-digit postal codes are extracted, and workplace arrangements are classified as `REMOTE`, `HYBRID`, or `ON_SITE`.
- **Umlaut-insensitive matching.** Keyword and location filters treat `München`/`Muenchen`/`Munich` and `Köln`/`Cologne` as equivalent, so no listings are missed due to spelling variants.
- **Cross-source deduplication.** Identical positions appearing on multiple boards are removed using both source IDs and normalized company/title/city fingerprints.

### Use cases

- **Job seekers**: monitor the German market for roles matching your skills, salary expectations, and preferred work arrangement, without checking a dozen portals.
- **Recruiters and HR analysts**: benchmark salaries, track hiring activity by company, region, or technology, and build talent market reports.
- **Job boards and aggregators**: feed a normalized, deduplicated stream of German listings directly into your product.
- **Market researchers**: analyze demand for technologies and skills across Bundesländer, industries, and seniority levels.
- **Lead generation**: identify companies that are actively hiring in a specific region or technology niche.

### Input parameters

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `keywords` | Array | `["Softwareentwickler", "Cloud Engineer"]` | Job titles or search terms in German or English. |
| `location` | String | `"Berlin"` | Target city or region. Leave empty to search all of Germany. |
| `radiusKm` | Integer | `50` | Search radius in kilometers around the target city (Arbeitsagentur only). |
| `sources` | Array | all five | Data sources to query: `arbeitsagentur`, `personio`, `greenhouse`, `lever`, `smartrecruiters`. |
| `workplaceType` | String | `"ALL"` | Filter: `ALL`, `REMOTE`, `HYBRID`, `ON_SITE`. Remote is also applied server-side on the Arbeitsagentur API. |
| `contractType` | String | `"ALL"` | Filter: `ALL`, `FULL_TIME`, `PART_TIME`, `MINIJOB`, `CONTRACT`, `DUAL_STUDY`. |
| `excludeAgencies` | Boolean | `true` | Exclude Zeitarbeit and recruiting agencies. |
| `extractSkills` | Boolean | `true` | Extract technologies and tools into the `skills` array. |
| `postedWithinDays` | Integer | `14` | Only return jobs published within the last N days (`0` = any time). |
| `maxItems` | Integer | `100` | Maximum number of jobs to return across all sources. |

Example input:

```json
{
  "keywords": ["Softwareentwickler", "Cloud Engineer"],
  "location": "Berlin",
  "radiusKm": 50,
  "sources": ["arbeitsagentur", "personio", "greenhouse", "lever", "smartrecruiters"],
  "workplaceType": "ALL",
  "contractType": "ALL",
  "excludeAgencies": true,
  "extractSkills": true,
  "postedWithinDays": 14,
  "maxItems": 200
}
```

Tip: ATS boards (Personio, Greenhouse, Lever, SmartRecruiters) often carry postings older than two weeks. Set `postedWithinDays` to `0` and leave `location` empty to see their full volume.

### Output schema

Each dataset item follows this structure:

```json
{
  "id": "ba-10000-1189324501-S",
  "title": "Senior Cloud Security Engineer (m/w/d)",
  "company": {
    "name": "FinTech AG",
    "domain": "fintech-example.de",
    "isAgency": false
  },
  "location": {
    "city": "10115 Berlin",
    "state": "Berlin",
    "country": "DE",
    "postalCode": "10115",
    "workplaceType": "HYBRID"
  },
  "compensation": {
    "min": 75000,
    "max": 90000,
    "currency": "EUR",
    "interval": "yearly",
    "rawText": "75.000 EUR - 90.000 EUR pro Jahr"
  },
  "contractType": "FULL_TIME",
  "experienceLevel": "SENIOR",
  "skills": ["AWS", "Terraform", "Kubernetes", "SOC 2", "TypeScript"],
  "applyUrl": "https://www.arbeitsagentur.de/jobsuche/jobdetail/10000-1189324501-S",
  "sourceUrl": "https://www.arbeitsagentur.de/jobsuche/jobdetail/10000-1189324501-S",
  "source": "arbeitsagentur",
  "postedAt": "2026-08-20T14:30:00Z",
  "refNumber": "10000-1189324501-S",
  "descriptionSnippet": "Beruf: Ingenieur/in - IT-Sicherheit"
}
```

Field notes:

- `compensation` is `null` when the source provides no salary information. Most German listings do not publish salaries; expect salary data on a minority of records.
- `experienceLevel` is one of `INTERN`, `JUNIOR`, `MID`, `SENIOR`, `LEAD`, or `UNKNOWN`.
- `location.workplaceType` is `UNKNOWN` when the listing does not state a work arrangement; such listings are kept when a workplace filter is active to avoid false negatives.

### Performance and cost

The Actor uses plain HTTP requests against JSON and XML endpoints rather than headless browsers, so compute consumption is minimal:

- Approximately 0.02 compute units per 1,000 ingested jobs
- Typical execution time of 10 to 30 seconds for several hundred jobs, since all sources are queried concurrently

### Running locally

```bash
npm install
npm run build
npm start
```

Provide input in `storage/key_value_stores/default/INPUT.json`. Run the test suite with:

```bash
npm test
```

### Integrations and API access

Like any Apify Actor, results can be:

- Exported from the dataset as JSON, CSV, Excel, or XML
- Fetched programmatically via the [Apify API](https://docs.apify.com/api/v2) or the JavaScript and Python clients
- Scheduled to run periodically and connected to Slack, email, Google Sheets, Zapier, or Make for automated job alerts
- Used as a data source in a larger Actor workflow

### FAQ

**Why do some sources return zero results for my query?**
Each source is filtered by your keywords, location, contract, and recency settings. Narrow keywords with a city filter and a short `postedWithinDays` window can legitimately produce zero matches on the smaller ATS boards while the Arbeitsagentur source still returns hundreds.

**Does it scrape LinkedIn, Indeed, or StepStone?**
No. Those platforms prohibit scraping and employ aggressive anti-bot measures. This Actor intentionally relies on the official federal API and public ATS feeds, which keeps it stable, legal, and fast.

**How current are the listings?**
Every run queries the sources live. Use `postedWithinDays` to restrict results to fresh postings.

**Can more companies be added?**
Yes. The ATS company lists are curated constants in the source code and can be extended with any company slug that exposes a public Personio, Greenhouse, Lever, or SmartRecruiters board.

### Compliance

This Actor queries public, documented endpoints only: the Bundesagentur für Arbeit open data API and the public postings APIs and feeds that applicant tracking systems expose for job distribution. It does not bypass authentication, does not collect personal data, and respects rate limits with retry backoff and per-feed request delays.

### License

Apache-2.0

# Actor input Schema

## `keywords` (type: `array`):

Keywords or job titles to search for (e.g., 'Softwareentwickler', 'Cybersecurity', 'Cloud Engineer', 'Account Executive').

## `location` (type: `string`):

Target German city or region (e.g. 'Berlin', 'München', 'Frankfurt am Main', 'Düsseldorf', 'Hamburg'). Leave empty for all Germany.

## `radiusKm` (type: `integer`):

Search radius around the specified city in kilometers (e.g., 25, 50, 100). Default is 50km.

## `sources` (type: `array`):

Select data sources to ingest jobs from.

## `workplaceType` (type: `string`):

Filter by work arrangement: Remote (Homeoffice), Hybrid, or On-Site.

## `contractType` (type: `string`):

Filter by German employment contract type.

## `excludeAgencies` (type: `boolean`):

If enabled, automatically filters out listings from recruiters, headhunters, and Zeitarbeit agencies (e.g., Randstad, Adecco, Hays, Ferchau).

## `extractSkills` (type: `boolean`):

Extract tech skills, programming languages, and tools into structured arrays (`skills`).

## `postedWithinDays` (type: `integer`):

Only retrieve jobs published within the last N days (0 = any time).

## `maxItems` (type: `integer`):

Maximum number of total jobs to return across all queries.

## `proxyConfiguration` (type: `object`):

Optional proxy settings.

## Actor input object example

```json
{
  "keywords": [
    "Softwareentwickler",
    "Software Engineer",
    "Cloud Engineer"
  ],
  "location": "Berlin",
  "radiusKm": 50,
  "sources": [
    "arbeitsagentur",
    "personio",
    "greenhouse",
    "lever",
    "smartrecruiters"
  ],
  "workplaceType": "ALL",
  "contractType": "ALL",
  "excludeAgencies": true,
  "extractSkills": true,
  "postedWithinDays": 14,
  "maxItems": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `jobListings` (type: `string`):

Dataset containing all normalized job search results.

## `runSummary` (type: `string`):

JSON report containing total jobs scraped, source breakdown, and execution parameters.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "Softwareentwickler",
        "Software Engineer",
        "Cloud Engineer"
    ],
    "location": "Berlin"
};

// Run the Actor and wait for it to finish
const run = await client.actor("wonderful_beluga/germany-job-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "Softwareentwickler",
        "Software Engineer",
        "Cloud Engineer",
    ],
    "location": "Berlin",
}

# Run the Actor and wait for it to finish
run = client.actor("wonderful_beluga/germany-job-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "Softwareentwickler",
    "Software Engineer",
    "Cloud Engineer"
  ],
  "location": "Berlin"
}' |
apify call wonderful_beluga/germany-job-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,wonderful_beluga/germany-job-finder"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iQy3FqqLH2jO1hbAI/builds/u97Fnr0ZI9ltvldI0/openapi.json
