# Ausbildung Jobs Search Scraper (`soft_alexist/ausbildung-jobs-search-scraper`) Actor

Scrape German apprenticeship listings from Ausbildung.de effortlessly. This scraper extracts job titles, company details, locations, start dates, salary info, and 35+ fields per vacancy — perfect for talent agents, HR teams, and education researchers analyzing the German Ausbildung market.

- **URL**: https://apify.com/soft\_alexist/ausbildung-jobs-search-scraper.md
- **Developed by:** [Soft Alexist](https://apify.com/soft_alexist) (community)
- **Categories:** Automation, Developer tools, Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are a software tools running on the Apify platform, for all kinds of web data extraction and automation use cases.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js/docs.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python/docs.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

## Ausbildung.de Jobs Scraper: Automate Apprenticeship Data Collection

---

### Understanding Ausbildung.de

Ausbildung.de is Germany's largest apprenticeship job board, connecting employers seeking trainees with prospective apprentices across all professions and regions. The platform hosts thousands of active Ausbildung (apprenticeship) positions ranging from crafts to IT, hospitality to healthcare. Manually browsing and extracting this data is labor-intensive — the **Ausbildung.de Jobs Scraper** automates the entire process, delivering structured, ready-to-analyze apprenticeship data at scale.

---

### Scraper Overview

The **Ausbildung.de Jobs Search Scraper** extracts vacancy listings from Ausbildung.de search results pages, converting raw HTML into clean, structured data. It's built for:

- **Vocational educators** tracking apprenticeship trends and employer demand
- **Recruitment agencies** managing candidate matches and market intelligence
- **Career researchers** studying vocational education pathways in Germany
- **HR professionals** benchmarking apprenticeship programs across industries
- **Data analysts** building intelligence platforms around German talent markets

Key features include flexible search filtering by profession, configurable item limits, and robust error handling via `ignore_url_failures` — allowing you to scrape hundreds of listings reliably in a single run.

---

### Input Format & Configuration

The scraper accepts a JSON configuration object controlling search parameters and extraction limits:

```json
{
  "urls": [
    "https://www.ausbildung.de/suche/?professions=Pflegefachmann%2F-frau"
  ],
  "ignore_url_failures": true,
  "max_items_per_url": 200
}
````

| Parameter | Purpose | Example |
|---|---|---|
| `urls` | Direct links to Ausbildung.de search results filtered by profession, location, or keyword | `https://www.ausbildung.de/suche/?professions=Pflegefachmann%2F-frau` (Nursing Apprenticeships) |
| `max_items_per_url` | Maximum vacancies to extract per search results page (1–200) | `200` extracts up to 200 listings per URL |
| `ignore_url_failures` | If `true`, scraper continues if a URL fails; if `false`, stops on first error | `true` recommended for large batch runs |

**Tips for URL construction:**

- Use Ausbildung.de's search UI to build filtered URLs (by profession, region, company)
- Paste the resulting URL into the `urls` array
- Each URL can yield up to 200 listings depending on search breadth

***

### Output Format & Field Definitions

**Sample output**

```json
{
  "corporation_name": "maxQ. im bfw – Unternehmen für Bildung. / Berufsfortbildungswerk Gemeinnützige Bildungseinrichtung des DGB GmbH (bfw)",
  "vacancy_public_id": "47515630-ae75-411b-83b4-312898cf5d5b",
  "slug": "ausbildung-pflegefachmann-frau-m-w-d-bei-maxq-im-bfw-unternehmen-fuer-bildung-berufsfortbildungswerk-gemeinnuetzige-bildungseinrichtung-des-dgb-gmbh-bfw-in-frankfurt-am-main-47515630-ae75-411b-83b4-312898cf5d5b",
  "title": "Ausbildung Pflegefachmann/-frau (m/w/d)",
  "vacancy_count": 18,
  "location": "60329 Frankfurt am Main",
  "related_branches_count": 17,
  "starts_no_earlier_than": null,
  "corporation_logo": "https://www.ausbildung.de/uploads/image/e2/e2af9f98-8fe1-45cc-a34c-6476f21c64c8/Logo-quadrat.png",
  "direct_application_on": false,
  "subsidiary_logo": null,
  "subsidiary_name": "maxQ. im bfw – Unternehmen für Bildung. / Berufsfortbildungswerk Gemeinnützige Bildungseinrichtung des DGB GmbH (bfw)",
  "in_spotlight": false,
  "corporation_display_vacancy_counts": true,
  "profession_title": "Pflegefachmann/-frau",
  "application_options": "online",
  "salesforce_category": "A",
  "non_eu_flow": 1,
  "subsidiary_public_id": "161d69e1-2c5a-4d07-8229-1464ed533ce1",
  "corporation_public_id": "49b03837-7317-4e1f-beec-e2d3ea48d992",
  "corporation_starving_state": 0,
  "valid_until": null,
  "expected_graduation": "mittlerer-schulabschluss",
  "apprenticeship_type": "klassische-duale-berufsausbildung",
  "duration": "3 Jahre",
  "ba_booking": "FALSE",
  "corporation_recent_review_count": 0,
  "corporation_recent_star_count_avg": 0,
  "top_rated_employer": false,
  "id": 1139424,
  "subsidiary_id": 134644,
  "profession_id": 668,
  "created_at": "2023-12-11T13:40:05.777Z",
  "updated_at": "2023-12-11T13:40:05.777Z",
  "from_url": "https://www.ausbildung.de/suche/?professions=Pflegefachmann/-frau"
}
```

Each scraped vacancy returns a comprehensive record with 35+ fields covering employer, position, logistics, and candidate fit:

#### Employer & Company Data

| Field | Meaning |
|---|---|
| `Corporation Name` | Official legal name of the hiring employer |
| `Corporation Public ID` | Unique Ausbildung.de identifier for the employer |
| `Corporation Logo` | URL to employer's company logo |
| `Corporation Display Vacancy Counts` | Whether the employer publicly shows vacancy counts |
| `Corporation Starving State` | Current hiring status flag (e.g., active, paused) |
| `Corporation Recent Review Count` | Number of recent trainee reviews for the company |
| `Corporation Recent Star Count Avg` | Average star rating (1–5) from recent trainee feedback |
| `Top Rated Employer` | Boolean flag: `true` if company meets Ausbildung.de quality standards |
| `Subsidiary Name` | Branch or subsidiary name (if applicable) |
| `Subsidiary Logo` | Logo for subsidiary entity |
| `Subsidiary Public ID` | Unique ID for the subsidiary |
| `Subsidiary ID` | Internal subsidiary database identifier |

#### Vacancy & Position Details

| Field | Meaning |
|---|---|
| `Vacancy Public ID` | Unique identifier for the specific apprenticeship listing |
| `ID` | Internal database ID for the vacancy |
| `Title` | Job title or apprenticeship role name (e.g., "Pflegefachmann") |
| `Slug` | URL-friendly version of the title |
| `Profession Title` | Standardized apprenticeship profession name in German |
| `Profession ID` | Internal ID for the profession category |
| `Vacancy Count` | Number of open positions within this listing |
| `Apprenticeship Type` | Type of training program (e.g., "dual apprenticeship", "school-based") |
| `Duration` | Training period length (typically 2–3.5 years) |
| `Salesforce Category` | Internal Ausbildung.de categorization |

#### Logistics & Applicant Information

| Field | Meaning |
|---|---|
| `Location` | City/region where apprenticeship takes place |
| `Starts No Earlier Than` | Earliest possible start date for the apprenticeship |
| `Expected Graduation` | Estimated apprenticeship completion date |
| `Valid Until` | Deadline for submitting applications |
| `Created At` | Date the listing was posted to Ausbildung.de |
| `Updated At` | Last modification timestamp |

#### Application & Special Flags

| Field | Meaning |
|---|---|
| `Application Options` | Available application methods (e.g., online form, email, direct contact) |
| `Direct Application On` | Whether direct online application is enabled |
| `In Spotlight` | Boolean: `true` if listing is featured/promoted |
| `BA Booking` | Booking status for Federal Employment Agency integration |
| `Non EU Flow` | Flag for non-EU applicant eligibility |
| `Related Branches Count` | Number of related apprenticeship branches offered by employer |

***

### How to Use the Scraper

1. **Navigate to Ausbildung.de** — Visit the website and use the search UI to filter by profession, location, or company.
2. **Copy the search URL** — Once results match your criteria, copy the full URL from your browser's address bar.
3. **Paste into configuration** — Add the URL to the `urls` array in your scraper input.
4. **Set extraction limits** — Adjust `max_items_per_url` based on your needs (e.g., `50` for quick samples, `200` for comprehensive datasets).
5. **Enable error tolerance** — Set `ignore_url_failures: true` for large batch runs to prevent interruptions.
6. **Execute and export** — Run the scraper and download results as JSON, CSV, or Excel.

**Best practices:**

- Start with `max_items_per_url: 50` to test your URLs
- Use specific profession filters for cleaner, more relevant data
- Schedule regular scrapes to track vacancies over time
- Store results in a database for trend analysis

***

### Real-World Applications

- **Vocational counseling:** Analyze which apprenticeships offer the most positions and regional opportunities
- **Market research:** Track employer demand for specific trades (e.g., nursing, IT, construction)
- **Competitive intelligence:** Monitor salary trends, company reviews, and hiring practices across sectors
- **Integration projects:** Feed German apprenticeship data into career portals, chatbots, or educational platforms
- **Policy analysis:** Build datasets supporting vocational education research and workforce planning

By automating data extraction, the Ausbildung.de Jobs Scraper transforms hundreds of hours of manual work into seconds, enabling data-driven decisions in talent, education, and HR.

***

### Conclusion

The **Ausbildung.de Jobs Search Scraper** is an essential tool for anyone navigating the German apprenticeship landscape at scale. With 35+ output fields capturing employer reputation, training logistics, and application details, it delivers the rich, structured data needed for intelligent hiring, research, and business strategy.

# Actor input Schema

## `urls` (type: `array`):

Add the URLs of the Jobs list urls you want to scrape. You can paste URLs one by one, or use the Bulk edit section to add a prepared list.

## `ignore_url_failures` (type: `boolean`):

If true, the scraper will continue running even if some URLs fail to be scraped.

## `max_items_per_url` (type: `integer`):

The maximum number of items to scrape per URL.

## Actor input object example

```json
{
  "urls": [
    "https://www.ausbildung.de/suche/?professions=Pflegefachmann%2F-frau"
  ],
  "ignore_url_failures": true,
  "max_items_per_url": 20
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.ausbildung.de/suche/?professions=Pflegefachmann%2F-frau"
    ],
    "ignore_url_failures": true,
    "max_items_per_url": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("soft_alexist/ausbildung-jobs-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.ausbildung.de/suche/?professions=Pflegefachmann%2F-frau"],
    "ignore_url_failures": True,
    "max_items_per_url": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("soft_alexist/ausbildung-jobs-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.ausbildung.de/suche/?professions=Pflegefachmann%2F-frau"
  ],
  "ignore_url_failures": true,
  "max_items_per_url": 20
}' |
apify call soft_alexist/ausbildung-jobs-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=soft_alexist/ausbildung-jobs-search-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "Ausbildung Jobs Search Scraper",
        "description": "Scrape German apprenticeship listings from Ausbildung.de effortlessly. This scraper extracts job titles, company details, locations, start dates, salary info, and 35+ fields per vacancy — perfect for talent agents, HR teams, and education researchers analyzing the German Ausbildung market.",
        "version": "0.0",
        "x-build-id": "HhMDCrgTW1HJhCRw8"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/soft_alexist~ausbildung-jobs-search-scraper/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-soft_alexist-ausbildung-jobs-search-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/soft_alexist~ausbildung-jobs-search-scraper/runs": {
            "post": {
                "operationId": "runs-sync-soft_alexist-ausbildung-jobs-search-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/soft_alexist~ausbildung-jobs-search-scraper/run-sync": {
            "post": {
                "operationId": "run-sync-soft_alexist-ausbildung-jobs-search-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "properties": {
                    "urls": {
                        "title": "URLs of the Jobs list urls to scrape",
                        "type": "array",
                        "description": "Add the URLs of the Jobs list urls you want to scrape. You can paste URLs one by one, or use the Bulk edit section to add a prepared list.",
                        "items": {
                            "type": "string"
                        }
                    },
                    "ignore_url_failures": {
                        "title": "Continue running even if some URLs fail to be scraped",
                        "type": "boolean",
                        "description": "If true, the scraper will continue running even if some URLs fail to be scraped."
                    },
                    "max_items_per_url": {
                        "title": "Max items per URL",
                        "type": "integer",
                        "description": "The maximum number of items to scrape per URL."
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
