# Gelbe Seiten Scraper: German Business Leads (`b2b_leads/gelbe-seiten-real-time-data-scraper`) Actor

Collect live German business data from Gelbe Seiten by city and category: names, addresses, phones, websites, ratings, and reviews — plus contact emails and social profiles for outreach teams. Rows stream to your dataset as structured JSON in real time.

- **URL**: https://apify.com/b2b\_leads/gelbe-seiten-real-time-data-scraper.md
- **Developed by:** [Emmanuel](https://apify.com/b2b_leads) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Gelbe Seiten Real-Time Data

**Live German business directory data from Gelbe Seiten (gelbeseiten.de) — streamed to your dataset in real time.**

> ⚠️ **Free-tier notice:** on a free Apify plan this Actor exports a small sample of results per run (2 rows). Upgrade to a paid Apify plan for full, unlimited output.

Search German businesses by category and city, enrich every row in place with full listing details (description, opening hours, payment methods, coordinates), collect public ratings and reviews, and discover contact emails and social profiles on business websites for outreach. Structured JSON, one row per business, delivered to your dataset as it is collected — plus optional real-time webhooks.

### Who is this for?

- **Lead generation & sales teams** building DACH-wide B2B outreach lists with phones, websites, and emails
- **Local SEO & marketing agencies** mapping competitors and branch presence city by city
- **CRM & SDR platform builders** feeding clean German business records into pipelines
- **Market researchers** tracking branch structures across German cities
- **Franchise & expansion analysts** counting and qualifying businesses per region
- **Data journalists** analyzing business density by category and region

### Use cases

1. **Build a local outreach list** — search `steuerberater` in `berlin`, enrich with emails and phones, push straight into your CRM via webhook.
2. **Map a vertical across cities** — run one search task per city (`zahnarzt` in `münchen`, `hamburg`, `köln`) and compare market density.
3. **Size a branch nationwide** — keyword searches across all of Germany to see how a trade is distributed before picking target cities.
4. **Enrich a list you already have** — paste Gelbe Seiten listing URLs and get the same structured rows back.
5. **Reputation snapshot** — collect ratings, review counts, and public review texts per business.
6. **Website-based lead enrichment** — discover contact emails and social profiles on the businesses' own websites.

### Feature matrix

| Feature | Input | What you get |
| --- | --- | --- |
| 🔎 Search 1 — City & category (default on) | `enableCitySearch` + `cityTasks` | Businesses in one category within one city — the classic branchenbuch lookup |
| 🔎 Search 2 — Keyword (nationwide) | `enableKeywordSearch` + `keywordTasks` | Businesses across all of Germany matching a free-text keyword |
| 🔗 Listing URLs | `enableScrapeByUrl` + `scrapeUrls` | Full structured row for any Gelbe Seiten listing URL |
| 🎯 Details & enrichment | `enableEnrichment` + three toggles | One master switch plus three enrichments (details / leads / reviews), all filling the same row |
| 🔔 Webhooks | `webhookUrl` + `webhookFormat` | Every row also POSTed to your URL (JSON or Slack) the moment it is saved |

Both search groups carry **prefilled demo rows** when the Actor form opens. City & category is switched on — click **Start** and rows stream into your dataset within seconds. The nationwide keyword group is off by default; flip its toggle and its demo rows are ready to run as they are.

Enrichment is **additive**: every discovered listing goes to the dataset, and when details are enabled the same row is filled in (`detailsFetched: true`). Nothing is filtered and no second rows are created — partial results are always exported.

### Input reference

#### 🔎 Search 1 — City & category (default on)

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `enableCitySearch` | boolean | `true` | Toggle for category + city searches. |
| `cityTasks[].category` | string | — | Branch to look up: `steuerberater`, `zahnarzt`, `hotel`, `autohändler`, … Plain words work. |
| `cityTasks[].city` | string | — | City: `berlin`, `münchen`, `hamburg`, `köln`, … Umlauts are fine. |
| `cityTasks[].maxResults` | integer | run limit | Max rows for this one search. Collection continues until this number is reached or the directory runs out of results — no fixed page cap. |

Example: tax advisors in Berlin, dentists in Munich —

```json
{
    "enableCitySearch": true,
    "cityTasks": [
        { "category": "steuerberater", "city": "berlin", "maxResults": 10 },
        { "category": "zahnarzt", "city": "münchen", "maxResults": 10 }
    ]
}
```

#### 🔎 Search 2 — Keyword (nationwide)

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `enableKeywordSearch` | boolean | `false` (demo rows prefilled) | Toggle for nationwide keyword searches. |
| `keywordTasks[].keyword` | string | — | Free-text term, e.g. `steuerberater`, `buchhaltung`, `dachdecker`. Searched across all of Germany, no city filter. |
| `keywordTasks[].maxResults` | integer | run limit | Max rows for this one search. Collection continues until this number is reached or the directory runs out of results. |

Example — 15 Steuerberater matches across Germany:

```json
{
    "enableKeywordSearch": true,
    "keywordTasks": [
        { "keyword": "steuerberater", "maxResults": 15 }
    ]
}
```

Each group has its own on/off switch, so you run exactly the searches you need. Enable both and city tasks run first, then keyword tasks in listed order.

#### 🔗 Listing URLs

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `enableScrapeByUrl` | boolean | `false` | Collect specific listing URLs. |
| `scrapeUrls` | string\[] | `[]` | Gelbe Seiten listing URLs (`https://www.gelbeseiten.de/gsbiz/...`). |

#### 🎯 Details & enrichment (one group, one master switch)

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `enableEnrichment` | boolean | `false` | Master switch for the whole group. Off = fast mode (results-page fields only). |
| `enableListingDetails` | boolean | `false` | Fill description, opening hours, payment methods, coordinates, fax, keywords. |
| `enableLeadDetails` | boolean | `false` | Discover emails + socials on the business website (needs a website on the row). |
| `enableReviews` | boolean | `false` | Aggregate rating + public review texts (capped per business). |

The three individual toggles sit in the same section as the master switch and only apply when it is on — one place to decide speed vs. depth.

#### ⚙️ Output & limits

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `concurrency` | integer | `2` | Parallel detail enrichments (1–4). |

#### 🔔 Webhooks

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `webhookUrl` | string | `""` | Your receiver URL (CRM, Zapier, Make, Slack incoming webhook, Google Sheets). Each saved row is also POSTed here in real time. |
| `webhookFormat` | select | `json` | `json` = the full record; `slack` = a Slack-friendly message. |

Delivery is best-effort: a failing webhook never stops a run or blocks dataset writes.

### Output field reference

One row per business (`type: "business"`, `platform: "gelbeseiten"`):

| Field | Description |
| --- | --- |
| `name`, `category`, `branch`, `description` | Business identity; description appears with full details |
| `profileUrl`, `profileId` | Listing link and listing id |
| `phone`, `fax`, `website` | Contact data as listed |
| `email`, `emails` | Contact emails discovered on the business website (lead details) |
| `socials`, `instagram`, `facebook`, `linkedin`, `twitter` | Social profiles (details/enrichment) |
| `street`, `zipCode`, `city`, `district`, `location` | Postal address |
| `lat`, `lng`, `distanceKm` | Coordinates and distance from the search center (when shown) |
| `openingHours`, `openingHoursToday` | Weekly hours and the open/closed status line |
| `paymentAccepted`, `keywords` | Payment methods, services/specialities |
| `rating`, `ratingSource`, `reviewCount`, `reviews` | Public ratings and review texts |
| `featureType` | `search` or `scrape_by_url` |
| `detailsFetched`, `leadDetails` | Enrichment flags for this row |
| `searchCity`, `searchCategory`, `searchTaskIndex`, `searchTaskLabel` | Which search produced this row |
| `scrapedAt` | ISO timestamp |

The run also writes a summary to the OUTPUT key-value entry: totals, duration, errors, spending-limit and paywall status.

### Webhook guide

1. **JSON** — set `webhookFormat: "json"`. Each record is POSTed as the same JSON object you get in the dataset. Works with Zapier, Make, n8n, your own receiver.
2. **Slack** — set `webhookFormat: "slack"` and use a Slack incoming webhook URL. You get a compact message per business: name, category, location, phone, website, email, rating, and a link to the listing.

```json
{
    "enableCitySearch": true,
    "cityTasks": [{ "category": "hotel", "city": "hamburg", "maxResults": 5 }],
    "webhookUrl": "https://hooks.slack.com/services/XXX/YYY/ZZZ",
    "webhookFormat": "slack"
}
```

### Using with AI agents (MCP)

This Actor works out of the box with the Apify MCP Server, so Claude or any MCP-capable agent can call it:

> *"Find me 10 Steuerberater in Berlin with their phone numbers and websites."*

The agent enables **Directory search**, runs the task, and reads the dataset — no code needed. See [Apify MCP docs](https://docs.apify.com/platform/integrations/mcp).

### FAQ

**Do I need my own proxies?**
No. A residential connection is configured by default. Adjust the proxy settings only if your plan or use case requires it.

**How current is the data?**
Every run queries the directory live, so results reflect the directory's state at run time — not a stored copy.

**What does the free plan get me?**
A small sample per run (2 rows) so you can verify the output shape. Paid plans export the full, uncapped result set.

**Why are some email fields empty?**
Emails appear only when the business publishes them on its own website, and only with lead details enabled. Fields are filled when discoverable — rows are never dropped.

**Can I run the same search for many cities?**
Yes — one search task per city. The per-search "Max results" governs each row; the run streams everything to one dataset.

**How deep does each search go?**
Until the per-search "Max results" is reached or the directory runs out of results for that search — whichever comes first. There is no fixed page cap.

***

### Kurzanleitung auf Deutsch

Dieser Actor sammelt Live-Daten aus dem Gelbe-Seiten-Branchenverzeichnis: Unternehmen nach Branche und Stadt, vollständige Detailangaben (Beschreibung, Öffnungszeiten, Koordinaten), öffentliche Bewertungen sowie Kontakt-E-Mails und Social-Media-Profile von den Unternehmens-Websites. Jeder Datensatz wird in Echtzeit als strukturiertes JSON in den Datensatz geschrieben — optional zusätzlich per Webhook an Ihre eigene URL.

Beispiel-Eingabe (Steuerberater in Berlin):

```json
{
    "enableCitySearch": true,
    "cityTasks": [
        { "category": "steuerberater", "city": "berlin", "maxResults": 10 }
    ]
}
```

**Hinweis für kostenlose Apify-Tarife:** Pro Lauf wird nur eine kleine Auswahl von Ergebnissen exportiert (2 Zeilen). Mit einem bezahlten Apify-Tarif erhalten Sie das vollständige, unbegrenzte Ergebnis.

Alle Felder und Optionen sind in den englischen Abschnitten oben dokumentiert — die Eingabemaske in der Apify Console führt ebenfalls auf Englisch durch die Einrichtung.

# Actor input Schema

## `enableCitySearch` (type: `boolean`):

Search a branch within a city — the classic branchenbuch lookup. On by default with a ready-to-run demo below. NOTE: free Apify plans are limited to a small sample of results per run — upgrade to a paid plan for full, unlimited data.

## `cityTasks` (type: `array`):

One row per category + city search. Click "+ Add" for every additional search. "Max results" governs that one row and collection continues until it is reached or the directory runs out of results.

## `enableKeywordSearch` (type: `boolean`):

Search the whole directory with free-text keywords, without a city filter — e.g. "steuerberater", "buchhaltung", "dachdecker". Off by default; the demo rows below are ready to run the moment you switch it on.

## `keywordTasks` (type: `array`):

One row per nationwide keyword search. Click "+ Add" for every additional keyword. "Max results" governs that one row.

## `enableScrapeByUrl` (type: `boolean`):

Collect full details for specific Gelbe Seiten listing URLs instead of (or in addition to) searches.

## `scrapeUrls` (type: `array`):

Gelbe Seiten listing URLs (https://www.gelbeseiten.de/gsbiz/...).

## `enableEnrichment` (type: `boolean`):

Master switch for the enrichments below. Off = fast mode (only the fields the results page already shows). Every result is still exported either way — enrichment fills fields in place, it never filters or drops rows.

## `enableListingDetails` (type: `boolean`):

Fill each result with description, opening hours, payment methods, coordinates, fax, and services. Applies when the master switch above is on.

## `enableLeadDetails` (type: `boolean`):

Fill contact emails and social profiles discovered on the business's own website (businesses with a website only). Applies when the master switch above is on.

## `enableReviews` (type: `boolean`):

Collect the aggregate rating and public customer reviews for each listing (capped per business). Applies when the master switch above is on.

## `concurrency` (type: `integer`):

How many detail enrichments run in parallel (1–4). Higher is faster but a bit harder on rate limits.

## `webhookUrl` (type: `string`):

Every record is always saved to the run's dataset — a webhook is an ADDITIONAL real-time push. When set, each new record is also POSTed to this URL (CRM, Zapier, Make, Google Sheets).

## `webhookFormat` (type: `string`):

json = full record object; slack = Slack-friendly message payload.

## `proxyConfiguration` (type: `object`):

Apify residential proxy is enabled by default for reliable results.

## Actor input object example

```json
{
  "enableCitySearch": true,
  "cityTasks": [
    {
      "category": "steuerberater",
      "city": "berlin",
      "maxResults": 10
    },
    {
      "category": "zahnarzt",
      "city": "münchen",
      "maxResults": 10
    }
  ],
  "enableKeywordSearch": false,
  "keywordTasks": [
    {
      "keyword": "buchhaltung",
      "maxResults": 10
    }
  ],
  "enableScrapeByUrl": false,
  "scrapeUrls": [],
  "enableEnrichment": false,
  "enableListingDetails": false,
  "enableLeadDetails": false,
  "enableReviews": false,
  "concurrency": 2,
  "webhookUrl": "",
  "webhookFormat": "json",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `allResults` (type: `string`):

Complete dataset with every row from all enabled features in this run.

## `search` (type: `string`):

Rows from directory search.

## `runSummary` (type: `string`):

Run metadata: totals, duration, errors, spending-limit and free-tier (paywall) status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "enableCitySearch": true,
    "cityTasks": [
        {
            "category": "steuerberater",
            "city": "berlin",
            "maxResults": 10
        },
        {
            "category": "zahnarzt",
            "city": "münchen",
            "maxResults": 10
        }
    ],
    "enableKeywordSearch": false,
    "keywordTasks": [
        {
            "keyword": "buchhaltung",
            "maxResults": 10
        }
    ],
    "enableEnrichment": false,
    "enableListingDetails": false,
    "enableLeadDetails": false,
    "enableReviews": false,
    "concurrency": 2,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("b2b_leads/gelbe-seiten-real-time-data-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "enableCitySearch": True,
    "cityTasks": [
        {
            "category": "steuerberater",
            "city": "berlin",
            "maxResults": 10,
        },
        {
            "category": "zahnarzt",
            "city": "münchen",
            "maxResults": 10,
        },
    ],
    "enableKeywordSearch": False,
    "keywordTasks": [{
            "keyword": "buchhaltung",
            "maxResults": 10,
        }],
    "enableEnrichment": False,
    "enableListingDetails": False,
    "enableLeadDetails": False,
    "enableReviews": False,
    "concurrency": 2,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("b2b_leads/gelbe-seiten-real-time-data-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "enableCitySearch": true,
  "cityTasks": [
    {
      "category": "steuerberater",
      "city": "berlin",
      "maxResults": 10
    },
    {
      "category": "zahnarzt",
      "city": "münchen",
      "maxResults": 10
    }
  ],
  "enableKeywordSearch": false,
  "keywordTasks": [
    {
      "keyword": "buchhaltung",
      "maxResults": 10
    }
  ],
  "enableEnrichment": false,
  "enableListingDetails": false,
  "enableLeadDetails": false,
  "enableReviews": false,
  "concurrency": 2,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call b2b_leads/gelbe-seiten-real-time-data-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,b2b_leads/gelbe-seiten-real-time-data-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cgUAlgnP67cV7kZtl/builds/AtE0qYaWdKNplFX3G/openapi.json
