# Impressum Scraper – Company Data, HRB & VAT (DACH) (`joseph_cloud/impressum-company-data`) Actor

Turn DE/AT/CH website lists into company records from the Impressum: legal name, legal form, address, HRB/FN register + court, VAT ID checked in EU VIES, managing directors, email, phone. Measured accuracy. $7 per 1,000.

- **URL**: https://apify.com/joseph\_cloud/impressum-company-data.md
- **Developed by:** [Abdülkadir Tekin](https://apify.com/joseph_cloud) (community)
- **Categories:** Lead generation, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$7.00 / 1,000 company record extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Impressum Scraper – Company Data, HRB & VAT (DACH)

Give it a list of **German, Austrian or Swiss websites** and get back the **legal company record** behind each one, read from its Impressum (the legal notice every DACH business site must publish): legal name, legal form, address, commercial register number and court, VAT ID — **checked live against the EU VIES register** — managing directors, email and phone.

It finds the Impressum page by itself, works without a browser (fast and cheap), and you pay only for websites that actually gave company data.

### What you get for each website

| Field | Example | Meaning |
|---|---|---|
| `companyName` | `NetCologne Gesellschaft für Telekommunikation mbH` | Legal name of the site operator |
| `legalForm` | `GmbH` | Normalised legal form: GmbH, UG (haftungsbeschränkt), AG, GmbH & Co. KG, KG, OHG, e.K., e.V., eG, public law body, … |
| `entityType` | `company` | `company`, `partnership`, `sole_trader`, `association`, `foundation`, `public_body` or `unknown` |
| `street`, `postalCode`, `city`, `country` | `Am Coloneum 9`, `50829`, `Köln`, `DE` | Registered address from the Impressum |
| `registerType`, `registerNumber`, `registerCourt` | `HRB`, `HRB 25580`, `Amtsgericht Köln` | German HRB/HRA/VR/GnR/PR, Austrian Firmenbuch (FN), Swiss UID |
| `vatId` | `DE811808435` | USt-IdNr. / ATU / CHE … MWST, normalised without spaces |
| `vatValid`, `vatRegisteredName`, `vatRegisteredAddress` | `true` | Result of the live EU VIES check (Germany does not return name/address) |
| `representatives` | `[{"name":"Timo von Lepel","role":"managing_director"}]` | Managing directors, board members, owners |
| `emails`, `phones`, `fax` | `info@netcologne.de` | Contact data published in the Impressum |
| `containsPersonalData` | `true` | The record names a natural person (see Data protection) |
| `impressumUrl` | `https://www.netcologne.de/impressum/` | Page the data came from — easy to check |
| `error` | `No Impressum / imprint page found, not charged` | Why a website gave no data |

### Use cases

- **B2B lead enrichment:** turn a list of domains into company names, addresses and decision makers for your CRM.
- **KYB and supplier checks:** confirm the legal entity, register entry and a valid VAT ID behind a webshop or partner.
- **Invoicing and master data:** fill in exact legal names and VAT IDs before you invoice.
- **Market research:** see which legal forms, regions or register courts dominate a list of sites.
- **AI agents:** works as a tool through the Apify MCP server — give it domains, get structured JSON back.

### Input

```json
{
  "urls": ["netcologne.de", "https://www.wnt.at", "bekb.ch"],
  "checkVat": true
}
```

### Measured accuracy

We tested the Actor on **50 randomly chosen DACH websites it had never seen during development**. 31 of them gave a company record (and would have been charged); the other 19 had no reachable Impressum, blocked bots or showed no company data (free). Every field of the 31 records was then checked one by one against the Impressum page it came from.

| Field | Accuracy | Precision |
|---|---|---|
| Company name | 83% | 93% |
| Legal form | 83% | 92% |
| Address | 87% | 86% |
| Register number + court | 77% | 73% |
| VAT ID | 94% | 100% |
| Representatives | 77% | 84% |
| Emails | 100% | 100% |
| Phones | 97% | 100% |
| **All fields** | **87%** | **91%** |

*Accuracy* = share of fields that came out right, including fields correctly left empty because the page has no such data. *Precision* = when a value is returned, how often it is fully correct. Partly right values (for example a register number without its court) count as not correct.

Known weak spots: GmbH & Co. KG pages that describe the general partner at length (the partner can be returned instead of the KG), Austrian Firmenbuch courts written in unusual ways, and sites that list a branch office or a P.O. box before the registered address.

### Pricing

Pay per event: **$7 per 1,000 company records**. There is no extra charge per run, and the VIES check is included.

You are **not charged** when a website does not exist or cannot be reached, blocks automated requests, shows a bot-protection page, disallows bots in robots.txt, has no Impressum we can find, or has an Impressum without any company name, register number or VAT ID. Such items are still returned, with the reason in `error`.

### How it works, and its limits

The Actor loads the homepage, follows the Impressum / Imprint / Legal notice / Offenlegung link (or tries the usual paths such as `/impressum`), and parses the page text. Sites that render their Impressum with JavaScript are often still read from the page's embedded data.

Please know the limits before you buy:

- **No browser is used.** An Impressum that only appears after JavaScript runs, or that is shown as an image, can be missed (you are not charged then).
- **Some sites block automated requests.** We respect that and robots.txt and do not try to get around it.
- **The Impressum is the source of truth.** If a site publishes an outdated address or an old managing director, you get exactly that. Use `impressumUrl` to check.
- **People are read from the operator's block only.** Editors, data protection officers and web agencies further down the page are ignored on purpose.

### Data protection

An Impressum is published because the law requires it, so that customers can identify and contact a business. For sole traders (Einzelunternehmen) and small partnerships it names a natural person, and managing directors are named for every company. Such data is personal data under the GDPR. Records that contain it are marked with `containsPersonalData: true`, so you can filter them out or treat them separately.

You are responsible for having a legal basis for what you do with the data — for example, German law (§ 7 UWG) restricts unsolicited advertising emails. Use the data for legitimate business purposes only.

### Support

Open an issue on the Actor's Issues tab; it is usually answered within one to two days. If a field comes out wrong for a site, include the URL — every report makes the parser better.

# Actor input Schema

## `urls` (type: `array`):

Domains or URLs of German, Austrian or Swiss websites (e.g. zalando.de or https://www.example.at). The Actor finds the Impressum / imprint page itself. Duplicates are removed.

## `checkVat` (type: `boolean`):

Checks every EU VAT ID found (DE, AT, ...) against the European Commission's VIES service and adds vatValid. Swiss UIDs are not in VIES.

## Actor input object example

```json
{
  "urls": [
    "netcologne.de",
    "narayana-verlag.de",
    "niedax.de",
    "wnt.at",
    "bekb.ch"
  ],
  "checkVat": true
}
```

# Actor output Schema

## `results` (type: `string`):

One item per website: legal name and form, address, register number and court, VAT ID with VIES check, representatives, emails, phones.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "netcologne.de",
        "narayana-verlag.de",
        "niedax.de",
        "wnt.at",
        "bekb.ch"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("joseph_cloud/impressum-company-data").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "netcologne.de",
        "narayana-verlag.de",
        "niedax.de",
        "wnt.at",
        "bekb.ch",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("joseph_cloud/impressum-company-data").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "netcologne.de",
    "narayana-verlag.de",
    "niedax.de",
    "wnt.at",
    "bekb.ch"
  ]
}' |
apify call joseph_cloud/impressum-company-data --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,joseph_cloud/impressum-company-data"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pj23MUbJeolMeHrpK/builds/bsdlXqX8oHlM3Pwpl/openapi.json
