# Imprint & Impressum Scraper — Email, Phone, VAT & Company Data (`perceptr0n/imprint-impressum-website-contact-scraper`) Actor

Domains to verified company records from the imprint / Impressum / legal notice: name, legal form, address, register (HRB, FN, CHE), VAT ID checked against the EU VIES register, e-mail, phone, fax, social profiles. Germany, Austria, Switzerland. No login, no cookies, no API key.

- **URL**: https://apify.com/perceptr0n/imprint-impressum-website-contact-scraper.md
- **Developed by:** [Perceptron Data](https://apify.com/perceptr0n) (community)
- **Categories:** Lead generation
- **Stats:** 7 total users, 3 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.01 / actor start

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Imprint & Impressum Scraper — company data, emails, phones, VAT and socials

Give it a list of domains, get back a clean company record for each one. The
data comes from the **imprint** — the *Impressum*, *legal notice*, *Offenlegung*
— that every commercial website in **Germany, Austria and Switzerland** is
required by law to publish, which is why it can be read reliably instead of
guessed.

Every VAT ID found is then **verified against the European Commission's official
VIES register**, so you learn not just what a company claims, but whether that
number is actually registered.

No login, no cookies, no API key, no anti-bot fight. **$5 per 1,000 websites.**

> Want to try it first? The **[free edition](https://apify.com/perceptr0n/free-website-contact-imprint-scraper)** returns the identical row for up to 10 websites per run — no credit card, nothing held back.

***

### Two ways in: a list, or a search

**You have the domains.** Paste them — one company record comes back per entry.

**You don't have the domains.** Type the trade and the place instead:

```json
{ "search": "Zahnärzte in Hamburg", "max_companies": 25 }
```

The businesses are looked up in **OpenStreetMap**, and every website found is
then enriched from its imprint. About 180 trades are recognised, in German and
English — `dentist`, `Steuerberater`, `roofer`, `Bäckerei`, `Autohaus`,
`Rechtsanwalt`, `Hotel`, `Elektriker`, `Physiotherapie` and so on, including
plurals. Places can be a city, a district, a county or a whole state
(`Hamburg`, `Landkreis Rosenheim`, `Bayern`).

For a trade the list does not know, pass the OpenStreetMap tag yourself:

```json
{ "osm_filter": "[\"office\"=\"company\"]", "location": "Wien" }
```

#### What a search actually returns

Nine searches across trades, city sizes and all three countries, run on the
platform on 28 September 2026, five companies requested each:

| Search | Rows | With name | With e-mail | With VAT ID | Runtime |
|---|---|---|---|---|---|
| Steuerberater Wien 🇦🇹 | 5 | 5 | 4 | 3 | 103 s |
| Rechtsanwälte München | 5 | 5 | 4 | 3 | 102 s |
| Apotheken Leipzig | 5 | 5 | 2 | 2 | 130 s |
| Autohäuser Stuttgart | 5 | 5 | 3 | 2 | 53 s |
| Elektriker Köln | 5 | 5 | 3 | 2 | 99 s |
| Restaurants Zürich 🇨🇭 | 5 | 5 | 4 | 2 | 28 s |
| Hotels Berlin | 5 | 4 | 1 | 1 | 34 s |
| Dachdecker Landkreis Rosenheim | **1** | 1 | 1 | 1 | 13 s |
| Bäckereien **Bayern** (whole state) | **failed** | — | — | — | 104 s |

What that table says, plainly: **city-sized searches work, rural districts are
thin, and a whole federal state is too much for the free query service.** Every
row always carries the website and what the map knows; the imprint fields
depend on the site actually publishing one. Swiss VAT numbers are never marked
valid because `CHE-` identifiers are outside the EU's VIES register.

#### What to expect from the search, honestly

- **Not every business is in OpenStreetMap, and not every entry has a website.**
  Measured on 28 September 2026: hotels in Berlin 76 % with a website, dentists
  in Hamburg 52 %, bakeries in Munich 35 %. The search finds what the map knows,
  which is a lot for local trades and little for pure office businesses.
- **OpenStreetMap's query service is volunteer-run and throttles hard.** A busy
  moment produces a clear message rather than an empty result, and the Actor
  retries across three mirrors within a fixed 100-second budget. A city answers
  where a federal state times out.
- **Attribution travels with the data.** Every discovered row carries
  `dataSource: "© OpenStreetMap contributors"` and `dataLicense: "ODbL 1.0"`,
  because that licence requires it — of you as well, if you republish.
- The map fields (`osmName`, `osmCategory`, `osmCity`, `osmPhone`) stay in the
  row next to the imprint fields, so you can see what came from where. Where
  they disagree, the imprint wins: it is the legally binding source.

### What you get per website

| Field | Example |
|---|---|
| `companyName`, `legalForm` | `Liebherr-International Deutschland GmbH`, `GmbH` |
| `street`, `postalCode`, `city`, `country` | `Hans-Liebherr-Straße 45`, `88400`, `Biberach an der Riß`, `DE` |
| `registerType`, `registerNumber`, `registerCourt` | `HRB`, `HRB 720327`, `Ulm` |
| `vatId`, `vatCountry` | `DE811907980`, `DE` |
| `vatChecked`, `vatValid` | `true`, `true` — checked live against VIES |
| `vatRegisteredName`, `vatRegisteredAddress` | the name the tax register holds, where the member state discloses it |
| `email`, `allEmails` | `info@liebherr.com`, up to 5 addresses |
| `phone`, `allPhones`, `fax` | up to 5 numbers |
| `socialProfiles` | LinkedIn, Xing, Facebook, Instagram, X, YouTube, TikTok |
| `impressumUrl`, `impressumFound`, `foundVia` | where the data came from, verbatim |
| `completeness`, `dataQuality` | 0–1 and `high` / `medium` / `low` |
| `representatives` | managing directors — **only if you switch them on** |

### Three countries, three legal bases, one schema

| Country | Duty to publish | Register identifier the scraper reads |
|---|---|---|
| 🇩🇪 Germany | § 5 DDG (successor to § 5 TMG) | `HRB`, `HRA`, `VR`, `GnR`, `PR` + Amtsgericht |
| 🇦🇹 Austria | § 5 ECG, § 25 MedienG | `FN 123456a` + Firmenbuchgericht |
| 🇨🇭 Switzerland | UWG Art. 3(1)(s) | `CHE-123.456.789` + Handelsregister of the canton |

Four-digit Austrian and Swiss postal codes are handled as well — a German-only
parser silently drops those addresses.

### The VAT check is the part competitors skip

An imprint *states* a VAT ID. Whether it is registered is a different question,
and only the member state can answer it. This actor asks
[VIES](https://ec.europa.eu/taxation_customs/vies/) for every ID it finds:

- `vatValid: false` on a live company is a genuine red flag for invoicing
- Austria, Italy, Spain and most other member states also return the
  **registered name and address**, so you can cross-check the imprint against
  the register in the same row
- Germany discloses validity only — no name, no address. That is a limit of the
  German tax authority, not of this actor, and it is reported as such
- Swiss `CHE-` numbers are outside the EU system and are never checked

It costs you nothing extra and adds roughly a quarter of a second per company.

### Input

```json
{
  "websites": ["liebherr.com", "frischpack.de", "tuwien.at", "sbb.ch"],
  "validate_vat": true,
  "allow_contact_page": true,
  "include_person_names": false
}
```

Paste domains, full URLs or a messy copy-paste — they are normalised. One row
comes back per entry, always, so your list and the result line up.

### Output (shortened)

```json
{
  "input": "tuwien.at",
  "website": "https://www.tuwien.at",
  "country": "AT",
  "companyName": "Technische Universität Wien",
  "street": "Karlsplatz 13",
  "postalCode": "1040",
  "city": "Wien",
  "vatId": "ATU37675002",
  "vatChecked": true,
  "vatValid": true,
  "vatRegisteredName": "Technische Universität Wien",
  "vatRegisteredAddress": "Karlsplatz 13 AT-1040 Wien",
  "email": "office@tuwien.ac.at",
  "socialProfiles": { "linkedin": "https://www.linkedin.com/school/tu-wien" },
  "impressumFound": true,
  "foundVia": "footer link",
  "dataQuality": "high"
}
```

### What does a run cost?

| Websites | Cost |
|---|---|
| 100 | $0.51 |
| 1,000 | **$5.01** |
| 10,000 | $50.01 |

$0.01 per run start, $0.005 per website returned. You are charged per **result
row**, and a row that could not be resolved still tells you why — see below.

### Honesty about coverage

Not every website has a reachable imprint, and pretending otherwise would waste
your money. Three things are reported rather than hidden:

- **`impressumFound: false`** — the site publishes no legal notice, or none that
  could be reached. For a German site that is itself a finding: publishing one
  is mandatory.
- **`foundVia: "contact page"`** — no imprint existed, so the contact page was
  read instead. You get the address and mail box, but the row is *not* marked as
  a legal notice. Switch `allow_contact_page` off if you only want real imprints.
- **`foundVia: "blocked"`** — the site is reachable but turns away automated
  clients (401, 403, 405, 406, 429). No cloud scraper will get its legal notice,
  and you can see that instead of guessing why the row is empty.
- **`foundVia: "timeout"`** — the site did not answer within 45 seconds. Each
  website gets that hard budget, so one stalling host can never hold up your
  run or inflate its cost.

You never get an empty row without an explanation of why it is empty.

### Legal and data protection

- Company data — name, legal form, address, register number, VAT ID — is **not**
  personal data and is published under a statutory duty.
- **Person names are off by default.** Managing directors and editorial contacts
  are personal data under GDPR; `include_person_names` must be switched on
  deliberately, and the input form says so.
- Only publicly reachable pages are fetched, at a polite request rate, with no
  login and no circumvention of any access control.
- Using contact data for cold e-mail in Germany requires prior consent under
  § 7 UWG. That is your call to make, not this actor's — but you should know it
  before you buy a list of addresses from anyone.

### Auf Deutsch: Impressum-Daten aus Websites extrahieren

Dieser Actor liest das **Impressum** von Firmen-Websites aus Deutschland,
Österreich und der Schweiz und macht daraus je Website einen sauberen
Firmendatensatz: Firmenname, Rechtsform, Anschrift, **Handelsregister** (HRB/HRA
mit Amtsgericht, österreichisches Firmenbuch, Schweizer UID), **USt-IdNr. live
gegen das EU-VIES-Register geprüft**, E-Mail, Telefon, Fax und Social-Media-Profile.

Ohne eigene Domainliste genügt eine Suche wie `Zahnärzte in Hamburg`: Die
Firmen werden über OpenStreetMap gefunden und anschließend aus ihrem Impressum
angereichert.

Typische Anwendungsfälle: Adresslisten bereinigen und anreichern,
Lieferanten- und Kundenstammdaten prüfen, USt-IdNr. verifizieren, CRM-Pflege.
Hinweis: Werbe-E-Mails an Firmen brauchen in Deutschland nach § 7 UWG eine
vorherige Einwilligung.

### FAQ

**Is there an API for German company data?**
Not a free official one. The imprint is the only company record that every
German, Austrian and Swiss business must publish itself, which is what makes
this approach work without a paid register licence.

**What is an Impressum?**
The legal notice a commercial website must carry: who runs it, where they are
registered, how to reach them. In Germany it is required by § 5 DDG, in Austria
by § 5 ECG, in Switzerland by UWG Art. 3(1)(s).

**Does it work for Austrian and Swiss sites?**
Yes — Firmenbuch numbers, Swiss UIDs and four-digit postal codes are all parsed,
and the `country` field tells you which set of rules applied.

**Can I get the managing director's name?**
Yes, via `include_person_names`, but read the GDPR note first. It is off by
default on purpose.

**Is the VAT number checked or just copied?**
Checked, live, against the EU's official VIES service — and `vatChecked` tells
you whether the check actually ran.

**What happens if a website has no imprint?**
You get a row with `impressumFound: false` and, if you allowed it, whatever the
contact page held. You are never left guessing.

**Can I feed it 50,000 domains?**
Yes. Requests run concurrently with a polite delay; cost scales linearly at
$5 per 1,000.

**Does it need a proxy?**
No. No login, no cookies, no API key, no CAPTCHA.

### Related actors

| Actor | What it does |
|---|---|
| [Free Website Contact & Imprint Scraper](https://apify.com/perceptr0n/free-website-contact-imprint-scraper) | Free edition: up to 10 websites per run, same fields |
| [Google Ads Transparency Scraper](https://apify.com/perceptr0n/google-ads-transparency-scraper) | Which companies advertise on Google, how many ads, who runs them — plus every ad and a monitoring mode |
| [PDF & DOCX to Markdown Converter](https://apify.com/perceptr0n/pdf-to-markdown-converter) | PDF, Word, Excel and PowerPoint to Markdown for AI — tables as JSON, OCR for scans |
| [Arbeitsagentur Jobs Scraper](https://apify.com/perceptr0n/arbeitsagentur-jobs-scraper) | Jobs from the Bundesagentur für Arbeit Jobbörse via the official API |
| [kununu Scraper](https://apify.com/perceptr0n/kununu-employer-data-scraper) | Employer reviews, ratings, salaries and benefits for DACH companies |
| [TED Tenders API](https://apify.com/perceptr0n/eu-tenders-ted-scraper) | EU public tenders from TED — all 27 member states, CPV and deadline filters |
| [Marktstammdatenregister (MaStR) Scraper](https://apify.com/perceptr0n/german-energy-assets-mastr) | Solar, wind and battery storage plants in Germany with operator and capacity |
| [EUDAMED Scraper](https://apify.com/perceptr0n/eudamed-medical-device-scraper) | EU medical devices, manufacturers and notified-body certificates (MDR/IVDR) |
| [Messe München Exhibitor List Scraper](https://apify.com/perceptr0n/messe-muenchen-exhibitor-list-scraper) | Exhibitor lists of bauma, IFAT, BAU, automatica, ceramitec and more |
| [Koelnmesse Exhibitor List Scraper](https://apify.com/perceptr0n/anuga-ism-koelnmesse-exhibitor-scraper) | Exhibitor lists of Anuga, ISM, imm cologne, interzum, ORGATEC and more |
| [App Store & Google Play Scraper](https://apify.com/perceptr0n/appstore-google-play-intelligence) | Ratings, reviews and keyword rankings from both app stores in one dataset |
| [DACH & EU Data Source Finder](https://apify.com/perceptr0n/dach-eu-data-source-finder) | Free: tells you which ready-made scraper covers your data source |

# Actor input Schema

## `search` (type: `string`):

Type what you are looking for instead of pasting domains — for example `Zahnärzte in Hamburg`, `dentists in Hamburg`, `Steuerberater München`, `roofers Cologne`. The businesses are looked up in **OpenStreetMap** (© OpenStreetMap contributors, ODbL) and every website found is then enriched from its imprint. About 180 trades are recognised in German and English. Leave empty to use the website list below instead. (Deliberately left empty by default: a run with no input uses the website list below, so an automated run never depends on OpenStreetMap being reachable.)

## `max_companies` (type: `integer`):

Only applies to the search above. Each company found is fetched and enriched, so this is also the cost lever: 25 companies cost about $0.14.

## `websites` (type: `array`):

Domains or URLs — e.g. `liebherr.com`, `https://www.sap.com`, `tuwien.at`, `sbb.ch`. Paste the whole list; one row of company data comes back per entry. Germany, Austria and Switzerland are all covered.

## `validate_vat` (type: `boolean`):

ON by default. Every VAT ID found is checked against the European Commission's VIES service, so you learn whether the number a company publishes is actually registered. Adds about a quarter of a second per company and costs nothing extra. Austria, Italy, Spain and most other member states also return the registered name and address, which lets you cross-check the imprint against the register. Germany discloses validity only. Swiss UIDs are outside the EU system and are never checked.

## `allow_contact_page` (type: `boolean`):

ON by default. If a site publishes no imprint, the contact page is read instead and the row is marked `foundVia: contact page` with `impressumFound: false` — you still get the address and the mail box, and you can see exactly where the data came from.

## `include_person_names` (type: `boolean`):

OFF by default. Person names in a legal notice are personal data under GDPR — switch this on only if you have a lawful basis for processing them. Company data, address, register and VAT number are returned either way.

## `max_items` (type: `integer`):

Optional cap. Leave empty to process the whole list.

## `osm_filter` (type: `string`):

For trades the search does not know. Pass a raw OSM tag filter such as `["amenity"="dentist"]` or `["office"="company"]` together with a location below. See taginfo.openstreetmap.org for the vocabulary.

## `location` (type: `string`):

City, district, county or state — e.g. `Hamburg`, `Landkreis Rosenheim`, `Bayern`. Needed only with the tag filter above.

## Actor input object example

```json
{
  "max_companies": 25,
  "websites": [
    "liebherr.com",
    "frischpack.de",
    "tuwien.at"
  ],
  "validate_vat": true,
  "allow_contact_page": true,
  "include_person_names": false
}
```

# Actor output Schema

## `companies` (type: `string`):

No description

## `companiesCsv` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "liebherr.com",
        "frischpack.de",
        "tuwien.at"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("perceptr0n/imprint-impressum-website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "liebherr.com",
        "frischpack.de",
        "tuwien.at",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("perceptr0n/imprint-impressum-website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "liebherr.com",
    "frischpack.de",
    "tuwien.at"
  ]
}' |
apify call perceptr0n/imprint-impressum-website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,perceptr0n/imprint-impressum-website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/YoHobeQht8StZit0A/builds/I1rZImQ3WPSch1ryH/openapi.json
