# Impressum Scraper – German, Austrian & Swiss Company Data (`gazidev/imprint-scraper`) Actor

Extract company data from the Impressum (imprint / legal notice) of any DE, AT or CH website: company name, legal form, managing directors, HRB/HRA/FN register number and court, USt-IdNr/UID VAT ID, address, phone, fax and email. No Google search. Pay only when an imprint is found.

- **URL**: https://apify.com/gazidev/imprint-scraper.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 imprint founds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Impressum Scraper – German, Austrian & Swiss Company Data

Give it a list of websites from **Germany, Austria or Switzerland**. For each one, the Actor finds the **Impressum** (imprint / legal notice / Offenlegung) and returns the company data in one clean row: **company name, legal form, managing directors (Geschäftsführer / Vorstand / Inhaber), register court and number (HRB, HRA, FN, CHE-UID), VAT ID (USt-IdNr., ATU…, CHE-… MWST), street, postcode, city, country, phone, fax, email and the person responsible for content**. Each row also has a **confidence score** and the **raw imprint text** the fields were read from.

- **Your domains, no Google search.** You bring the list (a CRM export, trade-fair exhibitors, a Google Maps scrape, a directory) and the Actor goes straight to each site. Nothing depends on search-engine scraping, so runs are fast, cheap and predictable.
- **Finds the imprint page itself.** It reads the homepage links named *Impressum, Imprint, Legal notice, Offenlegung, Anbieterkennzeichnung, Rechtliche Hinweise* and similar (also in footers). If there is none, it tries common paths such as `/impressum` and `/de/impressum`. It follows "Legal" overview pages to the actual imprint and also handles one-page sites with the imprint on the homepage.
- **Built for DACH legal forms and registers.** GmbH, UG (haftungsbeschränkt), AG, SE, KG, OHG, GbR, e.K., e.V., eG, PartG mbB, GmbH & Co. KG, SE & Co. KG, Austrian GmbH/Ges.m.b.H./OG/KG, Swiss AG/SA/GmbH/Sàrl. Register entries from Amtsgericht (HRB/HRA/VR/GnR/PR), Austrian Firmenbuch (FN 123456 a, Landesgericht/Handelsgericht) and Swiss UID (CHE-123.456.789) and Handelsregisteramt. For a GmbH & Co. KG, it returns the company's own HRA, not the general partner's HRB.
- **Honest output.** Every row has `confidence` (0–1), `confidenceLevel` and `fieldsFound`, plus `rawSnippet` (the first 1,500 characters of the imprint block) so you can verify the data or post-process it with an LLM.
- **Pay only for hits.** $1.50 per 1,000 websites where an imprint with a company name or a register/VAT number was found. Sites without an imprint, dead domains and blocked sites are free.
- **Polite.** HTTP only (no browser), **robots.txt respected** (including Crawl-delay), a configurable delay between requests to the same host (1 s by default), sequential requests per site, and an honest User-Agent with a contact address.

**Related Actors:** [Website Contact Finder](https://apify.com/gazidev/website-contact-finder) for emails, phones and social profiles of any website worldwide, and [Domain Checker – WHOIS/RDAP, DNS, SSL](https://apify.com/gazidev/domain-intel) to check the domains themselves.

### What can I use it for?

- **B2B lead enrichment in DACH:** turn a list of company domains into legal name, managing director, address and register number for your CRM
- **KYC / KYB and supplier onboarding:** check a supplier's legal entity, register number and VAT ID before you contract with them, then confirm them in the official register or VIES
- **Data cleaning:** fix company names and legal forms in your CRM (e.g. "Muster" → "Muster Maschinenbau GmbH & Co. KG")
- **Market research:** map which legal forms, cities and registers the companies in a niche use
- **Compliance monitoring:** check that your own shops or your franchise partners publish a complete Impressum (register, VAT ID, responsible person)
- **AI agents:** give an agent a "who operates this German website?" tool (see below)

### Input

| Field | Description |
|---|---|
| `urls` | Domains or URLs (`muster.de`, `https://www.firma.at/`), or the imprint URL directly |
| `bulkText` | Paste a big list or a CSV export (the first domain-like cell per row is used) |
| `sourceFileUrl` | Public TXT/CSV file with domains (e.g. a Google Sheets "publish to web" CSV link) |
| `onlyFound` | Leave out rows for sites where no imprint data was found (default: off) |
| `includeRawSnippet` | Add the raw imprint text (default: on) |
| `maxSites`, `maxPagesPerSite` (default 6), `respectRobotsTxt` (default on), `minDelayPerHostSecs` (default 1), `maxConcurrency` (default 10), `requestTimeoutSecs`, `proxyConfiguration` | Advanced |

```json
{
  "urls": ["heise.de", "zotter.at", "kaffeemacher.ch"],
  "onlyFound": false,
  "includeRawSnippet": true
}
```

### Output

One row per website. The Output tab has three tables: **Company data**, **Contacts & people** and **Register & tax IDs**. Here is a real row (`rawSnippet` shortened):

```json
{
  "domain": "heise.de",
  "website": "https://www.heise.de/",
  "imprintUrl": "https://www.heise.de/impressum.html",
  "imprintFound": true,
  "companyName": "Heise Medien GmbH & Co. KG",
  "legalForm": "GmbH & Co. KG",
  "managingDirectors": ["Ansgar Heise", "Beate Gerold"],
  "representatives": [{"name": "Ansgar Heise", "role": "Geschäftsführer"}, {"name": "Beate Gerold", "role": "Geschäftsführer"}],
  "registerCourt": "Amtsgericht Hannover",
  "registerType": "HRA",
  "registerNumber": "26709",
  "registerId": "HRA 26709",
  "vatId": "DE813501887",
  "street": "Karl-Wiechert-Allee 10",
  "postalCode": "30625",
  "city": "Hannover",
  "country": "DE",
  "phone": "+49 511 53520",
  "fax": "+49 511 5352129",
  "email": "webmaster@heise.de",
  "responsiblePerson": "Dr. Volker Zota",
  "confidence": 1.0,
  "confidenceLevel": "high",
  "rawSnippet": "Impressum\nVerantwortlich für dieses Angebot:\nHeise Medien GmbH & Co. KG\nKarl-Wiechert-Allee 10\n30625 Hannover…",
  "error": null
}
```

An Austrian row has `registerType: "FN"`, for example `"registerId": "FN 220619s"` with `"registerCourt": "Landesgericht für ZRS Graz"` and `"vatId": "ATU53816900"`. A Swiss row has `"registerId": "CHE-256.970.360"`, `"swissUid"` and `"vatId": "CHE-256.970.360 MWST"`. Other fields: `taxNumber` (Steuernummer), `emails` (all emails in the imprint), `companyNameSource`, `imprintSource`, `pagesFetched`, `robotsTxt`, `blocked`, `durationMs`. See [SAMPLE_OUTPUT.json](SAMPLE_OUTPUT.json) for 13 real rows, including a sole trader, an e.V., a robots.txt-disallowed shop and a blocked site.

**Accuracy (local test, October 2026).** We ran 52 real DE/AT/CH websites, from sole traders and blogs to Kärcher, Tchibo, BILLA and Ricola. 44 returned imprint data. Of the other 8, 6 were bot-protected (HTTP 403/406 or challenge pages), 1 Shopify shop disallows its legal-notice page in robots.txt and 1 site was too slow. We checked the fields by eye against the live imprint for 27 sites:

- Company name: 27/27 correct
- Register number: 23/23 correct
- VAT ID: 23/23 correct
- Managing directors: 20/21 complete (one named only in a parenthesis was missed)
- Address: 26/27 complete (one street came back only partly)
- Phone, fax and email: all correct where present

Swiss imprints often omit the UID. In that case you get name, address and contacts at `confidence` 0.5.

### Pricing

Pay per event:

| Event | Price |
|---|---|
| Imprint found (company name or register/VAT number extracted) | **$0.0015** ($1.50 per 1,000) |
| Actor start | $0.00005 |
| No imprint, dead domain, blocked site, robots-disallowed page | free |

How that compares with other imprint scrapers in the Apify Store (listed prices, October 2026):

| Actor | Listed price | Notes |
|---|---|---|
| winningsolutions/german-imprint-scraper | $5 / 1,000 | finds sites through Google search |
| dominic-quaiser imprint scraper | $1.20 / 1,000 | |
| **This Actor** | **$1.50 / 1,000 hits; misses are free** | your own domain list, DE + AT + CH registers, confidence + raw text |

Set a **maximum cost per run** in the run options. The Actor saves as many hits as fit and then stops cleanly.

### GDPR and responsible use

An Impressum is published because the law requires it (§ 5 DDG in Germany, § 5 ECG and § 25 MedienG in Austria, Art. 3 UWG in Switzerland). It is meant to let people identify and contact the business. It still often contains **personal data**, such as managing directors' names, sole traders' home addresses and personal emails.

- This Actor is intended for **B2B use**: identifying the company behind a website, KYB/supplier checks, and enriching business records.
- **You are the controller** of the data you collect and are responsible for a lawful basis (usually legitimate interest, Art. 6(1)(f) GDPR), data minimisation, retention limits, and informing data subjects where required (Art. 14 GDPR).
- An Impressum is **not consent to marketing**. Cold emails or calls to businesses are restricted in Germany (§ 7 UWG) and Austria (§ 174 TKG). Check the rules for your channel before you use the contacts for outreach.
- Use `includeRawSnippet: false` and drop the fields you do not need. Verify register numbers in the official registers (Handelsregister, Firmenbuch, Zefix) and VAT IDs in VIES before you rely on them.
- The Actor reads only publicly accessible pages, respects robots.txt, and does not log in or bypass bot protection.

### FAQ

**How does it find the imprint?**
It reads the homepage links whose text or URL says Impressum, Imprint, Legal notice, Offenlegung, Rechtliche Hinweise and so on. If none is found, it tries `/impressum`, `/de/impressum`, `/imprint`, `/kontakt` and similar paths. If the first page has no company data (for example a "Legal" overview page), it follows the imprint links on that page. It fetches at most `maxPagesPerSite` pages per site (default 6, usually 2 are enough).

**Why is a site marked "Blocked by bot protection"?**
Some large shops (Cloudflare, Akamai and similar) block every non-browser client. This Actor does not use browsers or try to bypass protection. Those rows are free.

**Why is the imprint of a Shopify shop "disallowed by robots.txt"?**
Shopify's default robots.txt disallows `/policies/`, which is where Shopify shops keep their legal notice. With `respectRobotsTxt` on (the default and our recommendation) the page is not fetched, and you pay nothing for that row.

**What is `confidence`?**
A 0–1 score built from what was found: company name with a legal form, register number and court, VAT ID, a full address, managing directors and contact data. 0.7 or higher is `high`. Sole traders without a register usually score 0.3–0.5. The score is higher there when name and address are present.

**Does it work for other countries?**
It is tuned for the German-speaking imprint format (DE, AT, CH, LI). English-language legal notices of DACH companies ("Registration court", "VAT No.", "Managing directors") work too. For other countries use [Website Contact Finder](https://apify.com/gazidev/website-contact-finder).

**Can I get CSV or Excel?**
Yes. Every Apify dataset can be exported to CSV, Excel, JSON or XML, or sent to Google Sheets with an integration. Array fields such as `managingDirectors` become joined columns.

### Use with AI agents / Apify MCP

The input is just a domain and the output is structured company data at $0.0015 per hit, which makes this Actor a good agent tool. Connect it through the [Apify MCP server](https://mcp.apify.com) (`https://mcp.apify.com?actors=gazidev/imprint-scraper`) and Claude, ChatGPT, Cursor and other MCP clients can answer prompts such as *"Who is the managing director of the company behind muster-shop.de, and what is its HRB number?"* or *"Enrich these 200 Austrian domains with FN number and UID"*. Over the API: `POST https://api.apify.com/v2/acts/gazidev~imprint-scraper/run-sync-get-dataset-items` with the input JSON.

### Website toolkit by gazidev

Other low-cost, HTTP-only Actors for working with lists of websites:

- [Website Contact Finder](https://apify.com/gazidev/website-contact-finder): emails, phones and social profiles of any website
- [Sitemap Extractor](https://apify.com/gazidev/sitemap-url-extractor): every URL from sitemap.xml, with an optional status check
- [Wappalyzer Alternative – Tech Stack Detector](https://apify.com/gazidev/tech-stack-detector): the CMS, shop system, analytics and frameworks a site uses
- [Domain Checker – WHOIS/RDAP, DNS, SSL](https://apify.com/gazidev/domain-intel): registration, DNS, SSL and email security
- [Website to Markdown](https://apify.com/gazidev/website-to-markdown): clean Markdown of any page or site for LLMs and RAG
- [SEO Audit Crawler](https://apify.com/gazidev/seo-audit-crawler): broken links, meta tags and redirects
- [Phone Number Validator](https://apify.com/gazidev/phone-validator): validate and format the phone numbers you collect

### Categories

Lead generation · Business · Automation

# Actor input Schema

## `urls` (type: `array`):

Company websites from Germany, Austria or Switzerland, e.g. `muster-gmbh.de` or `https://www.firma.at/`. `https://` is added automatically. You can also give the imprint URL directly. Each domain is processed once and returns one row.

## `bulkText` (type: `string`):

Paste a big list: one domain per line, comma/semicolon separated, or a whole CSV export (the first cell that looks like a domain/URL in each row is used).

## `sourceFileUrl` (type: `string`):

Public URL of a .txt or .csv file with domains/URLs (e.g. a Google Sheets 'publish to web' CSV link).

## `onlyFound` (type: `boolean`):

Leave out rows for websites where no usable imprint was found (unreachable, blocked, no Impressum). Those rows are free either way.

## `includeRawSnippet` (type: `boolean`):

Add `rawSnippet`: the first 1,500 characters of the imprint block the fields were read from, so you can verify or post-process them (e.g. with an LLM).

## `maxSites` (type: `integer`):

Stop after this many websites (0 = no limit). Useful to test a big list cheaply.

## `maxPagesPerSite` (type: `integer`):

Homepage + imprint candidates (links named Impressum/Imprint/Legal notice, then common paths like `/impressum`). 6 is enough for almost every site. The price per website is the same whatever you choose.

## `respectRobotsTxt` (type: `boolean`):

Skip pages disallowed by the site's robots.txt and honor its Crawl-delay. Recommended. Note: Shopify shops disallow `/policies/` by default, so their legal notice is then reported as disallowed.

## `minDelayPerHostSecs` (type: `integer`):

Politeness throttle: minimum gap between two requests to the same website (a larger robots.txt Crawl-delay wins, up to 10 s).

## `maxConcurrency` (type: `integer`):

How many different websites are processed at the same time. Requests to one website are always sequential.

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout. Failed requests are retried.

## `proxyConfiguration` (type: `object`):

Optional. Not needed for most sites.

## Actor input object example

```json
{
  "urls": [
    "heise.de",
    "zotter.at",
    "kaffeemacher.ch"
  ],
  "onlyFound": false,
  "includeRawSnippet": true,
  "maxSites": 0,
  "maxPagesPerSite": 6,
  "respectRobotsTxt": true,
  "minDelayPerHostSecs": 1,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `contacts` (type: `string`):

No description

## `register` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "heise.de",
        "zotter.at",
        "kaffeemacher.ch"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/imprint-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "heise.de",
        "zotter.at",
        "kaffeemacher.ch",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("gazidev/imprint-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "heise.de",
    "zotter.at",
    "kaffeemacher.ch"
  ]
}' |
apify call gazidev/imprint-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/imprint-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DtND2TJ0sA3QGhqdL/builds/zbta4JdrQ0jrRwd5y/openapi.json
