# Company Enrichment by Domain (Official Registers) (`truswen/company-enrichment-by-domain`) Actor

Turn domains into company profiles verified in official registers: legal name, company number, status, founding year, industry, headcount band, revenue, officers, VAT check, LEI and parent company, plus email provider, DMARC and tech stack. Every fact carries its source.

- **URL**: https://apify.com/truswen/company-enrichment-by-domain.md
- **Developed by:** [Benjamin Zsigri](https://apify.com/truswen) (community)
- **Categories:** Lead generation, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $25.00 / 1,000 company enricheds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Company Enrichment by Domain (Official Registers)

Give it a list of domains. Get back one company profile per domain, built from the company's own website and checked against **official business registers**: legal name, company number, legal form, register status, founding date, industry code, headcount band, revenue, officers, VAT validity, LEI and parent company. On top of that come the domain's email provider and DMARC policy from DNS, and the technologies the website runs.

Every value says where it came from (`legal_name_source`, `address_source`, `registry_matched_by`, `data_sources`). When the evidence is not strong enough to name the legal entity, the row says so instead of guessing.

### What makes the rows trustworthy

- **The website tells us who the company is, the register tells us the facts.** The crawler reads the legal notice first (Impressum, mentions légales, company information, site info) and takes the company number, VAT number and legal name the company publishes about itself. That evidence is looked up in the national register.
- **No guessed numbers.** Headcount and revenue only appear when a register publishes them (for example the INSEE headcount band and the filed accounts in France, the statistical register in the Czech Republic). Nothing is estimated from LinkedIn or from the size of the website.
- **A name alone is never enough.** A register hit found by name is accepted only when the register address (postcode and town) also appears on the company's website and exactly one candidate fits. Otherwise the row stays unmatched.
- **Confidence on every match.** `registry_match_confidence` is `high` when an identifier from the company's own website led to the record and the record agrees with the site, and `medium` when name and address agreed.

### Coverage

| Country | Source | Found by | Officers | Size and financials |
|---|---|---|---|---|
| France | Recherche d'entreprises (INSEE Sirene, RNE) | SIREN, SIRET or VAT on the website; or name and address | Yes (natural persons) | INSEE headcount band; revenue and net income when filed publicly |
| United Kingdom | Companies House (public service pages, or the REST API with your key) | Company number on the website; or name and address | Active directors and secretaries | Not published |
| Czech Republic | ARES (core record, statistical register, commercial register) | IČO or DIČ on the website; or name and address | Current statutory body | CZSO headcount band |
| Finland | YTJ open data (PRH) | Y-tunnus or VAT on the website; or name and address | Not in the open data | Not in the open data |
| Poland | KRS extract and the Ministry of Finance VAT list | KRS, NIP or REGON on the website | Not available (masked in the extract) | Not available |
| Switzerland | UID register (Federal Statistical Office) | UID (CHE number) on the website; or name and address | Not available | Not available |
| Norway | Enhetsregisteret (Brønnøysund) | Organisation number on the website; or name and address | CEO and board | Exact employee count |
| Germany, Netherlands, Belgium, Austria and others | GLEIF LEI index | Commercial register number (HRB, KvK, ...) or legal name with a matching address | Not available | Not available |
| All EU countries | VIES VAT check | VAT number on the website | | |

LEI records (with direct and ultimate parent) and VIES checks work for every country. Companies outside these registers still get the website, DNS and technology part of the profile.

### Output

One row per domain. The overview columns:

| Column | Example |
|---|---|
| `domain` | `lumiere-conseil.fr` |
| `legal_name`, `legal_form` | `LUMIERE CONSEIL`, `SAS` |
| `registry_verified` | `true` |
| `registry_id`, `registry_id_type`, `registry_status` | `912345675`, `FR_SIREN`, `active` |
| `registry_match_confidence`, `registry_matched_by` | `high`, `company_number_on_website` |
| `country`, `city`, `postal_code`, `street` | `FR`, `PARIS`, `75002`, `12 RUE DE LA PAIX` |
| `founded_year` | `2022` |
| `industry`, `industry_code`, `industry_section` | `Activities of head offices; management consultancy activities`, `70.22Z`, `Professional, scientific and technical activities` |
| `employees_range`, `employees_count`, `employees_source` | `10-19`, `null`, `INSEE tranche d'effectif` |
| `revenue`, `net_income`, `financial_year`, `currency` | `1250000`, `85000`, `2024`, `EUR` |
| `vat_id`, `vat_valid` | `FR65912345675`, `true` |
| `lei`, `parent_company`, `ultimate_parent_company` | from GLEIF when the company has an LEI |
| `key_person_name`, `key_person_role` | `Claire Martin`, `Président` |
| `email_provider`, `email_security_gateway`, `dmarc_policy` | `Google Workspace`, `null`, `reject` |
| `cms`, `ecommerce_platform`, `analytics`, `marketing_tools` | `WordPress`, `null`, `Google Analytics`, `null` |
| `primary_email`, `primary_phone`, `linkedin` | from the website |

The full row adds the complete register record (`registry`), the LEI record (`lei_record`), the VIES answer (`vat_validation`), all key people with their source (`key_people`), the DNS footprint (`email_and_dns`), every detected technology with its evidence (`technologies`), contacts, the identifiers found on the website, a per-register lookup trace (`registry_lookup`) and the crawled pages. The example above is a fictional company.

### Input

- **Domains or websites**: `acme.com`, `https://www.acme.com/about`, either works. Duplicates are merged.
- **Dataset import**: point to any dataset with a website or domain field (CRM export, Google Maps Scraper). A company name or country in the item helps the register match, and the item's identifying fields are copied to `source_item`. Items without a website are skipped; a Google Maps listing link is never taken for the website.
- **Switches** for each source (registers, LEI, VIES, DNS, technologies, contacts), so you only wait for what you use.
- **Companies House API key** (optional): with your own free key, UK lookups use the official REST API.

### Where your list can come from

- **Google Maps Scraper**: set `datasetId` to its run's dataset. This Actor never searches Google Maps itself; it works on the dataset you bring.
- **A CSV file or spreadsheet**: paste the website or domain column into `websites`, one per line.
- **Your CRM**: export the accounts and paste the website or domain column, or push the rows into an Apify dataset with the API and set `datasetId` to it. `datasetUrlField` names the column when it is not one of the usual names.
- **Make, n8n, Zapier or Clay**: start the Actor through the Apify API (Apify also offers ready-made modules for some of these tools), pass `websites` in the input and read the run's dataset as the output.

### Pricing

Pay per event, shown on the Actor's pricing tab:

- `company-enriched`: one per domain whose website answered.
- `registry-match`: one per domain identified in an official register, in the LEI index, or by a VIES-confirmed VAT number whose registered name matches the website.

Domains whose website cannot be read (blocked by robots.txt, offline, not a company website) produce a free row with the reason. When your maximum cost per run is reached, the run stops cleanly; register data that could not be paid for is left out of the row and marked `registry_limited_by_cost_limit`.

### Access rules

- Every website and every register API is read under its own robots.txt. If a robots.txt cannot be read, that site or register is skipped for the run, not worked around. Login pages, CAPTCHAs and other access controls are never bypassed.
- Each register is queried below its published rate limit. A register that keeps failing is switched off for the rest of the run, so one provider's outage never stops the run.
- Emails hidden by Cloudflare protection are flagged, not decoded.

### Personal data

Officers come from public registers with **name and role only**. Dates of birth, private and correspondence addresses, personal ID numbers and bank account lists that some registers publish are never copied. Website contacts are limited to what the company publishes about itself. You are responsible for having a lawful basis (for example legitimate interest in B2B prospecting) and for honouring objections when you process the people in these rows.

### Known limits

- Websites that render their content only with JavaScript expose few links to the crawler; their register match depends on the legal notice being reachable as plain HTML.
- Some companies publish their legal notice on a separate domain (for example a `.legal` domain); such pages are not followed.
- Norway: at the time of writing, the register API's robots.txt could not be read from our crawler, so Norwegian register lookups are skipped while that lasts. LEI and website data still work.
- Germany has no free register API; German companies are verified through their LEI when the Impressum shows the HRB number, and through VIES, which confirms the VAT number but does not disclose the name.
- Poland's free VAT list allows few lookups per day; the Actor caps its use per run and prefers a KRS number found on the website.

### Typical uses

- Fill CRM company records (legal name, company number, industry, size, address) from a domain list.
- Qualify inbound leads and B2B sign-ups by register status, age, size and parent company.
- Segment accounts by email provider, security gateway, DMARC policy, CMS and ecommerce platform.
- Run KYB pre-checks: does the company exist, is it active, is its VAT number valid, who are its officers.

### Need it done for you?

Want this connected to your CRM or workflow, or adapted to your market? Tribloc, the team behind this Actor, builds these setups. Get in touch at [tribloc.co.uk](https://tribloc.co.uk/).

# Actor input Schema

## `websites` (type: `array`):

One per line. A bare domain (acme.com) works. Each domain is processed and charged once, even if it appears several times.

## `datasetId` (type: `string`):

ID of a dataset whose items contain a website or domain field, for example a CRM export or a Google Maps Scraper run. Fields like title, company\_name, country and id are copied into `source_item` so you can join the rows back; a company name or country in the item also helps the register match. Pick the dataset (or paste its ID); the Actor is given read access to that one dataset only.

## `datasetUrlField` (type: `string`):

Name of the field that holds the website or domain. Leave empty to try the usual names: website, url, websiteUrl, domain, companyWebsite, homepage.

## `registryLookup` (type: `boolean`):

Identify the company in the national register on the evidence its own website gives (company number, VAT number, legal name and address in the legal notice). Covered: France, United Kingdom, Czech Republic, Finland, Poland, Switzerland, Norway (when reachable). A match is a separately charged event.

## `includeLei` (type: `boolean`):

Find the company's Legal Entity Identifier and its direct and ultimate parent in the open GLEIF index. This also covers companies in countries without a free register API, such as Germany, when the website shows the commercial register number.

## `validateVat` (type: `boolean`):

Check the VAT number in the European Commission's VIES service and compare the registered name with the website.

## `checkDns` (type: `boolean`):

Read the domain's public MX, SPF and DMARC records: email provider, security gateway, DMARC policy and email services the domain authorises.

## `detectTechnologies` (type: `boolean`):

CMS, ecommerce platform, analytics, marketing and chat tools, hosting and frameworks, from the pages fetched anyway.

## `includeContacts` (type: `boolean`):

Published emails, phone numbers, social profiles and the decision makers named on the website.

## `companiesHouseApiKey` (type: `string`):

Without a key, UK companies are read from the public Companies House service pages. With your own free REST API key (developer.company-information.service.gov.uk) the official API is used instead.

## `defaultCountry` (type: `string`):

Two-letter code (FR, GB, DE, ...) used when neither the dataset item, the website's identifiers nor its country domain tell the country. Also used to read local phone numbers.

## `maxPagesPerWebsite` (type: `integer`):

Homepage included. The crawler reads the legal notice first, then contact, about and team pages. 5 is enough for most companies.

## `maxConcurrency` (type: `integer`):

How many domains are processed at the same time. Each website and each register is always requested at its own polite rate.

## `proxyConfiguration` (type: `object`):

Optional, for the company websites only (registers are always read directly). Use it only if many websites return blocked or HTTP errors.

## Actor input object example

```json
{
  "websites": [
    "alan.com",
    "kiwi.com",
    "octopus.energy",
    "tietoevry.com"
  ],
  "registryLookup": true,
  "includeLei": true,
  "validateVat": true,
  "checkDns": true,
  "detectTechnologies": true,
  "includeContacts": true,
  "maxPagesPerWebsite": 5,
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

One row per company: legal name, register match, country, industry, size, revenue, key person, email provider and CMS.

## `results` (type: `string`):

Complete rows: register record, LEI record, VAT check, key people, DNS footprint, technologies and contacts, each with its source.

## `summary` (type: `string`):

Domains by status, register matches, per-register request counts and charged events.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "alan.com",
        "kiwi.com",
        "octopus.energy",
        "tietoevry.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("truswen/company-enrichment-by-domain").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "alan.com",
        "kiwi.com",
        "octopus.energy",
        "tietoevry.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("truswen/company-enrichment-by-domain").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "alan.com",
    "kiwi.com",
    "octopus.energy",
    "tietoevry.com"
  ]
}' |
apify call truswen/company-enrichment-by-domain --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,truswen/company-enrichment-by-domain"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7FYFKZBKAIcgFhGeK/builds/zFIRkEiQ776Y3FQPR/openapi.json
