# North Data Scraper: German & European Companies (`enisbodlli/northdata-company-scraper`) Actor

Look up companies on North Data by name, register number or URL: register ID, legal form, status, address, LEI, purpose, capital and published financials, one row per company. Company data only: officers are counted, never returned. Plain requests, no proxy. You pay only for companies found.

- **URL**: https://apify.com/enisbodlli/northdata-company-scraper.md
- **Developed by:** [Enis Bodlli](https://apify.com/enisbodlli) (community)
- **Categories:** Business, Lead generation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 company records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## North Data Scraper: German & European Companies

This **North Data scraper** looks up German and European companies by name, register number or
company page URL and returns one clean row per company: register entry, legal form, status, address,
LEI, corporate purpose, share capital and published financials. It is built for compliance, KYB and
enrichment teams that need **company data without people's names**: officers and other people are
counted, never returned. To try it, keep the sample input and press **Start**: two companies come
back in well under a minute, with plain requests, no proxy, and a charge only for companies found.

- **Company-level data only.** No officer, director, signatory or shareholder name is ever in the
  output. `officersOmitted` tells you how many people the page names; that is all.
- **A stable row.** Every field is always present (`null` or `[]` when the page has nothing), with
  `sourceUrl` and `scrapedAt` on every row, so nothing breaks downstream when a company has less data.
- **Honest matching and billing.** A company counts as found only when every word of your entry is
  in its name, its town or an earlier name. An entry that finds nothing gets a row with
  `found: false` and is not charged.

### What you get

One row per company. Shortened here; the full examples are further down.

```json
{
    "query": "Siemens AG München",
    "found": true,
    "matchRank": 1,
    "name": "Siemens AG",
    "legalForm": "AG",
    "status": "active",
    "registerId": "Amtsgericht München HRB 6684",
    "address": { "street": "Werner-von-Siemens-Str. 1", "postalCode": "80333", "city": "München", "country": "DE" },
    "lei": "W38RGI023J3WT1HWRP32",
    "financials": [{ "year": 2025, "revenue": 78914000000, "earnings": 9620000000, "totalAssets": 166202000000, "employees": null, "currency": "EUR" }],
    "officersOmitted": 142,
    "sourceUrl": "https://www.northdata.com/Siemens%20AG,%20M%C3%BCnchen/HRB%206684"
}
```

### How to scrape North Data company profiles

1. Put one company per line into **Companies to look up**. Three kinds of entry are understood:
   - a company name, best with its town: `Siemens AG München`
   - a German register number with the town of its court: `HRB 6684 München`
   - the URL of a company page on northdata.com (the .de and .fr sites work too)
2. Leave **Companies per entry** at 1 to get the best match only, or raise it (up to 10) to get the
   next matches as well, each with its `matchRank`.
3. Start the run. Rows appear in the dataset as each company is read; the `RUN_SUMMARY` record in
   the key-value store says what happened to every entry.

Use the registered name. `Bayerische Motoren Werke` finds the car maker; `BMW AG München` does not,
because the registered name does not contain "BMW". A register number or a URL is the surest way to
name exactly one company.

### Pricing

You pay per company found, through one event called `company-record`:

| Apify plan | Per 1,000 companies | Per company |
|---|---|---|
| Free, Bronze | $3.50 | $0.0035 |
| Silver | $3.00 | $0.003 |
| Gold and above | $2.50 | $0.0025 |

Plus Apify's standard start fee of $0.00005 per run. Platform usage is included; there is nothing
else to pay.

- **200 companies found:** $0.70 on the Free and Bronze plans, $0.60 on Silver, $0.50 on Gold.
- **One lookup:** $0.0035 on the Free and Bronze plans.
- **An entry that finds nothing:** no `company-record` charge. Its row is stored free.

Financials and history are included in the price. A company is charged at the moment its row is
stored, and one company is charged once per run even when several entries lead to it.

**When a run stops early or fails** you pay for the rows that are in the dataset and nothing else.
A run that is restarted or resurrected continues where it stopped: entries that are finished are
skipped and a company that already has a row is not stored or charged again. After a crash at most
one lookup is repeated, and it is not charged twice. Lookups the source did not answer are asked
again when a run is resurrected.

You can set a maximum charge per run. The Actor starts a lookup only while that limit has room for
it, stops cleanly when it is reached and lists the entries it did not get to.

### Input

```json
{
    "queries": ["Siemens AG München", "HRB 6089 Ansbach"],
    "maxResultsPerQuery": 1,
    "includeFinancials": true,
    "includeHistory": false,
    "maxResults": 100,
    "requestDelaySeconds": 1.5
}
```

| Field | What it does | Default |
|---|---|---|
| `queries` | Company names, register numbers or North Data company URLs. 1 to 1,000 entries; repeated entries are looked up once. | required |
| `maxResultsPerQuery` | Matching companies returned per entry, best first. 1 to 10. | 1 |
| `includeFinancials` | Revenue, earnings and total assets per year, newest year first, where the public page shows them. | true |
| `includeHistory` | Dated events of the company as a date and an event type. | false |
| `maxResults` | The run stops after this many companies. 1 to 10,000. | 1,000 |
| `requestDelaySeconds` | Pause after each request. 1 to 30 seconds; it cannot go under 1. | 1.5 |

### Output

Three rows from runs on 2026-10-07. Long texts are cut with "…" here, the first row shows two of
its nine financial years, and the second row shows the first three of its 22 history events.

```json
[
    {
        "query": "Siemens AG München",
        "found": true,
        "matchRank": 1,
        "name": "Siemens AG",
        "legalForm": "AG",
        "status": "active",
        "registerCourt": "Amtsgericht München",
        "registerType": "HRB",
        "registerNumber": "6684",
        "registerId": "Amtsgericht München HRB 6684",
        "address": { "street": "Werner-von-Siemens-Str. 1", "postalCode": "80333", "city": "München", "country": "DE" },
        "lei": "W38RGI023J3WT1HWRP32",
        "vatId": null,
        "industry": { "codes": [], "text": "Manufacture of electronic components" },
        "purpose": "Development, manufacture, supply, operation and distribution of as well as trade in products, systems, plants and solutions …",
        "foundedOn": "1996-08-28",
        "capital": { "amount": 2350000000, "currency": "EUR", "date": "2026-03-20" },
        "financials": [
            { "year": 2025, "revenue": 78914000000, "earnings": 9620000000, "totalAssets": 166202000000, "employees": null, "currency": "EUR" },
            { "year": 2024, "revenue": 75930000000, "earnings": 8301000000, "totalAssets": null, "employees": null, "currency": "EUR" }
        ],
        "website": null,
        "phone": null,
        "email": null,
        "officersOmitted": 142,
        "history": [],
        "sourceUrl": "https://www.northdata.com/Siemens%20AG,%20M%C3%BCnchen/HRB%206684",
        "scrapedAt": "2026-10-07T21:37:02.009Z"
    },
    {
        "query": "https://www.northdata.com/Siemens AG Österreich, Wien/060562m",
        "found": true,
        "matchRank": 1,
        "name": "Siemens AG Österreich",
        "legalForm": null,
        "status": "active",
        "registerCourt": "Firmenbuch",
        "registerType": null,
        "registerNumber": "060562m",
        "registerId": "Firmenbuch 060562m",
        "address": { "street": "Siemensstraße 90", "postalCode": "1210", "city": "Wien", "country": "AT" },
        "lei": "52990021T5LVTQOGSU18",
        "vatId": null,
        "industry": { "codes": [], "text": "Manufacture of other general-purpose machinery n.e.c." },
        "purpose": "Foreign entity.",
        "foundedOn": "1993-11-12",
        "capital": { "amount": 125900000, "currency": "EUR", "date": "2000-08-04" },
        "financials": [],
        "website": null,
        "phone": null,
        "email": null,
        "officersOmitted": 80,
        "history": [
            { "date": "2025-12-31", "type": "patent" },
            { "date": "2025-07-02", "type": "trademark" },
            { "date": "2025-03-05", "type": "patent" }
        ],
        "sourceUrl": "https://www.northdata.com/Siemens%20AG%20%C3%96sterreich,%20Wien/060562m",
        "scrapedAt": "2026-10-07T21:45:20.481Z"
    },
    {
        "query": "BMW AG München",
        "found": false,
        "matchRank": null,
        "name": null,
        "legalForm": null,
        "status": null,
        "registerCourt": null,
        "registerType": null,
        "registerNumber": null,
        "registerId": null,
        "address": { "street": null, "postalCode": null, "city": null, "country": null },
        "lei": null,
        "vatId": null,
        "industry": { "codes": [], "text": null },
        "purpose": null,
        "foundedOn": null,
        "capital": null,
        "financials": [],
        "website": null,
        "phone": null,
        "email": null,
        "officersOmitted": 0,
        "history": [],
        "sourceUrl": "https://www.northdata.com/BMW%20AG%20M%C3%BCnchen",
        "scrapedAt": "2026-10-07T21:45:23.648Z"
    }
]
```

| Field | Meaning |
|---|---|
| `query` | The entry this row answers. |
| `found` | `true` for a company (charged), `false` when the entry found none (free). |
| `matchRank` | 1 for the best match of the entry, 2 for the next, and so on. |
| `name`, `legalForm`, `status` | Registered name; legal form read from the name (AG, GmbH, SE, Ltd and others); `active`, `in_liquidation` or `terminated`. |
| `registerCourt`, `registerType`, `registerNumber`, `registerId` | The register entry in parts, and the full citation. |
| `address` | `street`, `postalCode`, `city`, `country` (two-letter code). |
| `lei`, `vatId` | Legal Entity Identifier; VAT number when the page lists one. |
| `industry` | `text` is the industry the source files the company under; `codes` is a list of classification codes. |
| `purpose`, `foundedOn`, `capital` | Corporate purpose (up to 5,000 characters), founding date, and the latest share capital entry with its date. |
| `financials` | One object per financial year, **newest first**, so `financials[0]` is the latest published year. |
| `website`, `phone`, `email` | Company contact points when the page lists them; `email` only for a shared mailbox such as info@. |
| `officersOmitted` | How many people the page names for this company. Their names and roles are left out. |
| `history` | Events as `date` and `type` (registration, name_change, capital_change, liquidation, merger_or_acquisition, trademark and others), newest first. |
| `sourceUrl`, `scrapedAt` | The company's page on North Data (unique per company) and when it was read (UTC). |

The `RUN_SUMMARY` record holds the totals (`companiesFound`, `chargedEvents`, `soleTradersSkipped`,
`possibleSoleTradersSkipped`, `peopleRowsIgnored`, `requests`), why a run stopped early
(`stoppedBy`) and one outcome per entry:
`found`, `not_found`, `duplicate` (the company already has a row from another entry) or `failed`,
with the reason.

### Limits

**People are never returned.** This is the rule the Actor is built around, and it holds for every
field:

- Officers, directors, managing directors, board members and signatories are not returned. You get
  the count in `officersOmitted`.
- There is no people search and no person record. Search results about people are never opened. A
  person's URL in the input is refused, and if a page about a person is reached all the same,
  nothing from it is kept.
- Sole traders (German e.K., e.Kfm., e.Kfr., Austrian e.U. and businesses whose page names a
  proprietor) are legally a natural person and carry that person's name. They are skipped and
  counted in `soleTradersSkipped`.
- A business gets a row only when something on its page shows a company or another organisation:
  a legal form in its name (GmbH, AG, KG, B.V., SAS, Ltd and so on), a register that holds no
  natural persons (German HRB, GnR, VR, PR and GsR, UK Companies House), or a director, board
  member or partner among its representatives. The source's public page states no legal form, and
  a one-person business abroad can look like any other entry: on a Dutch page checked on
  2026-10-08 the owner's name was the business name and nothing said so. An entry that shows none
  of the three is skipped and counted in `possibleSoleTradersSkipped`. This also leaves out some
  real organisations, such as a foundation or association outside Germany whose name carries no
  legal form and whose page lists nobody.
- History events are a date and a type. The source's text for an event can name people and is not
  returned; appointments and departures of people are left out entirely.
- `email` is returned only for a shared mailbox (info@, kontakt@, office@ and similar).

**Not offered:**

- No corporate network and no related companies: that list names shareholders and shared officers.
- No risk level or risk grade. The source's terms prohibit producing credit assessments from its
  information, and a grade is an assessment, not register data.
- No proxy and no proxy option. Requests go out one at a time from one address, so large runs are
  slower than tools that run in parallel through proxies.
- No concurrency setting. The run is sequential by design; `requestDelaySeconds` sets the pace.
- No country filter in this version. Add the town or the register number to the entry to pin a company.
- No geo coordinates, stock symbols, news, balance-sheet lines, EU ID or ELF code in this version.
- No status, legal-form, keyword or size filters, no change monitoring, no notifications and no
  choice of output language. `status` and `legalForm` are fields you can filter on yourself.
- The source's paid API, its Power Search and everything behind its Premium login are not used.
- Names the source lists without a register entry (some foundations, associations and public
  bodies) are not returned; such an entry can be a private person in business.

**What the public page does not show**, so these stay empty:

- `employees` is `null`: the source shows employee numbers to its paying subscribers only.
- `industry.codes` is `[]`: the public page names the industry in words and shows no code.
- `vatId`, `website`, `phone` and `email` were `null` on every page checked on 2026-10-07. They are
  filled only if a page lists them.
- `financials` is `[]` for companies that publish no figures, and a year can lack single figures.
  Estimates of the source are left out; only published figures are returned.

**Accuracy.** The source compiles its data by automated reading of public announcements and says
itself that it can be wrong. `capital` is the amount of its latest capital entry and was clearly off
for one company checked. `legalForm` is read from the end of the name and is `null` for names that
end otherwise. `foundedOn` is often the date of the current register entry, not the year a company
was started.

**Speed and caps.** A name takes two requests (the search, then the company page); a URL, or a
register number with its court, takes one. Measured on Apify on 2026-10-08 at the default 512 MB: 12 names in 146 seconds and
30 names in about 8 minutes, so 4 to 5 companies a minute by name (about 6 at 1,024 MB). 1,000
names take about four hours: the default timeout is set to four hours for that reason, and a
longer list is best split into several runs. A run takes up to 1,000 entries and
10 companies per entry, and reads at most 3 result pages (45 results) per name. The Actor takes no
new entry in the last minute before the run's timeout and lists what it did not reach.

**When the source says no.** A 403, a CAPTCHA or another bot check is a final answer: the run stops,
keeps what it has stored, ends as failed and says so. Nothing is done to get around it. Three
lookups in a row without an answer stop the run as well.

### FAQ

#### Is it legal to scrape North Data?

The Actor reads public pages that any visitor sees without an account. It does not log in, uses no
proxy, solves no CAPTCHA and works around no blocking, and it returns company data only. As read on
2026-10-07, the site's usage terms say nothing about automated access; the terms of its paid
products forbid customers to read the database with scripts and state that they do not cover
visitors of the free website; robots.txt shuts out 20 named crawlers and sets no rule for others.
For bulk or contractual use the source offers its own paid interfaces, and this Actor is not a
replacement for them. You are responsible for how you use the data. This is not legal advice.

#### May I republish the data?

Name North Data as the source and link the company's page; every row carries it in `sourceUrl`.

#### Can I use it for credit scoring?

No. The source prohibits credit assessments based on its information, and this Actor returns no
risk grade.

#### Why was my company not found?

Every word of the entry must be in the company's name, its town or an earlier name, so a brand name,
an abbreviation or a typo gives `found: false` instead of a wrong company. Use the registered name
with the town, a register number such as `HRB 6684 München`, or the company's URL. A match is also
left out when it is a sole trader, or when nothing on its page shows a company (see Limits). The
reason is in `RUN_SUMMARY`.

#### Why did I get a different company with the same name?

Namesakes exist, and a company can have entries at two courts. Add the town or the court to the
entry; it decides which one is `matchRank` 1.

#### The run failed with "the source refused the requests". What now?

The source declined to answer from the address the run used. Wait, then resurrect the run to
continue it, or start a new one; the rows that were stored are kept and you paid only for those.

#### Where do I report a problem?

Open an Issue on the Actor's page and include the run ID. Wrong or missing fields are fixed fastest
with the entry that produced them.

### More Actors from this developer

- [Handelsregister Scraper](https://apify.com/enisbodlli/handelsregister-scraper): official German company register search
- [European Company Registry Search](https://apify.com/enisbodlli/eu-company-registry-search): official registers of eight countries
- [Brazil CNPJ Scraper](https://apify.com/enisbodlli/brazil-cnpj-company-search): company search and lookup
- [US Business Entity Search](https://apify.com/enisbodlli/us-business-registry-search): official state filings
- [Company Jobs Search](https://apify.com/enisbodlli/company-jobs-search)
- [ATS Job Postings](https://apify.com/enisbodlli/ats-job-postings)
- [Workday Jobs Scraper](https://apify.com/enisbodlli/workday-jobs-scraper)
- [Greenhouse Jobs Scraper](https://apify.com/enisbodlli/greenhouse-jobs-scraper)
- [Lever Jobs Scraper](https://apify.com/enisbodlli/lever-jobs-scraper)
- [Ashby Jobs Scraper](https://apify.com/enisbodlli/ashby-jobs-scraper)
- [Email Validator](https://apify.com/enisbodlli/email-validator)
- [Website Contact Scraper](https://apify.com/enisbodlli/website-contact-scraper)

# Changelog

This Actor's version history is a separate document: https://apify.com/enisbodlli/northdata-company-scraper/changelog.md

# Actor input Schema

## `queries` (type: `array`):

One entry per line. Three kinds are understood: a company name, best with its town ("Siemens AG München"); a German register number with the town of its court ("HRB 6684 München"); or the URL of a company page on northdata.com. A company counts as found only when every word of the entry is in its name, its town or an earlier name, so a misspelt name gives "not found" instead of a wrong company. Up to 1,000 entries per run; repeated entries are looked up once.

## `maxResultsPerQuery` (type: `integer`):

How many matching companies to return for one entry, best match first. 1 returns the best match only. A register number with its court and a URL always name one company. Each company returned is charged once.

## `includeFinancials` (type: `boolean`):

Return revenue, earnings and total assets per year, newest year first, where the public North Data page shows them. Employee numbers are a paid feature of the source and stay empty. When switched off, "financials" is an empty list. The price is the same either way.

## `includeHistory` (type: `boolean`):

Return the dated events of the company's history, newest first, as a date and an event type (registration, name change, capital change, merger and so on). The source's text for an event is not returned, because it can name people. When switched off, "history" is an empty list.

## `maxResults` (type: `integer`):

The run stops once this many companies have been returned, which also caps what the run can cost. Entries that were not reached are listed in the run summary.

## `requestDelaySeconds` (type: `number`):

Seconds to wait after each request to North Data. Requests are sent one at a time, never in parallel. The default of 1.5 gives about 12 to 25 companies a minute; a longer pause is slower and gentler on the source. It cannot be set under 1 second.

## Actor input object example

```json
{
  "queries": [
    "Siemens AG München",
    "HRB 6089 Ansbach"
  ],
  "maxResultsPerQuery": 1,
  "includeFinancials": true,
  "includeHistory": true,
  "maxResults": 100,
  "requestDelaySeconds": 1.5
}
```

# Actor output Schema

## `companies` (type: `string`):

One row per company found, in the run's default dataset: register entry, legal form, status, address, LEI, purpose, capital, financials and the count of people left out. An entry that found nothing has a row with found = false.

## `runSummary` (type: `string`):

Totals for the run and the outcome of every entry: companies found (charged), entries not found, sole traders skipped, and what was left out and why.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "Siemens AG München",
        "HRB 6089 Ansbach"
    ],
    "maxResultsPerQuery": 1,
    "maxResults": 100,
    "requestDelaySeconds": 1.5
};

// Run the Actor and wait for it to finish
const run = await client.actor("enisbodlli/northdata-company-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": [
        "Siemens AG München",
        "HRB 6089 Ansbach",
    ],
    "maxResultsPerQuery": 1,
    "maxResults": 100,
    "requestDelaySeconds": 1.5,
}

# Run the Actor and wait for it to finish
run = client.actor("enisbodlli/northdata-company-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "Siemens AG München",
    "HRB 6089 Ansbach"
  ],
  "maxResultsPerQuery": 1,
  "maxResults": 100,
  "requestDelaySeconds": 1.5
}' |
apify call enisbodlli/northdata-company-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,enisbodlli/northdata-company-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QksNBLRBr2JSWOosb/builds/a788lgGdaBAxMstGw/openapi.json
