# Dutch KvK & BTW Website Scraper — Company Data & Contacts (`scrapersdelight/nl-kvk-website-contact-scraper`) Actor

Turn Dutch company domains into verified B2B/KYB leads from each site's statutory art. 3:15d disclosure: legal name, KvK-nummer, btw-identificatienummer validated with both Dutch check digits, address, e-mail and phone — plus the page each field came from. $0.009 per delivered lead.

- **URL**: https://apify.com/scrapersdelight/nl-kvk-website-contact-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$9.00 / 1,000 per dutch company lead returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 🇳🇱 Dutch KvK & BTW Website Scraper — company data from the art. 3:15d disclosure

Paste a list of Dutch company domains (or point this Actor at another Actor's dataset) and get back **one clean B2B/KYB record per company**, read off the statutory disclosure each Dutch website is obliged to publish: the legal name and rechtsvorm, the **KvK-nummer**, the **btw-identificatienummer checked against BOTH Dutch check-digit algorithms**, the postal address, an e-mail, a phone number and the IBAN — for **$9 per 1,000 leads**.

**Why a website and not the Handelsregister?** Because **Burgerlijk Wetboek art. 3:15d lid 1** obliges every provider of an information-society service in the Netherlands to make easily, directly and permanently accessible its identity and address (sub a), contact details **including its e-mail address** (sub b), the trade register it is entered in **together with its registration number** (sub c), and, where it is liable for VAT, its **btw-identificatienummer** (sub f). The corpus is therefore *the whole Dutch commercial web* — not one directory's member list — and every field is a public disclosure the company published about itself, by law. No login, no API key, no paid KvK key.

***

### ✅ What does it do?

For each domain: plain HTTP hops through the Apify datacenter proxy. **No browser is ever launched.**

1. **Fetch the homepage** — for its links as much as for its fields.
2. **Rank and follow the statutory links** — `/contact`, `/algemene-voorwaarden`, `/privacy`, `/over-ons`, `/disclaimer`, `/klantenservice` — matching on both the `href` and the anchor text, with assets filtered out. If nothing is linked, fall through to a **10-path URL ladder**.
3. **Read every page that loads, score them, and MERGE across them** — then report, per field, **which page it came from**.

#### The one thing that makes the Dutch corpus different

**Dutch law names no disclosure page.** Germany has the *Impressum*, France has the *mentions légales*, Spain has the *aviso legal* — one page, one name, one link. Art. 3:15d says only that the information must be "easily, directly and permanently accessible", so a Dutch MKB site **scatters the block**: the phone on `/contact`, the KvK-nummer in the small print of `/algemene-voorwaarden`, the BTW number in the footer of `/privacy`.

Measured on the 300-domain benchmark below: **of the 89 KvK numbers found, 8 were on the homepage. The other 81 — 91.0 % — required a second page.** An Actor that reads the homepage and stops returns a tenth of this lane's register ids. **33 of the 180 delivered rows (18.3 %) were assembled from two or three different pages**, and every one of them says so in `fieldSources`.

#### The four things a naive Dutch scraper gets wrong

- 🧮 **A VAT validator that only knows the elfproef rejects every sole trader registered since 2020.** The classic Dutch BTW number is derived from the RSIN and carries the 11-test. The **btw-identificatienummer** issued to *eenmanszaken* since 1 January 2020 is deliberately **random**, is not derivable from the BSN, and carries a **mod-97** check instead. Measured on the benchmark: **35 of the 49 valid numbers passed the elfproef and 14 passed mod-97** — a validator with only the first would have marked 29 % of real, live VAT numbers invalid. Both are implemented, `btwCheckMethod` names which one answered.
- 🔢 **A bare 8-digit token is not a KvK number.** The same corpus prints AGB healthcare codes, order numbers and dates in exactly that shape. Every KvK number here is **label-anchored** — it must be the first number after `KvK` / `Kamer van Koophandel` / `Handelsregister`, with no other field's label in between. That also survives the real spellings: `KvK onder nummer 34271453`, `Kamer van Koophandel onder registratienummer 34115621`, `KvK | Oost-Nederland nummer 08211558`, and a transposed contact table where the number sits three lines below its own label.
- 🏦 **An IBAN is a phone-number factory.** `IBAN NL21 ABNA 0848 0406 27` contains `0848040627` — ten digits starting with a zero, a perfectly well-formed Dutch number. That shipped as a company's phone number once. Every IBAN-shaped token is now blanked before any phone is read, whether or not it passes mod-97.
- 🕳 **A 200 carrying a block page is a transport failure, not an empty result.** A bot-walled Dutch retailer answering with a challenge interstitial must never read as *"this company publishes no KvK number"*. Block bodies are sniffed, retried on a fresh IP, and reported as `blocked` — never as an absent disclosure, and never charged.

***

### 📊 Measured, from real runs

Everything below is from actual runs on the live Actor, not estimates.

#### End-to-end over 300 real Dutch MKB domains

Sourced from Telefoonboek.nl (the Places.nl network) across **50 trades × 16 cities** — loodgieters, elektriciens, aannemers, tandartsen, makelaars, accountants, bakkerijen, autobedrijven and so on, in Amsterdam, Rotterdam, Den Haag, Utrecht, Eindhoven, Groningen, Tilburg, Breda, Nijmegen, Apeldoorn, Zwolle, Maastricht, Arnhem, Haarlem, Enschede and Leiden. A realistic buyer list of small Dutch companies, **not** a hand-picked one, and deliberately not the bot-walled `bol.com` / `wehkamp` class of retailer. Apify **datacenter** proxy, concurrency 8, otherwise default settings, 512 MB.

| Outcome | Domains | Share | Billed? |
|---|---:|---:|:--:|
| ✅ `ok` — a parsed, billable lead | **180** | **60.0 %** | **yes** |
| 🟡 `partial` — below your minimum-fields bar | 13 | 4.3 % | no |
| ⚪ `no_disclosure_found` — reached, publishes nothing readable | 23 | 7.7 % | no |
| 💀 `unreachable` — dead host, or a homepage serving nothing | 72 | 24.0 % | no |
| 🔴 `blocked` by an anti-bot wall | 12 | 4.0 % | no |

**216 of the 300 domains were reachable, and 180 of those 216 (83.3 %) became billable leads.**
Run-to-run variance on the same list is real: three runs of this identical 300 returned **181, 182 and 180** leads, with the unreachable count moving between 59 and 74. Stale directory listings come and go; expect a few points either way.

#### Per-field fill, over the 180 delivered leads

| Field | Fill | | Field | Fill |
|---|---:|---|---|---:|
| `companyName` | **91.1 %** | | `socialProfiles` | 70.0 % |
| `email` | **95.6 %** | | `mobile` | 39.4 % |
| `phone` | **93.3 %** | | `legalName` | 36.1 % |
| `postalCode` | **90.6 %** | | `legalForm` | 32.8 % |
| `city` | **90.6 %** | | `btwNumber` | 27.8 % |
| `street` | **81.7 %** | | `iban` | 8.9 % |
| `houseNumber` | 77.8 % | | `postbus` | 5.6 % |
| `kvkNumber` | 49.4 % | | `privacyEmail` | 3.3 % |

**Any statutory id (KvK or BTW): 53.9 % · both: 23.3 % · any contact (e-mail or phone): 97.8 % · full postal address: 90.0 %.**
Median row carries **10** of the 21 value fields (p90 = 13, best = 15).

#### Is 49 % a parser gap, or is it the corpus?

A fill rate on its own cannot tell those two apart, so here is the measurement that can. Over the
457 captured pages, for every page where a **register label actually appears in the raw bytes**
(`KvK`, `Kamer van Koophandel`, `Handelsregister`), did a number come out?

| | label present | value emitted | |
|---|---:|---:|---:|
| **KvK**, per domain | 85 domains | **83** | **97.6 %** |
| **BTW**, per domain | 70 domains | 45 | 64.3 % |

**Where a Dutch site names its register, this Actor reads the number 97.6 % of the time.** The 49 %
fill is the corpus, not the parser.

The BTW figure looks worse and is not: `btw` is an ordinary Dutch word. All 25 gap domains were read
by hand — **23 of them mention VAT in running prose and publish no number at all**
("prijzen zijn inclusief btw", "Vrij van btw"), or are accountancy firms listing *btw-advies* as a
service. That leaves **two** real ones, and both are now handled:

- `hetfamilierechtkantoor.nl` publishes `BTW: 8519.31.169` — nine digits, **no sub-number**. Those
  nine digits pass the elfproef and are the real fiscal base, so they are returned as
  `btwBaseNumber`. `btwNumber` stays **null**: `B01` is only the commonest sub-number, `B02`, `B13`,
  `B25` and `B41` all occur in this corpus, and completing it would be inventing data. (VIES does
  confirm `NL851931169B01` is live — which is exactly how a fabricated field passes review.)
- `schoonmaakbedrijf-pschaap.nl` publishes `BTW-NR: NL001711538825` — twelve digits where the format
  has nine, a `B`, then two. Its real number is `NL001711538B25`, so the site typed an `8` where the
  `B` belongs. **This Actor refuses it.** Substituting a character at a fixed position and keeping
  the result because a checksum accepted it is not parsing, it is guessing.

#### The same numbers, from a completely different corpus

Every figure above comes from a Telefoonboek.nl-sourced list, so it inherits whatever bias that
directory has. The whole benchmark was therefore re-run on an **independent sampling frame**: 298
domains drawn at random from **37,565 Dutch business websites in OpenStreetMap** (`shop`, `office`
and `craft` elements carrying a `website` tag). Same Actor, same build, same settings.

| | Telefoonboek (300) | OpenStreetMap (298) |
|---|---:|---:|
| billable `ok` | 60.0 % | 64.1 % |
| …of reachable | 83.3 % | 78.0 % |
| `unreachable` | 24.0 % | 12.1 % |
| reachable sites publishing nothing | 10.6 % | 7.3 % |
| `email` | 95.6 % | **96.9 %** |
| `phone` | 93.3 % | **97.4 %** |
| `postalCode` | 90.6 % | **95.3 %** |
| `street` | 81.7 % | **82.7 %** |
| `kvkNumber` | 49.4 % | 43.5 % |
| `btwNumber` | 27.8 % | **28.3 %** |
| KvK needing a page other than the homepage | 91.0 % | **97.6 %** |
| delivered == charged | 180 == 180 | 191 == 191 |
| margin at $0.009 | 97.9 % | 97.7 % |

Two unrelated sampling frames, essentially the same answer on every field that matters — `btwNumber`
within half a point, `street` within one. The OSM list is fresher (half the dead hosts) and its KvK
rate is a little lower, and on it **97.6 % of KvK numbers needed a page other than the homepage**.
The `rsin` field, which fills at 0 % on a trades list, fills on 2 of 191 OSM rows — foundations
exist in that frame and publish it, exactly as documented below.

#### The BTW validator, cross-checked against the authority

**50 BTW numbers read · 48 passed a check digit (96.0 %) · 34 via the elfproef, 14 via mod-97.**

The two that failed **both** algorithms were **published by the companies themselves** on their own pages. They are returned with `btwValid: false` rather than dropped — and **EU VIES independently answers `INVALID` for both** (`NL340695223B01`, `NL868487393B01`). During the build a further sample of 20 captured numbers was checked the same way: **19 of 20 passed locally and VIES agreed on every one it was asked about.**

The honest converse, and the reason `verifyBtwWithVies` exists as a separate switch: `btwValid` is a verdict on the **number**, not on the **registration**. `NL183108917B01` passes the elfproef and VIES answers `INVALID` — the number is well formed, the registration is closed. Two different questions; this Actor answers the first for free and the second on request.

#### Where the fields actually are

Per-field provenance over the 180 delivered leads, straight out of `fieldSources`:

| Field | contact | algemene-voorwaarden | privacy | homepage | over-ons | disclaimer / other |
|---|---:|---:|---:|---:|---:|---:|
| `kvkNumber` | **60.7 %** | 12.4 % | 11.2 % | 9.0 % | 3.4 % | 3.3 % |
| `btwNumber` | **64.0 %** | 14.0 % | 2.0 % | 14.0 % | — | 6.0 % |
| `legalName` | **49.2 %** | 18.5 % | 13.8 % | 6.2 % | 6.2 % | 6.1 % |
| `email` | **65.7 %** | 7.6 % | 9.3 % | 10.5 % | 5.8 % | 1.2 % |
| `phone` | **64.9 %** | 7.1 % | 7.7 % | 11.3 % | 4.8 % | 4.2 % |
| `postalCode` | **63.8 %** | 6.1 % | 8.6 % | 11.7 % | 5.5 % | 4.2 % |

That table is why the crawl order is `/contact` → `/algemene-voorwaarden` → `/over-ons` → `/privacy` → `/disclaimer`, and it is a measurement rather than a habit. It is also why `/privacy` — a page most contact scrapers never open — carries **11.2 %** of the KvK numbers.

#### Transport and cost

Datacenter proxy, got-scraping, **Chromium never launched**. Over the 300-domain run: **977 pages fetched for 300 domains (3.26 per domain, 4.53 per delivered lead)**, 12 domains blocked (4.0 %), peak memory **266 MB of the 512 MB allocated**, wall clock **700 s**.

Marginal cost measured as the **slope** between a 10-domain run (6 leads, $0.00170) and this 300-domain run (180 leads, $0.03472): **$0.000190 per delivered lead**. Every page fetched for a domain that never becomes a lead is paid for by this Actor, not by you — at $0.000116 per domain examined, this would still break even if only **1 domain in 78** produced a lead. It produces one in **1.67**.

***

### ⚠️ Honest limits — read these before you buy

- **A quarter of a directory-sourced Dutch domain list is dead.** 72 of 300 (24.0 %) did not serve a usable homepage: NXDOMAIN, TLS failures, refused connections, and a few hosts answering HTTP 200 with a 354-byte placeholder. Every row carries `homepageStatus` so you can tell a dead host from a 404 from a block. **Never charged.**
- **`kvkNumber` fills at 49 %, not at 100 %, and that is the corpus, not the parser.** Art. 3:15d is widely under-complied with by Dutch micro-businesses, and the KvK number is the field they most often omit. Where a site names its register at all, this Actor reads the number on **97.6 %** of domains (see *Is 49 % a parser gap?* above), and where the raw bytes carry a labelled number it finds **100 %** of them — checked page by page against an independently written parser across 457 captured pages, in both directions: every row that emits *no* number is also asserted to have none hiding in its bytes.
- **`btwNumber` fills at 28 %**, and measured identically (27.8 % / 28.3 %) on two unrelated corpora. Art. 3:15d sub f only bites on a provider that is *btw-plichtig*, and plenty of small providers simply do not print it. Of the domains that mention VAT but yield no number, 23 of 25 mention it only in prose.
- **`legalName` fills at 36 %.** It is only ever emitted when the page states a rechtsvorm (`B.V.`, `V.O.F.`, `Stichting` …), because that is what makes a name a *legal* name. `companyName` — the trading name from `og:site_name`, JSON-LD or the title — fills at **91 %** and is a separate field, so neither has to lie. It returns null rather than a page name: a site whose title is just "Contact" has not told us what it is called.
- **`rsin` filled on 0 of 180 rows on the trades list, and it is not a dead field.** The RSIN is published by *stichtingen* — an ANBI must — and a trades-MKB list contains none. It fills on **2 of 191** rows of the OpenStreetMap corpus, which does contain them, and a direct probe of 10 Dutch ANBI pages found it on 2 of the 6 that were reachable. Expect it on foundations, not on loodgieters.
- **`kvkRegion` (0.5 %) and `privacyEmail` (3.3 %) are genuinely rare**, and `director` is **off by default**: Dutch law does **not** require a company to name its bestuurder, unlike §5 TMG in Germany, so it filled on only 3.2 % of pages in a separate run with `includePersonalNames` enabled. All three are reported at their real rate rather than left out.
- **This Actor does not confirm anything against the Handelsregister.** It reports what the company published, plus local check digits. A KvK cross-check needs a paid KvK API key and is a different product.
- **A CDN-fronted enterprise tail is genuinely walled.** 3.7 % of this list answered with a challenge page on every attempt. `Escalate to residential on a block` is available and off by default. This Actor is built for the MKB long tail, which is unaffected.
- **`robots.txt` is not consulted by default.** The art. 3:15d block is a statutory public disclosure that the law requires to be *permanently accessible*. Turn `Respect robots.txt` on if your own policy requires it.

***

### 🧾 Output

One row per domain. Set **Output mode** to *Full* to add `emails[]`, `phones[]`, every KvK and BTW number found, the URLs tried, per-page scores and which fields came from structured data.

| Field | Description |
|---|---|
| `domain` | Registrable domain (eTLD+1, `www`-stripped) — the stable id. One domain = one row = at most one charge. |
| `website` | The homepage that was fetched. |
| `disclosureUrl` | The page the identity block was read from. |
| `disclosurePageType` | `contact` · `algemene-voorwaarden` · `privacy` · `over-ons` · `disclaimer` · `klantenservice` · `homepage` · `other`. |
| `discoveryMethod` | `link:<type>` (a ranked footer link), `guess:<path>` (the URL ladder), `homepage`, or `input-url`. |
| `companyName` | Trading name, from `og:site_name`, JSON-LD `Organization`, or the page title. |
| `legalName` | Statutory name **including the rechtsvorm** — only emitted when the page states one. |
| `legalForm` | `B.V.` · `N.V.` · `V.O.F.` · `C.V.` · `Stichting` · `Vereniging` · `Maatschap` · `Eenmanszaak` · `B.V. i.o.` |
| `kvkNumber` | The 8-digit KvK-nummer, taken only from a labelled position. |
| `kvkValid` | Eight digits, label-anchored. **The Handelsregister publishes no check digit**, so this is a shape verdict and the README says so rather than implying more. |
| `kvkRegion` | The KvK region when the page names one (`Oost-Nederland`, `Den Haag` …). Rare. |
| `btwNumber` | Normalised to `NL` + 9 digits + `B` + 2, from any on-page spelling — `820652623B01`, `8548.89.942.B01`, `NL 1831.08.917.B.01`. |
| `btwValid` | **Check-digit verdict**: the elfproef for rechtspersonen, mod-97 for the post-2020 btw-id. `false` means the company published a number that fails both. |
| `btwCheckMethod` | `elfproef` · `mod97` · `null` — which algorithm accepted it. |
| `btwSubNumber` | The two-digit sub-number (`01`, `13`, `41` …). `B00` is never issued and is flagged invalid. |
| `btwBaseNumber` | The nine-digit fiscal base, for the sites that publish a BTW number **without** a sub-number. Only ever set when `btwNumber` is null, and only when those nine digits pass the elfproef. |
| `btwViesValid` · `btwViesName` | Only with **Check the BTW against EU VIES**: is the registration open, and the registered name. `null` when VIES itself did not answer — never a false `invalid`. |
| `rsin` | RSIN, where published (stichtingen / ANBI). |
| `street`, `houseNumber` | Parsed separately, never a blob. |
| `postalCode` | `1234 AB`, normalised and upper-cased. |
| `city` | Trimmed to a place name by shape, so `6814 DA Arnhem BMV Makelaars` yields `Arnhem` and `Alphen aan den Rijn` survives intact. |
| `postbus` | The PO-box number, where the company publishes one. |
| `email` | Best contact address. Role addresses only, unless you opt in. `mailto:`, Cloudflare `data-cfemail` cloaking and `naam [at] domein [punt] nl` are all decoded. |
| `privacyEmail` | A dedicated `privacy@` / `avg@` / `fg@` address, where published. |
| `emailMxValid` | Only with **Verify the e-mail domain accepts mail (MX)**. |
| `phone`, `mobile` | E.164 (`+31…`), never derived from an IBAN. |
| `iban` | Validated with the ISO 13616 mod-97 check; a mis-read never ships. |
| `director` | Bestuurder / directeur / eigenaar. **Off by default** — personal data under the AVG. |
| `socialProfiles` | LinkedIn / Facebook / Instagram / X / YouTube / TikTok company pages; share widgets excluded. |
| `fieldSources` | **Per field, the page type it was read from.** The audit trail for every row. |
| `pagesFetched` · `homepageStatus` | How many pages were read, and what the homepage answered. |
| `status` | `ok` · `partial` · `no_disclosure_found` · `blocked` · `unreachable` · `not_attempted` · `error`. **Only `ok` is charged.** |
| `errorMessage` | Why a non-`ok` row is non-`ok`, in plain English. |
| `fetchedAt` | ISO timestamp. |

#### A real row, from the benchmark run

```json
{
  "domain": "darsilmedia.com",
  "website": "https://darsilmedia.com",
  "disclosureUrl": "https://www.darsilmedia.com/contact",
  "disclosurePageType": "contact",
  "discoveryMethod": "link:contact",
  "companyName": "darsilmedia",
  "legalName": "Darsil Media B.V.",
  "legalForm": "B.V.",
  "kvkNumber": "65404777",
  "kvkValid": true,
  "btwNumber": "NL856099168B01",
  "btwValid": true,
  "btwCheckMethod": "elfproef",
  "btwSubNumber": "01",
  "street": "Oude Baan",
  "houseNumber": "4",
  "postalCode": "4825 BL",
  "city": "Breda",
  "email": "sales@darsilmedia.com",
  "phone": "+31625331997",
  "mobile": "+31625331997",
  "iban": "NL40TRIO0391126792",
  "fieldSources": {
    "kvkNumber": "contact", "btwNumber": "contact", "legalName": "contact",
    "street": "contact", "postalCode": "contact", "city": "contact",
    "email": "contact", "phone": "contact", "iban": "contact"
  },
  "pagesFetched": 2,
  "homepageStatus": 200,
  "status": "ok",
  "fetchedAt": "2026-09-20T04:34:57.820Z"
}
```

And one assembled across three pages — the reason `fieldSources` exists:

```json
{
  "domain": "hansanders.nl",
  "legalName": "Hans Anders Nederland B.V.",
  "kvkNumber": "23061599",
  "street": "Papland", "houseNumber": "21", "postalCode": "4206 CK", "city": "Gorinchem",
  "email": "sponsoring@hansanders.com",
  "phone": "+31183697500",
  "fieldSources": {
    "kvkNumber": "algemene-voorwaarden",
    "legalName": "algemene-voorwaarden",
    "street": "disclaimer", "postalCode": "disclaimer", "city": "disclaimer",
    "phone": "disclaimer",
    "email": "over-ons"
  },
  "disclosurePageType": "algemene-voorwaarden",
  "pagesFetched": 8,
  "status": "ok"
}
```

***

### ⌨️ Input

#### Getting domains in

| Field | What it does |
|---|---|
| **Company domains** | Bare domains, full URLs, a direct `/contact` URL, or even an e-mail address — the domain is extracted. |
| **Start URLs** | The same list in Apify's Start-URLs shape, so the Console link-list UX and file uploads work. |
| **Source dataset ID** / **field** | Chain straight off another Actor — a Google Maps Netherlands run, a Telefoonboek scrape, a Funda makelaar list. |
| **Domain list URL** | A public `.txt` / `.csv` / `.tsv`. A UTF-8 BOM is stripped, so an Excel export works. |
| **Exclude domains** | A suppression list, applied **before any request**, so an excluded domain never costs a fetch. |

#### Bounds and network

`Max domains` (default 1000, `0` = no cap) · `Request concurrency` (8) · `Max pages per domain` (8) · `Request timeout` (20 s) · `Max retries per request` (3, each on a **fresh proxy IP**) · `Proxy` (datacenter) · `Proxy country` · `Escalate to residential on a block` (off) · `User-Agent override` · `Custom headers` · `Respect robots.txt` (off).

#### Discovery

`URL-guess ladder` · `Extended path ladder` (opt-in second tier) · `Read every statutory page, not just the first` (on) · `Merge fields across pages` (on) · `Read structured data too` (on — schema.org JSON-LD plus the inline ACF payloads Dutch WordPress themes render; it recovered fields on real domains where the visible text carried nothing).

#### Verification and the billing bar

`Validate the btw-identificatienummer` (on) · `Check the BTW against EU VIES` (off) · `Verify the e-mail domain accepts mail (MX)` (off) · `Include personal names` (off, AVG) · `Include personal e-mail addresses` (off) · **`Minimum fields for a billable lead`** (2 of KvK / BTW / e-mail / phone — anything below the bar is delivered free as `partial`).

#### Filters and output

`Only rows with a KvK number` · `Only rows with a BTW number` · `Only rows whose BTW passes the check digit` · `Filter by rechtsvorm` · `Filter by postcode prefix` · `Filter by city` · `Only push resolved leads` · `Deduplicate by domain` (on) · `Deduplicate by KvK number` (off) · `Include the raw disclosure text` (off) · `Output mode` (lead / full).

```json
{
  "domains": ["aannemersbedrijfvanderhelm.nl", "brandforlife.nl", "bloemenkiosk-jancop.nl"],
  "maxItems": 3,
  "outputMode": "lead"
}
```

***

### 💰 Pricing

**$0.009 per delivered lead — $9 per 1,000.** Pay-per-event, one event, **no start fee** and no per-dataset-item surcharge.

**You are charged for a `status: ok` row and nothing else.** Delivery and billing happen in the same call, so a charge cap truncates delivery too and can never leave you holding a row you did not pay for — or paying for one you did not receive. Across five 300-domain benchmark runs, delivered `ok` rows and `chargedEventCounts` matched **exactly every time** (180 == 180, 182 == 182, 191 == 191, 192 == 192, 181 == 181).

Never charged: unreachable hosts · blocked sites · sites that publish no readable disclosure · rows below your `minFieldsRequired` bar · rows a filter removed · duplicates collapsed before the fetch · domains a run timeout never reached.

At the default bar of 2 fields, the benchmark list of 300 domains costs **$1.62** and returns 180 leads.

***

### ❓ FAQ

**Is this legal?** The art. 3:15d block is information the company is **required by Dutch law** to make publicly and permanently accessible about itself. That makes it the most defensible starting point for Dutch B2B prospecting there is. You remain responsible for each site's Terms of Service, and for outreach compliance — the AVG/GDPR and, for unsolicited electronic communication, Telecommunicatiewet art. 11.7.

**Do I need a KvK API key?** No. Nothing here touches kvk.nl. This reads the company's own website.

**How is this different from a KvK Handelsregister scraper?** Opposite direction. A register scraper starts from a company name or a KvK number and looks it up. This starts from a **domain** — which is what you actually have after a Maps scrape, a directory export or a CRM dump — and gives you the identity behind it, plus the contact details the register does not hold.

**What does `btwValid: false` mean?** The company published a number that fails **both** Dutch check-digit algorithms. Usually a typo on their site. The number is still returned so you can see what they published; it is flagged, not silently dropped.

**What is `btwBaseNumber`?** Some sites print the nine-digit fiscal number and stop, with no `B01`
on the end. Those nine digits are real and are check-digit verified, but the full
btw-identificatienummer needs the sub-number, and guessing it would be inventing data. So the base
goes in its own field and `btwNumber` stays null.

**Why is `kvkValid` true for every KvK number?** Because the Handelsregister publishes no check digit for it. `kvkValid` means "eight digits, taken from a labelled position next to a register label" — which is a real guard, since that is exactly what an AGB code or an order number is not. It is not a claim that the company exists.

**Can I get the VIES registration status?** Turn on **Check the BTW against EU VIES**. It is off by default because VIES is an external service with hard rate limits; a member state that does not answer returns `null`, never `false`.

**Does it render JavaScript?** No — and on this corpus it does not need to. Dutch MKB sites are server-rendered; the enterprise tail that is not is also the tail behind the bot walls. Rows that could not be read come back `blocked` or `no_disclosure_found`, free.

**Which pages does it read?** The homepage plus, in measured yield order, `/contact`, `/algemene-voorwaarden`, `/over-ons`, `/privacy`, `/disclaimer`, `/klantenservice` — discovered from the site's own links first, then guessed. You can replace both ladders.

**Will it charge me twice for the same company?** No. Exact duplicate hostnames are always collapsed before any request; `Deduplicate by domain` additionally collapses subdomains; `Deduplicate by KvK number` collapses several brand domains that share one registration.

**How fast?** 300 domains in about 12 minutes at the default concurrency of 8, 512 MB, no browser.

**What if my list is one company per row with other columns?** Point **Domain list URL** at the CSV; the column named `website` / `url` / `domain` is used, or the most domain-looking column if there is no header.

***

### ⚖️ Legal & source

Source: the statutory disclosure published on each company's own website under **Burgerlijk Wetboek art. 3:15d lid 1** (the Dutch implementation of Directive 2000/31/EC art. 5). Only publicly served pages are read; no login, no paywall, no personal profile.

`robots.txt` is not consulted on the default path — the pages carrying this block are ones the law requires to be permanently accessible — and the `Respect robots.txt` switch is there if your own policy differs.

`director` is personal data under the AVG/GDPR and is **off by default**; enabling it makes you the controller for that field. Personal e-mail addresses are likewise opt-in, and only role addresses are returned otherwise. You are responsible for complying with each site's Terms of Service and with all applicable outreach law when you contact these companies.

# Actor input Schema

## `domains` (type: `array`):

One entry per Dutch company. Paste bare domains ("brandforlife.nl"), full homepage URLs ("https://www.b2-advocaten.nl"), a direct /contact or /algemene-voorwaarden URL, or even an e-mail address — everything is normalised to the registrable domain (eTLD+1), so one company = one row = at most one charge. Leave empty to run the documented three-domain sample.

## `startUrls` (type: `array`):

The same input in Apify's Start-URLs shape, so the Console link-list UX and "link list from a file" uploads work. Merged with Company domains.

## `sourceDatasetId` (type: `string`):

Chain straight off another Actor's dataset — a Google Maps Netherlands run, a Telefoonboek.nl scrape, a Funda makelaar list — instead of pasting a list. Every item's website field becomes an input domain.

## `sourceDatasetField` (type: `string`):

Which field of the source dataset holds the website. Falls back automatically through website, url, domain, web, site, homepage, link, companyWebsite.

## `domainsFileUrl` (type: `string`):

Public URL of a newline-delimited .txt or a .csv/.tsv of domains, for very large lists. A CSV header naming the source field (or website/url/domain) picks the column; otherwise the most domain-looking column is used. A UTF-8 BOM is stripped, so an Excel export works.

## `excludeDomains` (type: `array`):

Already-contacted or unwanted domains. Applied BEFORE any request is made, so an excluded domain never costs a fetch and can never be billed.

## `maxItems` (type: `integer`):

Hard cap on domains processed this run — your billing guard. Default 1000. Set 0 to process the whole input list.

## `requestConcurrency` (type: `integer`):

Domains processed in parallel. 8 is the measured stable setting on this corpus; the target sites are small business hosts, so keep it modest.

## `maxRequestsPerDomain` (type: `integer`):

Caps the ladder: homepage + discovered statutory links + URL guesses. The Dutch disclosure is scattered across /contact, /algemene-voorwaarden, /privacy and /over-ons, so this is the single biggest lever on both fill rate and cost. Measured average on a real MKB list is well under the default.

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout. Each attempt also carries an external hard cap, because a stalled proxied request can outlive this timer.

## `maxRetriesPerRequest` (type: `integer`):

Attempts per URL, each on a FRESH proxy IP. Retrying a flagged IP through the same IP is only a slower way to get the same answer, which is why this loop exists instead of the HTTP client's own retry.

## `proxyConfiguration` (type: `object`):

Apify DATACENTER proxy is the default and was MEASURED sufficient across a 716-domain Dutch MKB corpus. Switch to RESIDENTIAL only for the minority of sites behind an enterprise bot wall.

## `proxyCountry` (type: `string`):

Country pin. Applies ONLY when the proxy above uses the RESIDENTIAL group — datacenter pools ignore a country pin.

## `escalateToResidentialOnBlock` (type: `boolean`):

Retry a 403/429/challenge page on RESIDENTIAL-NL. Off by default because it costs residential traffic on every blocked page; turn it on when a run reports a high blocked count.

## `userAgent` (type: `string`):

Advanced: replace the default desktop Chrome User-Agent.

## `customHeaders` (type: `object`):

Advanced: extra request headers merged over the defaults (Accept-Language is nl-NL,nl;q=0.9,en;q=0.7).

## `respectRobotsTxt` (type: `boolean`):

Read each site's robots.txt and skip URLs its wildcard user-agent group disallows. Off by default; the art. 3:15d block is a statutory public disclosure and the pages carrying it are normally crawlable.

## `disclosurePaths` (type: `array`):

Tried when the homepage exposes no statutory link. Ordered by measured yield on a real Dutch MKB corpus — /contact first, then /algemene-voorwaarden, then /over-ons and /privacy. Extend or replace freely.

## `extendedPaths` (type: `array`):

A second, slower tier tried only when the first ladder finds nothing. Empty by default because it costs extra requests on every unresolved domain. Suggested: /klantenservice, /colofon, /privacybeleid, /bedrijfsgegevens, /wie-zijn-wij.

## `scoreAllCandidates` (type: `boolean`):

Evaluate EVERY discovered statutory page and keep the best, instead of stopping at the first one that loads. Keep this ON: first-hit-wins is how a franchise's /contact page hands you the franchisor's KvK number.

## `mergeAcrossPages` (type: `boolean`):

Dutch law names no disclosure page, so a site scatters the block: /contact carries the phone, /algemene-voorwaarden carries the KvK number. With this ON the row is assembled from all of them and fieldSources records which page each field came from. The legal name, KvK and BTW are still taken from ONE page, so two companies can never be stitched into one row.

## `useStructuredData` (type: `boolean`):

Also read schema.org JSON-LD (vatID / taxID / legalName / PostalAddress) and the inline JSON payloads Dutch WordPress themes render (an ACF block holding kvk\_number and btw\_number). Fills ONLY what the visible text did not, so a stale JSON-LD block can never override a live footer.

## `validateBtwNumber` (type: `boolean`):

Run both Dutch check-digit algorithms locally, with zero external calls, and emit btwValid + btwCheckMethod: the classic elfproef for rechtspersonen, and the mod-97 check for the btw-id issued to eenmanszaken since 2020, which is random and passes no elfproef. A validator that knows only the first rejects every sole trader registered in the last six years.

## `verifyBtwWithVies` (type: `boolean`):

A second, different question: btwValid says the number is well formed and its check digit is right; VIES says whether the registration is still open. They disagree in real life — a number can pass the elfproef while VIES answers INVALID. Off by default: it is an external service with its own rate limits. A member state that does not answer returns null, NEVER false, so you can never get a false invalid.

## `verifyMx` (type: `boolean`):

DNS MX lookup on the extracted e-mail's domain; emits emailMxValid. One DNS round-trip per unique mail domain.

## `includePersonalNames` (type: `boolean`):

Extract the named director or owner when the site publishes one. Personal data under the AVG/GDPR — OFF by default, and enabling it makes you the controller for that field. Note that Dutch law does not require a director to be named, so this fills at a far lower rate than the German Geschäftsführer field it mirrors.

## `includePersonalEmails` (type: `boolean`):

Include voornaam.achternaam@ style addresses. OFF by default — only role addresses (info@, contact@, administratie@, privacy@ …) are returned.

## `minFieldsRequired` (type: `integer`):

How many of {KvK number, BTW number, e-mail, phone} a company must yield before the row counts as a resolved lead (status: ok) and is therefore charged. Anything below the bar is delivered as status: partial, free.

## `requireKvk` (type: `boolean`):

Drop every row whose KvK-nummer could not be read. Dropped rows are never pushed and never billed.

## `requireBtw` (type: `boolean`):

Drop every row whose btw-identificatienummer could not be read. Useful when you need a VAT number for invoicing or reverse-charge checks.

## `onlyValidBtw` (type: `boolean`):

Keep only rows where btwValid is true. A published number that fails both Dutch check algorithms is usually a typo on the site — this drops it rather than shipping it as a fact.

## `legalFormFilter` (type: `array`):

Keep only companies whose legal form matches, e.g. \["BV"] for besloten vennootschappen or \["Stichting"] for foundations. Accepted: BV, NV, VOF, CV, Stichting, Vereniging, Maatschap, Eenmanszaak. Empty = keep everything.

## `postalCodeFilter` (type: `array`):

Postcode prefixes, e.g. \["10", "30"] for the Amsterdam and Rotterdam ranges, or \["1012"] for one district. Empty = keep everything.

## `cityFilter` (type: `array`):

Exact city names, e.g. \["Amsterdam", "Utrecht"]. Case-insensitive. Empty = keep everything.

## `onlyResolved` (type: `boolean`):

Drop partial, not-found, blocked and unreachable rows entirely. OFF by default so you can see exactly what happened to every domain you submitted — those rows are always free.

## `dedupeByDomain` (type: `boolean`):

Collapse subdomains onto their registrable domain (eTLD+1) before fetching. Exact duplicate hostnames are ALWAYS collapsed regardless of this switch — that is the double-charge guard, not a preference.

## `dedupeByKvk` (type: `boolean`):

Collapse several brand domains that turn out to share one KvK-nummer into a single row. Off by default, because you usually want a row per domain you submitted.

## `includeRawText` (type: `boolean`):

Attach the winning page's plain text (rawText) for audit or LLM post-processing. Off by default — it multiplies dataset size.

## `outputMode` (type: `string`):

"Lead" is one flat row per company. "Full" adds emails\[], phones\[], every KvK and BTW number found, the URLs tried, the per-page scores and which fields came from structured data.

## Actor input object example

```json
{
  "domains": [
    "aannemersbedrijfvanderhelm.nl",
    "brandforlife.nl",
    "bloemenkiosk-jancop.nl"
  ],
  "startUrls": [],
  "sourceDatasetId": "",
  "sourceDatasetField": "website",
  "domainsFileUrl": "",
  "excludeDomains": [],
  "maxItems": 1000,
  "requestConcurrency": 8,
  "maxRequestsPerDomain": 8,
  "requestTimeoutSecs": 20,
  "maxRetriesPerRequest": 3,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "proxyCountry": "none",
  "escalateToResidentialOnBlock": false,
  "userAgent": "",
  "customHeaders": {},
  "respectRobotsTxt": false,
  "disclosurePaths": [
    "/contact",
    "/contact/",
    "/algemene-voorwaarden",
    "/over-ons",
    "/privacy",
    "/disclaimer",
    "/contactgegevens",
    "/algemene-voorwaarden/",
    "/privacyverklaring",
    "/voorwaarden"
  ],
  "extendedPaths": [],
  "scoreAllCandidates": true,
  "mergeAcrossPages": true,
  "useStructuredData": true,
  "validateBtwNumber": true,
  "verifyBtwWithVies": false,
  "verifyMx": false,
  "includePersonalNames": false,
  "includePersonalEmails": false,
  "minFieldsRequired": 2,
  "requireKvk": false,
  "requireBtw": false,
  "onlyValidBtw": false,
  "legalFormFilter": [],
  "postalCodeFilter": [],
  "cityFilter": [],
  "onlyResolved": false,
  "dedupeByDomain": true,
  "dedupeByKvk": false,
  "includeRawText": false,
  "outputMode": "lead"
}
```

# Actor output Schema

## `records` (type: `string`):

The dataset of Dutch company records read from each site's statutory art. 3:15d disclosure (one item per domain).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "aannemersbedrijfvanderhelm.nl",
        "brandforlife.nl",
        "bloemenkiosk-jancop.nl"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/nl-kvk-website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "aannemersbedrijfvanderhelm.nl",
        "brandforlife.nl",
        "bloemenkiosk-jancop.nl",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/nl-kvk-website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "aannemersbedrijfvanderhelm.nl",
    "brandforlife.nl",
    "bloemenkiosk-jancop.nl"
  ]
}' |
apify call scrapersdelight/nl-kvk-website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/nl-kvk-website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZH8fcsyHeeDSNd5cX/builds/2crSyc9zjfMdBwa0s/openapi.json
