# Polish Company Data Scraper - KRS, NIP, REGON & Contacts (`scrapersdelight/pl-dane-spolki-website-contact-scraper`) Actor

Turn Polish company domains into registry-grade B2B leads from each site's statutory disclosure (KSH art. 206 / art. 374): firma, KRS number and registry court, checksum-validated NIP and REGON, share capital, registered office, email and phone. $0.007 per disclosure. No login, no browser.

- **URL**: https://apify.com/scrapersdelight/pl-dane-spolki-website-contact-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$7.00 / 1,000 per company disclosure returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 🇵🇱 Polish Company Data Scraper — KRS, NIP, REGON, Share Capital & Contacts

Give it a list of Polish company domains. Get back, for each one, the **statutory disclosure the
company is legally required to publish on its own website**: the registered *firma*, the **KRS**
entry number and the *sąd rejestrowy* that holds the file, a **checksum-validated NIP**, a
**checksum-validated REGON**, the **kapitał zakładowy** as a number, the registered office, email
and phone.

No login. No API key. No browser. Plain HTTP through the Apify proxy.

**You are never charged for a domain that is dead, blocked, or publishes no disclosure.**

***

### ⚖️ Why this data exists, and why it is clean

**Kodeks spółek handlowych art. 206 §1** requires every *spółka z ograniczoną odpowiedzialnością*,
and **art. 374 §1** every *spółka akcyjna*, to state — in its commercial letters and orders **and
`na stronach internetowych spółki`** —

1. the **firma**, the **siedziba** and the **adres**;
2. the **sąd rejestrowy** holding the company's file, and the **number** under which it is entered
   in the register (the **KRS**);
3. the **NIP**;
4. the **wysokość kapitału zakładowego** — and, for an S.A. under art. 374 §1 pt 4, the capital
   actually **paid up**.

Art. 127 §5 carries the same duty to a *spółka komandytowo-akcyjna*, and art. 300(101) to a *prosta
spółka akcyjna*.

So this is not scraped inference and it is not a directory's copy of a company's details. It is a
disclosure the company itself is **obliged by statute** to publish about itself, on its own site,
and to keep accurate.

#### The honest half of that: the duty does not bind everyone

The KSH website duty binds **sp. z o.o.**, **S.A.**, **S.K.A.** and **P.S.A.** — and **nobody
else**. A *jednoosobowa działalność gospodarcza* (a sole trader, and the commonest legal form in
Poland by a wide margin), a *spółka jawna*, a *spółka komandytowa* and a *spółka cywilna* have **no
such obligation at all**.

That is why every row carries **`statutoryScope`**, naming the provision that binds that company —
`KSH art. 206 §1`, `KSH art. 374 §1`, `KSH art. 127 §5 (via art. 374)` — or
`none (no KSH website-disclosure duty)`. A domain with no KRS is very often **correct output**
rather than a parser failure, and this field is how you tell the two apart.

> **Feed it business domains and the hit rate rises by about a third — measured.** Across two
> corpora built on opposite principles, the yield went from **45.5%** of reachable domains on a
> maps/POI-sourced list to **61.9%** on a traffic-ranked one: **+36%**, not the doubling we
> expected before measuring it. A maps-sourced list is full of sole traders, schools and market
> stalls who owe you nothing. A list of *sp. z o.o.* and *S.A.* websites — from a KRS export, a
> trade-association member list, a tender register, an exhibitor list — should do better than
> either, but we have not measured that one and so do not claim a number for it.

***

### 🔍 What it does, per domain

Two to five hops, plain HTTP (`got-scraping` + `cheerio`, **no browser**), Apify **datacenter**
proxy with a residential and an optional Unblocker escalation.

1. **Fetch the homepage and read its footer.**

2. **Climb the ladder.** This is the part that matters in Poland. Measured over real Polish
   business domains, **69-72% of the disclosures found needed a page past the homepage** — measured
   on two independent corpora that agree to within 3.3 points — the opposite of the UK, where the footer usually carries it, and close to the
   Dutch pattern. So the Actor ranks the site's own links by **anchor text *and* href** and walks
   them in the order Polish sites actually use:

   `dane spółki` › `kontakt` › `regulamin` › `nota prawna` › `polityka prywatności` › `o nas`

   then falls back to guessed conventional paths (`/kontakt`, `/regulamin`, `/o-nas`,
   `/dane-spolki`, `/polityka-prywatnosci`, …), the XML sitemap, and the WordPress page index.

   The **polityka prywatności page is a first-class source here, not a fallback**: a Polish RODO
   clause names the administrator's full firma, siedziba, KRS and NIP in one sentence.

3. **Merge field by field**, recording which page each field came from. The KRS on `/regulamin`,
   the email on `/kontakt` and the share capital on `/o-nas` is an ordinary Polish site.

Every statutory field is read inside a **statement window** anchored on a disclosure phrase, never
by a page-wide scan.

***

### 📦 What you get — every field

#### Identity

| Field | What it is |
|---|---|
| `registeredName` | The **firma** as published — `Żabka Polska sp. z o.o.`, `PESA Bydgoszcz SA` |
| `registeredNameSource` | How it was resolved: `label`, `suffix`, `clause` |
| `tradingName` | A brand or trading name stated as distinct from the firma |
| `companyName` | Best available name (firma, else JSON-LD `legalName`, else `og:site_name`) |
| `legalForm` | `sp. z o.o.` · `S.A.` · `P.S.A.` · `S.K.A.` · `sp. z o.o. sp.k.` · `sp.k.` · `sp.j.` · `sp.p.` · `s.c.` · `fundacja` · `stowarzyszenie` · `spółdzielnia` · `jednoosobowa działalność gospodarcza` |
| `statutoryScope` | **Which KSH provision binds this company's website**, or that none does |

#### Registry

| Field | What it is |
|---|---|
| `krs` | The **KRS** entry number, canonical 10 digits, zero-padded |
| `krsRaw` | Exactly as the page printed it |
| `krsSource` | `clause` (from the long statutory sentence), `label`, or `bare` |
| `krsFormatValid` | **Format check only.** The KRS carries **no checksum** — see *Honest limits* |
| `krsScheme` | `PL-KRS` |
| `registerUrl` | Direct deep link into the **Ministry of Justice's own public KRS API**, which returns the current extract (*odpis aktualny*) as JSON |
| `krsCourt` | The **sąd rejestrowy** — `Sąd Rejonowy dla Wrocławia-Fabrycznej we Wrocławiu` |
| `krsCourtDivision` | Its commercial division — `VI Wydział Gospodarczy` |
| `nip` | The **NIP**, 10 digits |
| `nipValid` | **Weighted checksum result** (weights 6,5,7,2,3,4,5,6,7 mod 11; a remainder of 10 is invalid) |
| `nipFormatted` | `527-212-86-91` |
| `nipSource` | `label`, `prefix` (a `PL`-prefixed form), `json-ld`, or `loose-checksummed` |
| `vatNumber` | The intra-EU VAT identifier — `PL` + the NIP, ready for VIES |
| `regon` | The **REGON**, 9 digits (entity) or 14 (a local unit of one) |
| `regonValid` | **Its own checksum**, a different rule from the NIP's |
| `regonLength` | `9` or `14`, so you can tell an entity from a branch |
| `shareCapital` | **Kapitał zakładowy as a NUMBER** you can filter and sort on |
| `shareCapitalCurrency` | `PLN` (or `EUR`/`USD` where a company states one) |
| `shareCapitalRaw` | The string it was parsed from, so you can check the parse yourself |
| `shareCapitalPaid` | The capital stated as **paid up** |
| `shareCapitalPaidInFull` | `true` when the page says *wpłacony w całości*; `null` when it says nothing — **never `false`** |
| `registeredOffice` | The **siedziba** as a postal line — `ul. Legnicka 48A, 54-202 Wrocław` |
| `registeredOfficePostcode` | `54-202` |
| `registeredOfficeCity` | `Wrocław` |

#### Contact

| Field | What it is |
|---|---|
| `email` | Best email — on-domain preferred |
| `emails` | Up to 10, deduped, Cloudflare `[email protected]` decoded |
| `emailConflict` | Set when the primary email is on a **different** domain from the site and is not a freemail provider — a signal the site is run by an agency |
| `phone` | Best phone, E.164 `+48…` |
| `phones` | Up to 10 |
| `addressLine` / `postcode` / `city` | A contact/trading address where the page gives one distinct from the siedziba |
| `officerName` / `officerRole` | A named *prezes zarządu*, *właściciel*, *dyrektor* or *inspektor ochrony danych* where the site states one |
| `socialLinks` | LinkedIn, Facebook, Instagram, X, YouTube, TikTok profiles |
| `termsUrl` / `privacyPolicyUrl` | The site's *regulamin* and *polityka prywatności* |

#### Provenance & accounting

| Field | What it is |
|---|---|
| `domain` / `inputUrl` / `resolvedUrl` | What you gave us and what answered |
| `disclosureUrl` | **The exact page the disclosure was read from** |
| `disclosurePagePath` | Which rung of the ladder that was — `/`, `/kontakt`, `/regulamin`, … |
| `disclosureSource` | `homepage`, `legal-page` or `merged` (split across pages) |
| `discoveryChannel` | `homepage`, `anchor`, `guess`, `sitemap`, `wpjson`, `jsassets` |
| `pagesParsed` / `pagesChecked` | What it cost us |
| `fieldSources` | *(optional)* the page path each individual field came from |
| `fieldCount` | How many of the 22 value fields are populated |
| `stableId` | The KRS when well-formed, else `NIP…`, else the registrable domain |
| `status` / `missReason` | `ok`, or exactly why not |
| `elapsedMs` / `fetchedAt` | |

***

### 📊 Measured results — real numbers, not claims

#### Two corpora, on purpose

**One corpus measures a parser. It does not measure a population.** Whether a label yields a value
is a fact about this Actor; what share of Polish companies publish a KRS at all is a fact about the
*list you feed it*. So every population number below was measured twice, on two corpora built in
completely different ways, with biases that point in **opposite** directions:

| | Frame A | Frame B |
|---|---|---|
| **source** | OpenStreetMap `website` tags | the Tranco .pl tail (ranked past 100,000) |
| **sampling** | someone walked past it and mapped it | traffic |
| **built from** | `office` + `shop` + `craft` + `amenity` + `healthcare` + `tourism` across 30 Polish cities | a 1M-row popularity list, `.pl` only |
| **size** | **16,400 unique hosts** | 7,994 candidates, 4,000 sampled |
| **skews towards** | local shops, workshops, schools, **sole traders — who owe no KSH duty at all** | larger companies — **disproportionately sp. z o.o. and S.A., who do** |

They bracket the population instead of agreeing by construction. A number that holds across both is
a fact about Poland; a number that moves is a fact about the sampling frame, and is quoted below as
a **range**.

Frame A's suffix split — `.pl` 81.1%, `.com.pl` 6.6%, `.com` 5.8%, `.eu` 3.5%, rest 1.8%; **`.pl`
family 87.6%, everything else 12.4%**. No TLD filter is applied, deliberately: the non-`.pl` eighth
skews towards the larger companies (`cersanit.com`, `asseco.com`, `maspex.com`, `solarisbus.com` all
publish a full KRS statement), so filtering to `.pl` would silently discard them.

*Frame A was also checked for cross-border contamination, because a country-wide Overpass area query
can silently pull in a neighbour's businesses. It cannot here — the query is 30 city-sized bounding
boxes, none of which reaches a border. Confirmed empirically: 35 foreign ccTLD hosts in 16,400
(0.21%), spread diffusely over 15 countries, and of the 13 on the six land-neighbour ccTLDs, 7 are
diplomatic missions physically located in Polish cities and the rest are international brands with a
Polish branch.*

#### Offline validation, against real captured bytes

Every fixture was captured **live, through the shipping proxy**, and the **shipping parser** is run
over it by `_validate.mjs`, which asserts each advertised field against the raw bytes it came from:

```
FIXTURE FILES: 547   DOMAINS: 237   WITH A DISCLOSURE: 124
ASSERTIONS: 5506   FAILURES: 0
```

plus a separate **165-assertion unit suite** over the checksums, the number formats, the address
assembler and the legal-form map — every case a string copied out of a named real page.

##### Checksums, each cross-checked against a *second, independently written* implementation

| | valid | published but **invalid** | not published | pass rate of those published |
|---|---|---|---|---|
| **NIP** | 118 | 4 | 2 | **97%** |
| **REGON** | 76 | 0 | 48 | **100%** |
| **KRS** | — | — | — | *format check only — the KRS has no checksum* |

The validator's checksums are written from the published rule in a **different shape** from the
shipping ones (a ten-term dot product versus an indexed nine-weight loop), so a bug in either
cannot validate itself. Each is also put through a **single-digit mutation sweep**: every one-digit
corruption of a real published number must fail.

A NIP that is published but fails its checksum is emitted with `nipValid: false`. That is a fact
about the site, not a parser failure, and hiding it would be worse than reporting it.

##### Conditional emit rate — *of the pages that carry the label, what share yielded a value*

This is the measurement that tells a sparse field apart from a broken regex. Raw fill cannot.

| Label present on | domains | value emitted | conditional rate |
|---|---|---|---|
| NIP | 118 | 117 | **99%** |
| KRS | 64 | 60 | **94%** |
| REGON | 79 | 76 | **96%** |
| registry court | 35 | 35 | **100%** |
| share capital | 19 | 18 | **95%** |

##### Negative direction — *did we miss something the bytes plainly contain?*

An independently written scan re-reads every page where the parser emitted `null` and looks for a
checksum-valid number it should have caught.

**NIP: 0 misses.** It was not 0 to begin with. The sweep found exactly one — and it earned its
keep, because the cause was general: `wtzdeotymy.ksnaw.pl` publishes `NIP 527- 21- 28 -691`, a
hyphen **and** a space between the same pair of digits, which a one-character separator class
cannot cross. The same sweep also found `KRS0000215585Regon011122045NIP5272128691`, where the tag
strip welds a label onto its own value so a `\bNIP\b` boundary can never match. Both are fixed.

**REGON: 13 domains flagged, all luck, none a miss.** A random nine-digit run passes the REGON
checksum about one time in eleven, so the raw sweep count is an **upper bound**, and every hit was
eyeballed. They are phone numbers: `pesa.pl` → `tel 52 586 85 00`, `mokate.com.pl` →
`+48 781 850 334`, `wielton.com.pl` → `+48 789 560 848`. None carries a REGON label.

##### The checksum-gated loose pass

A second, permissive pass allows up to 40 characters between the label and the number, and accepts
a match **only when the checksum passes** — so a loosely-associated number has to earn its place.

On this corpus it added **0 NIPs**, with the pass rate flat at 97%. Reported as measured: it costs
nothing, it is there for the clause-separated forms a wider corpus will contain, and it is *not*
carrying the fill number.

##### Which rung of the ladder the disclosure came from

On the fixture set (frame A only — the two-frame version of this claim is below), 87 of 124
disclosures came from a sub-page. What matters more here is the **per-field** split, because it is
what justifies the crawl rather than just describing it:

| Field | fill | from homepage | from a sub-page |
|---|---|---|---|
| `nip` / `vatNumber` | 98% | 40 | 82 |
| `email` | 89% | 76 | 34 |
| `phone` | 93% | 98 | 17 |
| `registeredName` | 70% | 32 | 55 |
| `legalForm` / `statutoryScope` | 69% | 31 | 55 |
| `addressLine` | 64% | 35 | 44 |
| `regon` | 61% | 21 | 55 |
| `krs` / `registerUrl` | 52% | 19 | 45 |
| `registeredOffice` | 48% | 12 | 47 |
| `krsCourt` | 30% | 7 | 30 |
| `krsCourtDivision` | 28% | 6 | 29 |
| `officerName` | 17% | 4 | 17 |
| `shareCapital` | 15% | 5 | 13 |

`fieldCount`: median **12**, p90 **18**, max **19** of 22.

Read the two ends of that table together. **Contact details are a homepage fact** — the phone comes
off the homepage 98 times out of 115. **Registry identifiers are not** — the registry court comes
off a sub-page 30 times out of 37, the registered office 47 out of 59. An Actor that fetched only
the homepage would return a contact scraper's output and call it company data.

##### Statutory scope of what was found

| Scope | rows |
|---|---|
| | frame A | frame B |
|---|---|---|
| `KSH art. 206 §1` (sp. z o.o.) | 33.3% | 43.4% |
| not stated (no legal form published) | 34.3% | 24.9% |
| `none (no KSH website-disclosure duty)` | 16.7% | 14.2% |
| `KSH art. 374 §1` (S.A.) | 7.9% | 15.1% |
| `KSH art. 127 §5 (via art. 374)` (S.K.A.) | 7.9% | 2.5% |

The mix moves with the frame in exactly the direction it should: as the corpus gets more corporate,
`art. 206 §1` rises and "no legal form published" falls. That is the field doing its job.

#### On the platform — both frames, same build, 1,184 domains

```
                                    frame A (OSM)   frame B (Tranco)
domains attempted                            584                600
BILLED rows                                  216                325
chargedEventCounts                           216                325     delivered == charged, both
```

**What is frame-INDEPENDENT — i.e. a fact about the parser:**

| | frame A | frame B | spread |
|---|---|---|---|
| NIP checksum pass rate | 97.7% | 98.4% | 0.8pt |
| REGON checksum pass rate | 98.1% | 96.2% | 1.9pt |
| `nip` fill | 99.1% | 98.5% | 0.6pt |
| `email` fill | 92.6% | 94.5% | 1.9pt |
| `officerName` fill | 15.7% | 16.6% | 0.9pt |

**What MOVES — i.e. a fact about the list you feed it, quoted as a range:**

| | frame A | frame B | spread |
|---|---|---|---|
| **yield, % of reachable** | **45.5%** | **61.9%** | **16.4pt** |
| yield, % of attempted | 37.0% | 54.2% | 17.2pt |
| unreachable or blocked | 18.7% | 12.5% | 6.2pt |
| `krsCourt` fill | 26.9% | 47.1% | 20.2pt |
| `krs` fill | 48.6% | 66.5% | 17.9pt |
| `registeredName` fill | 69.9% | 84.6% | 14.7pt |
| `registeredOffice` fill | 38.9% | 51.1% | 12.2pt |
| `regon` fill | 71.8% | 81.2% | 9.5pt |
| `shareCapital` fill | 13.4% | 20.9% | 7.5pt |
| `phone` fill | 89.8% | 80.0% | 9.8pt *(the only one that moves the other way — OSM is local businesses, who publish a phone)* |

**So: expect 46-62% of reachable domains to yield a billed row**, depending on what your list is
made of. A traffic-ranked list of Polish commercial sites sits at the top of that band; a list
scraped off a maps product sits at the bottom, because it is full of sole traders, schools and
market stalls that owe no disclosure duty in the first place. That is a real **+36%** between the
two frames — measured, not asserted. A list filtered to *registered companies* should do better
than either, but we have not measured that and so do not claim it.

`statutoryScope` moves with the frame in exactly the way it should: `art. 206 §1` 33.3% → 43.4%
and "no legal form published" 34.3% → 24.9% as the corpus gets more corporate.

#### The claim the multi-page crawl rests on — and it holds on both frames

| | frame A | frame B |
|---|---|---|
| disclosure needed a page past the homepage | **69.0%** | **72.3%** |

**3.3 points apart across two corpora built on opposite principles.** This is the one population
number stable enough to state flatly: **roughly 70% of Polish website disclosures are not on the
homepage.** The rungs that actually paid, both frames:

```
A: /=67 · /kontakt/=27 · /kontakt=20 · /polityka-prywatnosci/=16 · /regulamin/=9 · /regulamin=6
B: /=90 · /kontakt=30  · /kontakt/=27 · /regulamin=18 · /kontakt.html=10 · /polityka-prywatnosci=9
```

An Actor that fetched only the homepage would have returned 67 rows instead of 216 on frame A, and
90 instead of 325 on frame B.

**Checksums on the two runs:** NIP 97.7% / 98.4% pass · REGON 98.1% / 96.2%. **Conditional emit,
measured by re-fetching delivered pages** and re-scanning them with independently written detectors:
NIP 100% · KRS 98.9% · REGON 99.3% · registry court 100% · share capital 100%, with **0
non-deterministic re-parses**.

#### We re-probed the misses, rather than just counting them

The 217 rows marked "publishes no disclosure" are the largest unbilled bucket, so a sample of 40
was re-probed from scratch — homepage plus every guessed path, parsed again:

```
35  genuinely publish nothing
 4  unreachable on the re-probe
 1  found after all  ->  a real ladder gap, on /polityka-prywatnosci/
```

**87.5% of that bucket is genuine.** The one miss ran out of its 10-request per-domain budget before
reaching the page; raising `maxDiscoveryRequestsPerDomain` recovers some of it, and you pay nothing
extra for the ones that still return nothing.

***

### 🧰 Every input

| Input | Default | What it does |
|---|---|---|
| `domains` | demo batch | The list. Bare domains, homepage URLs, direct legal-page URLs or email addresses |
| `startUrls` | — | The same list in Apify request-list format |
| `domainsFileUrl` | — | A public CSV / TSV / TXT / JSON / JSONL URL |
| `sourceDatasetId` + `domainFieldName` | — | Read the domains out of another Actor's dataset |
| `skipDomains` | — | Never fetched, never charged |
| `previousDatasetId` | — | Suppress everything an earlier run already delivered |
| `maxItems` | 1000 | **Billing cap.** This × the row price is your maximum spend |
| `maxDomains` | 0 | Stop after reading this many input rows |
| `maxPagesParsed` | 4 | Pages of one site that may be parsed |
| `maxDiscoveryRequestsPerDomain` | 10 | Request budget per domain |
| `perDomainTimeoutSecs` | 60 | Hard per-domain deadline |
| `requestConcurrency` | 10 | Auto-clamped to the run's memory (~80 MB per parallel domain) |
| `requestTimeoutSecs` / `maxRequestRetries` | 25 / 2 | Each retry on a fresh proxy IP |
| `discoveryChannels` | all five | `homepage`, `anchor`, `pathGuess`, `sitemap`, `wpJson` |
| `deepJsDiscovery` | false | Read inline JSON and JS bundles for a client-rendered footer. No browser |
| `followWwwAndRootVariants` | true | Retry with/without `www.` |
| `proxyConfiguration` / `proxyCountry` | Apify datacenter / none | Country pinning forces residential — see the input hint |
| `escalateToResidentialOnBlock` | true | One escalation on a genuine block |
| `escalateToUnblockerOnBlock` | false | Second escalation for the anti-bot tail |
| `customUserAgent` / `extraHttpHeaders` | — | |
| `respectRobotsTxt` | false | See *Legal & fair use* |
| `validateTaxId` | true | Run the NIP and REGON checksums |
| `extractRegon` / `extractCourt` / `extractShareCapital` / `extractOfficer` / `extractSocials` / `extractPolicyUrls` | true | Per-field switches |
| `emailPolicy` | all | `all` · `exclude-role` (drop `biuro@`, `kontakt@`, …) · `role-only` |
| `requireRegistryId` / `requireKrs` / `requireValidNip` / `requireContact` | false | Quality gates — **filtered rows are never charged** |
| `legalFormFilter` | — | e.g. only `sp. z o.o.` and `S.A.` |
| `minShareCapital` | 0 | Filter on the published kapitał zakładowy |
| `minFieldsRequired` | 0 | Minimum populated value fields |
| `tldFilterMode` / `tldFilter` | none | Opt-in suffix filter |
| `dedupeBy` | `krs-then-domain` | Deduplication happens **before** billing |
| `includeMissRows` | false | Deliver an unbilled row for every domain that produced nothing |
| `includeFieldSources` | false | Which page each field came from |
| `flattenOutput` | true | Flat row (CSV-friendly) vs nested objects |

***

### 💷 Pricing

**Pay per event: one charge per company disclosure returned.** No monthly fee. **No start fee.**

You are charged **only** for a delivered row carrying at least one registry identifier (KRS, NIP or
REGON) or a named registry court. Delivery and billing are atomic — `pushData(item, event)` — so
you can never be charged for a row you did not receive, and never receive one you were not charged
for.

**Never charged:**

- a host that does not resolve or refuses the connection
- a host that answers with a 403 or a challenge page
- a domain that publishes **no** disclosure — *including every sole trader, who has no duty to*
- a duplicate, collapsed before billing
- a row removed by **your own** quality filters
- an unbilled coverage row (`includeMissRows`)

`RUN_SUMMARY` in the key-value store keeps those apart — `unreachable_host`,
`blocked_or_challenged`, `publishes_no_disclosure`, `filtered_out_by_your_settings`, `duplicate`,
`suppressed`, `skipped_by_robots`, `error` — so you can reconcile your whole input list against
your invoice.

Set **`maxItems`** to cap a run's cost absolutely.

***

### ❓ FAQ

**Do I need a KRS number to use this?**
No — and that is the point. Every other Polish company-data product on the store is keyed by a NIP
or a KRS you must already have. This one starts from a **domain**, which is what a lead list
actually contains.

**How is this different from scraping the KRS register?**
The register tells you about a company you have already identified. This tells you **which company
a website belongs to**, and adds the email and phone the register does not hold. Use both: the
`registerUrl` on every row is a direct link into the Ministry of Justice's own KRS API for the full
extract.

**Why is the NIP fill so much higher than the KRS fill?**
Because a NIP is what a Polish business puts in its footer for invoicing, while the KRS is what the
*statute* requires — and only from companies the statute binds. Sole traders have a NIP and no KRS
at all. `statutoryScope` tells you which case each row is.

**Is `nipValid: false` a bug?**
No. It means the company published a NIP that fails the official checksum — a typo in their own
footer, most often. You are seeing the site as it is. Use `requireValidNip` to drop those rows.

**Why is there no KRS checksum?**
Because the KRS does not have one. See *Honest limits*.

**Can I run this monthly without paying twice for the same companies?**
Yes. Pass the previous run's dataset id as `previousDatasetId` and every KRS, NIP and domain it
delivered is suppressed before a single request is made.

**Does it use a browser?**
No. `got-scraping` + `cheerio`, which is what keeps a run cheap. `deepJsDiscovery` reads inline
JSON and JavaScript bundles for a client-rendered footer without launching one.

**What if a site blocks the Apify proxy?**
A genuine block escalates to a residential IP automatically, and `escalateToUnblockerOnBlock` adds
a second escalation. You can also supply your own `proxyConfiguration`. Blocked domains are
reported in `RUN_SUMMARY` and are never charged.

***

### ⚠️ Honest limits

- **The KRS carries no checksum.** `krsFormatValid` is a **format and range check** — ten digits,
  not all zeros — and nothing more. Claiming a checksum here would be a lie you could test in five
  minutes. The NIP and REGON checks *are* real checksums, and they are the ones to filter on.
- **The duty binds only some legal forms.** A sole trader, a *sp.j.*, a *sp.k.* and a *s.c.* have
  no KSH website-disclosure obligation. On a mixed domain list, a large share of the misses are
  correct output. `statutoryScope` is how you tell. Feed it limited-company domains.
- **A town read out of a `z siedzibą w …` clause is in the Polish locative case**, exactly as the
  company printed it — `Bytowie` for Bytów, `Krakowie` for Kraków. Reversing a Polish declension
  would be a guess, and a wrong guess ships as a *populated wrong field*, which is worse than a
  published one. The **postcode is canonical** either way; filter on that.
- **`registeredOffice` is sometimes postcode + town only**, when the company publishes no street in
  its statutory clause. That is what the page says.
- **`officerName` fill is low (16-17%, stable across both frames).** Polish companies publish the *zarząd* in the KRS, not
  usually on the website. Where a site does name one, it is captured; where it does not, the field
  is null rather than inferred.
- **`shareCapital` fill is 13-21% depending on the list.** The art. 206 §1 pt 4 duty is widely honoured on the *regulamin*
  page and widely ignored in footers. This is compliance reality, not a parser gap — the
  conditional emit rate where the label *is* present is 95% offline and 100% on the live re-fetch. Note the denominator there is only 19 pages, which is too small to be evidence on its own - the validator prints it as not-enforced rather than reporting a flattering percentage.
- **The REGON negative-direction sweep flags ~10% of rows.** Every one inspected was a phone number
  passing a nine-digit checksum by luck. Stated so you can reproduce the check rather than take the
  0-misses claim on trust.
- **A site can be wrong about itself.** Everything here is what the company published. Where it
  publishes a stale address or a mistyped NIP, that is what you get — with `nipValid: false`
  telling you so.
- **Rating and last-updated of the source pages are not knowable** from the bytes; if a company
  has not touched its *regulamin* since 2019, its disclosure is 2019's.

***

### 🧾 Legal & fair use

These are **statutory public disclosures** that Polish companies are required by the Kodeks spółek
handlowych to publish on their own websites. The Actor reads only pages a company publishes openly;
it never logs in, never bypasses a paywall, and never touches personal data behind an account.

`respectRobotsTxt` is **off by default** and available as an input, because some buyers' own
compliance policy asks for it. Skipped pages are reported and never charged.

You are responsible for complying with each site's Terms of Service and with **GDPR / RODO** for
any personal data in the output — a named *prezes zarządu*, a personal email address. Fields that
can carry personal data (`officerName`, `officerRole`) can be switched off with `extractOfficer`.

***

*Part of a family of statutory website-disclosure scrapers: 🇩🇪 Impressum · 🇬🇧 UK trading
disclosures · 🇳🇱 KvK · 🇧🇪 ondernemingsnummer · 🇮🇹 note legali · 🇪🇸 aviso legal · 🇫🇷 mentions
légales · 🇳🇴 Brønnøysund.*

# Actor input Schema

## `domains` (type: `array`):

Polish company domains to process. A bare domain, a homepage URL or a direct link to a /kontakt, /regulamin or /polityka-prywatnosci page all work. An email address is accepted too - the domain is taken from it. Leave this empty to run the built-in 10-domain demo batch.

## `startUrls` (type: `array`):

The same list in Apify's request-list format, so Make, Zapier, Clay or a Google Sheet can hand over a URL list natively.

## `domainsFileUrl` (type: `string`):

A public URL holding the domain list. CSV, TSV, one-per-line TXT, JSON array and JSONL are all read; a header row naming domain / website / url is detected automatically.

## `sourceDatasetId` (type: `string`):

Read the domains out of an existing Apify dataset - the output of a Maps, directory or register scraper, for example.

## `domainFieldName` (type: `string`):

Which field of the source dataset or CSV holds the website. Leave empty to auto-detect domain / website / websiteUrl / url / site / homepage.

## `skipDomains` (type: `array`):

Domains to drop before any request is made. They are never fetched and never charged.

## `previousDatasetId` (type: `string`):

A dataset from an earlier run. Every KRS, NIP and domain it already delivered is suppressed, so a monthly re-run never re-buys a row you already paid for.

## `maxItems` (type: `integer`):

Stop after this many BILLED rows. This is the hard ceiling on what a run can cost you: rows that are not delivered are not charged, so this number times the per-row price is your maximum spend. 0 means no cap.

## `maxDomains` (type: `integer`):

Stop after reading this many domains from the input, whether or not they yielded a disclosure. 0 means all of them.

## `maxPagesParsed` (type: `integer`):

How many pages of one site may be parsed before giving up. A Polish disclosure is usually split across two or three pages (the KRS on /regulamin, the email on /kontakt), which is why the default is 4 rather than 1.

## `maxDiscoveryRequestsPerDomain` (type: `integer`):

Budget of HTTP requests spent looking for the disclosure on one site, homepage included.

## `perDomainTimeoutSecs` (type: `integer`):

Hard deadline for one domain. When it expires the domain is abandoned, reported as unreachable and never charged.

## `requestConcurrency` (type: `integer`):

How many domains to process in parallel. Each parallel domain needs roughly 80 MB, so this is automatically lowered to fit the memory the run was given - raise the Actor's memory if you want more parallelism.

## `requestTimeoutSecs` (type: `integer`):

Timeout for one HTTP request.

## `maxRequestRetries` (type: `integer`):

How many times a failed request is retried, each time on a fresh proxy IP.

## `discoveryChannels` (type: `array`):

Which channels to use, in order. homepage = the site-wide footer; anchor = the site's own links ranked by text and href (kontakt, regulamin, polityka prywatnosci, o nas, dane spolki); pathGuess = the conventional Polish paths; sitemap = the XML sitemap; wpJson = the WordPress page index.

## `deepJsDiscovery` (type: `boolean`):

For sites whose footer only appears after a client-side render: read the page's inline JSON payloads and JavaScript bundles for the legal-page URL. No browser is launched. Slower, off by default.

## `followWwwAndRootVariants` (type: `boolean`):

If the homepage does not answer, try the other of example.pl and www.example.pl before giving up.

## `proxyConfiguration` (type: `object`):

Apify Proxy, or your own. The datacenter pool is the measured-fastest default; a genuine block escalates to residential automatically.

## `proxyCountry` (type: `string`):

Pin the exit node to a country. Apify can only pin RESIDENTIAL proxies to a country, so choosing anything other than None switches the run to residential - slower and dearer. Polish sites do not require a Polish IP.

## `escalateToResidentialOnBlock` (type: `boolean`):

When a site refuses the datacenter IP with a 403 or a challenge page, retry once through a residential IP.

## `escalateToUnblockerOnBlock` (type: `boolean`):

Second escalation for the anti-bot tail, using Apify Unblocker. Costs more proxy credit; off by default.

## `customUserAgent` (type: `string`):

Override the browser User-Agent sent with every request.

## `extraHttpHeaders` (type: `object`):

Extra headers merged into every request.

## `respectRobotsTxt` (type: `boolean`):

Skip pages the site's robots.txt disallows for this User-Agent. Off by default: a statutory disclosure the company is legally required to publish about itself is not private, and skipped pages are reported and never charged.

## `validateTaxId` (type: `boolean`):

Run the NIP weighted checksum (weights 6,5,7,2,3,4,5,6,7 mod 11) and the REGON checksum on every number found, and report the result per row. Turning this off leaves nipValid and regonValid null rather than making them up.

## `extractRegon` (type: `boolean`):

Read the REGON statistical number (9-digit entity or 14-digit local unit) and check its own checksum.

## `extractCourt` (type: `boolean`):

Read the sad rejestrowy and its commercial division - the art. 206 §1 pt 2 element ('Sad Rejonowy dla Wroclawia-Fabrycznej we Wroclawiu, VI Wydzial Gospodarczy').

## `extractShareCapital` (type: `boolean`):

Read the kapital zakladowy as a number plus its currency, and whether it is stated as paid up in full - the art. 206 §1 pt 4 / art. 374 §1 pt 4 element.

## `extractOfficer` (type: `boolean`):

Read a named prezes zarzadu, wlasciciel, dyrektor or inspektor ochrony danych when the page states one. Fill is low: Polish companies publish the zarzad in the KRS, not usually on the website.

## `extractSocials` (type: `boolean`):

Collect LinkedIn, Facebook, Instagram, X, YouTube and TikTok profile links.

## `extractPolicyUrls` (type: `boolean`):

Record the URLs of the site's regulamin and polityka prywatnosci pages.

## `emailPolicy` (type: `string`):

Agencies disagree on whether biuro@ and kontakt@ count as leads, so it is your call. Role mailboxes are the Polish generic ones: biuro, kontakt, sekretariat, sprzedaz, rekrutacja, rodo and the rest.

## `requireRegistryId` (type: `boolean`):

Deliver only companies publishing a KRS, NIP or REGON. Rows removed by this filter are never charged.

## `requireKrs` (type: `boolean`):

Deliver only companies publishing a KRS entry number - in practice, only registered sp. z o.o. and S.A. companies.

## `requireValidNip` (type: `boolean`):

Deliver only rows whose NIP passes the weighted checksum. A published NIP that fails the checksum is a fact about the site, so this is off by default.

## `requireContact` (type: `boolean`):

Deliver only rows carrying at least one way to reach the company.

## `legalFormFilter` (type: `array`):

Keep only the listed legal forms. Leave empty for all of them.

## `minShareCapital` (type: `integer`):

Keep only companies whose published kapital zakladowy is at least this much. Rows that publish no share capital are removed by any value above 0.

## `minFieldsRequired` (type: `integer`):

Keep only rows carrying at least this many of the 22 value fields.

## `tldFilterMode` (type: `string`):

Off by default on purpose: many real Polish companies, and disproportionately the larger ones, publish on .com or .eu rather than .pl, so filtering to .pl silently discards them.

## `tldFilter` (type: `array`):

The suffixes the filter above applies to, for example pl, com.pl, eu.

## `dedupeBy` (type: `string`):

How repeated companies are collapsed. Deduplication happens BEFORE billing, so a duplicate can never be charged twice.

## `includeMissRows` (type: `boolean`):

Also deliver a row for every domain that produced nothing, with the reason kept apart: unreachable host, refused by the site, publishes no disclosure, removed by your filters, duplicate or suppressed. These rows are NEVER charged - they exist so you can account for your whole input list.

## `includeFieldSources` (type: `boolean`):

Add a fieldSources object giving the page path each field was read from, so you can see that the KRS came from /regulamin and the email from /kontakt.

## `flattenOutput` (type: `boolean`):

One flat row per company, which is what a CSV export and most CRMs want. Turn it off for a nested company / registry / contact / people / policies object.

## Actor input object example

```json
{
  "domains": [
    "selena.pl",
    "nowystylgroup.com",
    "kross.pl",
    "pesa.pl",
    "delia.pl",
    "drutex.pl",
    "cersanit.com",
    "evital.pl",
    "amica.pl",
    "mokate.com.pl"
  ],
  "maxItems": 1000,
  "maxDomains": 0,
  "maxPagesParsed": 4,
  "maxDiscoveryRequestsPerDomain": 10,
  "perDomainTimeoutSecs": 60,
  "requestConcurrency": 10,
  "requestTimeoutSecs": 25,
  "maxRequestRetries": 2,
  "discoveryChannels": [
    "homepage",
    "anchor",
    "pathGuess",
    "sitemap",
    "wpJson"
  ],
  "deepJsDiscovery": false,
  "followWwwAndRootVariants": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "proxyCountry": "none",
  "escalateToResidentialOnBlock": true,
  "escalateToUnblockerOnBlock": false,
  "respectRobotsTxt": false,
  "validateTaxId": true,
  "extractRegon": true,
  "extractCourt": true,
  "extractShareCapital": true,
  "extractOfficer": true,
  "extractSocials": true,
  "extractPolicyUrls": true,
  "emailPolicy": "all",
  "requireRegistryId": false,
  "requireKrs": false,
  "requireValidNip": false,
  "requireContact": false,
  "minShareCapital": 0,
  "minFieldsRequired": 0,
  "tldFilterMode": "none",
  "dedupeBy": "krs-then-domain",
  "includeMissRows": false,
  "includeFieldSources": false,
  "flattenOutput": true
}
```

# Actor output Schema

## `records` (type: `string`):

The dataset of Polish company leads (one item per domain that published a KSH art. 206 / art. 374 website disclosure).

## `runSummary` (type: `string`):

Coverage accounting for the run: domains attempted, rows billed, the ladder rung each disclosure came from, the statutory-scope distribution, the NIP and REGON checksum results, per-field fill, and the three kinds of nothing (unreachable host, blocked host, publishes no disclosure) kept apart - none of them charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "selena.pl",
        "nowystylgroup.com",
        "kross.pl",
        "pesa.pl",
        "delia.pl",
        "drutex.pl",
        "cersanit.com",
        "evital.pl",
        "amica.pl",
        "mokate.com.pl"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/pl-dane-spolki-website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "selena.pl",
        "nowystylgroup.com",
        "kross.pl",
        "pesa.pl",
        "delia.pl",
        "drutex.pl",
        "cersanit.com",
        "evital.pl",
        "amica.pl",
        "mokate.com.pl",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/pl-dane-spolki-website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "selena.pl",
    "nowystylgroup.com",
    "kross.pl",
    "pesa.pl",
    "delia.pl",
    "drutex.pl",
    "cersanit.com",
    "evital.pl",
    "amica.pl",
    "mokate.com.pl"
  ]
}' |
apify call scrapersdelight/pl-dane-spolki-website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/pl-dane-spolki-website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/AIOMy1Zb3eQtcM3OO/builds/2tCDukrZxHgcYZuIN/openapi.json
