Polish Company Data Scraper - KRS, NIP, REGON & Contacts
Pricing
$7.00 / 1,000 per company disclosure returneds
Polish Company Data Scraper - KRS, NIP, REGON & Contacts
Turn Polish company domains into registry-grade B2B leads from each site's statutory disclosure (KSH art. 206 / art. 374): firma, KRS number and registry court, checksum-validated NIP and REGON, share capital, registered office, email and phone. $0.007 per disclosure. No login, no browser.
Pricing
$7.00 / 1,000 per company disclosure returneds
Rating
0.0
(0)
Developer
Scrapers Delight
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
🇵🇱 Polish Company Data Scraper — KRS, NIP, REGON, Share Capital & Contacts
Give it a list of Polish company domains. Get back, for each one, the statutory disclosure the company is legally required to publish on its own website: the registered firma, the KRS entry number and the sąd rejestrowy that holds the file, a checksum-validated NIP, a checksum-validated REGON, the kapitał zakładowy as a number, the registered office, email and phone.
No login. No API key. No browser. Plain HTTP through the Apify proxy.
You are never charged for a domain that is dead, blocked, or publishes no disclosure.
⚖️ Why this data exists, and why it is clean
Kodeks spółek handlowych art. 206 §1 requires every spółka z ograniczoną odpowiedzialnością,
and art. 374 §1 every spółka akcyjna, to state — in its commercial letters and orders and
na stronach internetowych spółki —
- the firma, the siedziba and the adres;
- the sąd rejestrowy holding the company's file, and the number under which it is entered in the register (the KRS);
- the NIP;
- the wysokość kapitału zakładowego — and, for an S.A. under art. 374 §1 pt 4, the capital actually paid up.
Art. 127 §5 carries the same duty to a spółka komandytowo-akcyjna, and art. 300(101) to a prosta spółka akcyjna.
So this is not scraped inference and it is not a directory's copy of a company's details. It is a disclosure the company itself is obliged by statute to publish about itself, on its own site, and to keep accurate.
The honest half of that: the duty does not bind everyone
The KSH website duty binds sp. z o.o., S.A., S.K.A. and P.S.A. — and nobody else. A jednoosobowa działalność gospodarcza (a sole trader, and the commonest legal form in Poland by a wide margin), a spółka jawna, a spółka komandytowa and a spółka cywilna have no such obligation at all.
That is why every row carries statutoryScope, naming the provision that binds that company —
KSH art. 206 §1, KSH art. 374 §1, KSH art. 127 §5 (via art. 374) — or
none (no KSH website-disclosure duty). A domain with no KRS is very often correct output
rather than a parser failure, and this field is how you tell the two apart.
Feed it business domains and the hit rate rises by about a third — measured. Across two corpora built on opposite principles, the yield went from 45.5% of reachable domains on a maps/POI-sourced list to 61.9% on a traffic-ranked one: +36%, not the doubling we expected before measuring it. A maps-sourced list is full of sole traders, schools and market stalls who owe you nothing. A list of sp. z o.o. and S.A. websites — from a KRS export, a trade-association member list, a tender register, an exhibitor list — should do better than either, but we have not measured that one and so do not claim a number for it.
🔍 What it does, per domain
Two to five hops, plain HTTP (got-scraping + cheerio, no browser), Apify datacenter
proxy with a residential and an optional Unblocker escalation.
-
Fetch the homepage and read its footer.
-
Climb the ladder. This is the part that matters in Poland. Measured over real Polish business domains, 69-72% of the disclosures found needed a page past the homepage — measured on two independent corpora that agree to within 3.3 points — the opposite of the UK, where the footer usually carries it, and close to the Dutch pattern. So the Actor ranks the site's own links by anchor text and href and walks them in the order Polish sites actually use:
dane spółki›kontakt›regulamin›nota prawna›polityka prywatności›o nasthen falls back to guessed conventional paths (
/kontakt,/regulamin,/o-nas,/dane-spolki,/polityka-prywatnosci, …), the XML sitemap, and the WordPress page index.The polityka prywatności page is a first-class source here, not a fallback: a Polish RODO clause names the administrator's full firma, siedziba, KRS and NIP in one sentence.
-
Merge field by field, recording which page each field came from. The KRS on
/regulamin, the email on/kontaktand the share capital on/o-nasis an ordinary Polish site.
Every statutory field is read inside a statement window anchored on a disclosure phrase, never by a page-wide scan.
📦 What you get — every field
Identity
| Field | What it is |
|---|---|
registeredName | The firma as published — Żabka Polska sp. z o.o., PESA Bydgoszcz SA |
registeredNameSource | How it was resolved: label, suffix, clause |
tradingName | A brand or trading name stated as distinct from the firma |
companyName | Best available name (firma, else JSON-LD legalName, else og:site_name) |
legalForm | sp. z o.o. · S.A. · P.S.A. · S.K.A. · sp. z o.o. sp.k. · sp.k. · sp.j. · sp.p. · s.c. · fundacja · stowarzyszenie · spółdzielnia · jednoosobowa działalność gospodarcza |
statutoryScope | Which KSH provision binds this company's website, or that none does |
Registry
| Field | What it is |
|---|---|
krs | The KRS entry number, canonical 10 digits, zero-padded |
krsRaw | Exactly as the page printed it |
krsSource | clause (from the long statutory sentence), label, or bare |
krsFormatValid | Format check only. The KRS carries no checksum — see Honest limits |
krsScheme | PL-KRS |
registerUrl | Direct deep link into the Ministry of Justice's own public KRS API, which returns the current extract (odpis aktualny) as JSON |
krsCourt | The sąd rejestrowy — Sąd Rejonowy dla Wrocławia-Fabrycznej we Wrocławiu |
krsCourtDivision | Its commercial division — VI Wydział Gospodarczy |
nip | The NIP, 10 digits |
nipValid | Weighted checksum result (weights 6,5,7,2,3,4,5,6,7 mod 11; a remainder of 10 is invalid) |
nipFormatted | 527-212-86-91 |
nipSource | label, prefix (a PL-prefixed form), json-ld, or loose-checksummed |
vatNumber | The intra-EU VAT identifier — PL + the NIP, ready for VIES |
regon | The REGON, 9 digits (entity) or 14 (a local unit of one) |
regonValid | Its own checksum, a different rule from the NIP's |
regonLength | 9 or 14, so you can tell an entity from a branch |
shareCapital | Kapitał zakładowy as a NUMBER you can filter and sort on |
shareCapitalCurrency | PLN (or EUR/USD where a company states one) |
shareCapitalRaw | The string it was parsed from, so you can check the parse yourself |
shareCapitalPaid | The capital stated as paid up |
shareCapitalPaidInFull | true when the page says wpłacony w całości; null when it says nothing — never false |
registeredOffice | The siedziba as a postal line — ul. Legnicka 48A, 54-202 Wrocław |
registeredOfficePostcode | 54-202 |
registeredOfficeCity | Wrocław |
Contact
| Field | What it is |
|---|---|
email | Best email — on-domain preferred |
emails | Up to 10, deduped, Cloudflare [email protected] decoded |
emailConflict | Set when the primary email is on a different domain from the site and is not a freemail provider — a signal the site is run by an agency |
phone | Best phone, E.164 +48… |
phones | Up to 10 |
addressLine / postcode / city | A contact/trading address where the page gives one distinct from the siedziba |
officerName / officerRole | A named prezes zarządu, właściciel, dyrektor or inspektor ochrony danych where the site states one |
socialLinks | LinkedIn, Facebook, Instagram, X, YouTube, TikTok profiles |
termsUrl / privacyPolicyUrl | The site's regulamin and polityka prywatności |
Provenance & accounting
| Field | What it is |
|---|---|
domain / inputUrl / resolvedUrl | What you gave us and what answered |
disclosureUrl | The exact page the disclosure was read from |
disclosurePagePath | Which rung of the ladder that was — /, /kontakt, /regulamin, … |
disclosureSource | homepage, legal-page or merged (split across pages) |
discoveryChannel | homepage, anchor, guess, sitemap, wpjson, jsassets |
pagesParsed / pagesChecked | What it cost us |
fieldSources | (optional) the page path each individual field came from |
fieldCount | How many of the 22 value fields are populated |
stableId | The KRS when well-formed, else NIP…, else the registrable domain |
status / missReason | ok, or exactly why not |
elapsedMs / fetchedAt |
📊 Measured results — real numbers, not claims
Two corpora, on purpose
One corpus measures a parser. It does not measure a population. Whether a label yields a value is a fact about this Actor; what share of Polish companies publish a KRS at all is a fact about the list you feed it. So every population number below was measured twice, on two corpora built in completely different ways, with biases that point in opposite directions:
| Frame A | Frame B | |
|---|---|---|
| source | OpenStreetMap website tags | the Tranco .pl tail (ranked past 100,000) |
| sampling | someone walked past it and mapped it | traffic |
| built from | office + shop + craft + amenity + healthcare + tourism across 30 Polish cities | a 1M-row popularity list, .pl only |
| size | 16,400 unique hosts | 7,994 candidates, 4,000 sampled |
| skews towards | local shops, workshops, schools, sole traders — who owe no KSH duty at all | larger companies — disproportionately sp. z o.o. and S.A., who do |
They bracket the population instead of agreeing by construction. A number that holds across both is a fact about Poland; a number that moves is a fact about the sampling frame, and is quoted below as a range.
Frame A's suffix split — .pl 81.1%, .com.pl 6.6%, .com 5.8%, .eu 3.5%, rest 1.8%; .pl
family 87.6%, everything else 12.4%. No TLD filter is applied, deliberately: the non-.pl eighth
skews towards the larger companies (cersanit.com, asseco.com, maspex.com, solarisbus.com all
publish a full KRS statement), so filtering to .pl would silently discard them.
Frame A was also checked for cross-border contamination, because a country-wide Overpass area query can silently pull in a neighbour's businesses. It cannot here — the query is 30 city-sized bounding boxes, none of which reaches a border. Confirmed empirically: 35 foreign ccTLD hosts in 16,400 (0.21%), spread diffusely over 15 countries, and of the 13 on the six land-neighbour ccTLDs, 7 are diplomatic missions physically located in Polish cities and the rest are international brands with a Polish branch.
Offline validation, against real captured bytes
Every fixture was captured live, through the shipping proxy, and the shipping parser is run
over it by _validate.mjs, which asserts each advertised field against the raw bytes it came from:
FIXTURE FILES: 547 DOMAINS: 237 WITH A DISCLOSURE: 124ASSERTIONS: 5506 FAILURES: 0
plus a separate 165-assertion unit suite over the checksums, the number formats, the address assembler and the legal-form map — every case a string copied out of a named real page.
Checksums, each cross-checked against a second, independently written implementation
| valid | published but invalid | not published | pass rate of those published | |
|---|---|---|---|---|
| NIP | 118 | 4 | 2 | 97% |
| REGON | 76 | 0 | 48 | 100% |
| KRS | — | — | — | format check only — the KRS has no checksum |
The validator's checksums are written from the published rule in a different shape from the shipping ones (a ten-term dot product versus an indexed nine-weight loop), so a bug in either cannot validate itself. Each is also put through a single-digit mutation sweep: every one-digit corruption of a real published number must fail.
A NIP that is published but fails its checksum is emitted with nipValid: false. That is a fact
about the site, not a parser failure, and hiding it would be worse than reporting it.
Conditional emit rate — of the pages that carry the label, what share yielded a value
This is the measurement that tells a sparse field apart from a broken regex. Raw fill cannot.
| Label present on | domains | value emitted | conditional rate |
|---|---|---|---|
| NIP | 118 | 117 | 99% |
| KRS | 64 | 60 | 94% |
| REGON | 79 | 76 | 96% |
| registry court | 35 | 35 | 100% |
| share capital | 19 | 18 | 95% |
Negative direction — did we miss something the bytes plainly contain?
An independently written scan re-reads every page where the parser emitted null and looks for a
checksum-valid number it should have caught.
NIP: 0 misses. It was not 0 to begin with. The sweep found exactly one — and it earned its
keep, because the cause was general: wtzdeotymy.ksnaw.pl publishes NIP 527- 21- 28 -691, a
hyphen and a space between the same pair of digits, which a one-character separator class
cannot cross. The same sweep also found KRS0000215585Regon011122045NIP5272128691, where the tag
strip welds a label onto its own value so a \bNIP\b boundary can never match. Both are fixed.
REGON: 13 domains flagged, all luck, none a miss. A random nine-digit run passes the REGON
checksum about one time in eleven, so the raw sweep count is an upper bound, and every hit was
eyeballed. They are phone numbers: pesa.pl → tel 52 586 85 00, mokate.com.pl →
+48 781 850 334, wielton.com.pl → +48 789 560 848. None carries a REGON label.
The checksum-gated loose pass
A second, permissive pass allows up to 40 characters between the label and the number, and accepts a match only when the checksum passes — so a loosely-associated number has to earn its place.
On this corpus it added 0 NIPs, with the pass rate flat at 97%. Reported as measured: it costs nothing, it is there for the clause-separated forms a wider corpus will contain, and it is not carrying the fill number.
Which rung of the ladder the disclosure came from
On the fixture set (frame A only — the two-frame version of this claim is below), 87 of 124 disclosures came from a sub-page. What matters more here is the per-field split, because it is what justifies the crawl rather than just describing it:
| Field | fill | from homepage | from a sub-page |
|---|---|---|---|
nip / vatNumber | 98% | 40 | 82 |
email | 89% | 76 | 34 |
phone | 93% | 98 | 17 |
registeredName | 70% | 32 | 55 |
legalForm / statutoryScope | 69% | 31 | 55 |
addressLine | 64% | 35 | 44 |
regon | 61% | 21 | 55 |
krs / registerUrl | 52% | 19 | 45 |
registeredOffice | 48% | 12 | 47 |
krsCourt | 30% | 7 | 30 |
krsCourtDivision | 28% | 6 | 29 |
officerName | 17% | 4 | 17 |
shareCapital | 15% | 5 | 13 |
fieldCount: median 12, p90 18, max 19 of 22.
Read the two ends of that table together. Contact details are a homepage fact — the phone comes off the homepage 98 times out of 115. Registry identifiers are not — the registry court comes off a sub-page 30 times out of 37, the registered office 47 out of 59. An Actor that fetched only the homepage would return a contact scraper's output and call it company data.
Statutory scope of what was found
| Scope | rows |
|---|---|
| frame A | |
| --- | --- |
KSH art. 206 §1 (sp. z o.o.) | 33.3% |
| not stated (no legal form published) | 34.3% |
none (no KSH website-disclosure duty) | 16.7% |
KSH art. 374 §1 (S.A.) | 7.9% |
KSH art. 127 §5 (via art. 374) (S.K.A.) | 7.9% |
The mix moves with the frame in exactly the direction it should: as the corpus gets more corporate,
art. 206 §1 rises and "no legal form published" falls. That is the field doing its job.
On the platform — both frames, same build, 1,184 domains
frame A (OSM) frame B (Tranco)domains attempted 584 600BILLED rows 216 325chargedEventCounts 216 325 delivered == charged, both
What is frame-INDEPENDENT — i.e. a fact about the parser:
| frame A | frame B | spread | |
|---|---|---|---|
| NIP checksum pass rate | 97.7% | 98.4% | 0.8pt |
| REGON checksum pass rate | 98.1% | 96.2% | 1.9pt |
nip fill | 99.1% | 98.5% | 0.6pt |
email fill | 92.6% | 94.5% | 1.9pt |
officerName fill | 15.7% | 16.6% | 0.9pt |
What MOVES — i.e. a fact about the list you feed it, quoted as a range:
| frame A | frame B | spread | |
|---|---|---|---|
| yield, % of reachable | 45.5% | 61.9% | 16.4pt |
| yield, % of attempted | 37.0% | 54.2% | 17.2pt |
| unreachable or blocked | 18.7% | 12.5% | 6.2pt |
krsCourt fill | 26.9% | 47.1% | 20.2pt |
krs fill | 48.6% | 66.5% | 17.9pt |
registeredName fill | 69.9% | 84.6% | 14.7pt |
registeredOffice fill | 38.9% | 51.1% | 12.2pt |
regon fill | 71.8% | 81.2% | 9.5pt |
shareCapital fill | 13.4% | 20.9% | 7.5pt |
phone fill | 89.8% | 80.0% | 9.8pt (the only one that moves the other way — OSM is local businesses, who publish a phone) |
So: expect 46-62% of reachable domains to yield a billed row, depending on what your list is made of. A traffic-ranked list of Polish commercial sites sits at the top of that band; a list scraped off a maps product sits at the bottom, because it is full of sole traders, schools and market stalls that owe no disclosure duty in the first place. That is a real +36% between the two frames — measured, not asserted. A list filtered to registered companies should do better than either, but we have not measured that and so do not claim it.
statutoryScope moves with the frame in exactly the way it should: art. 206 §1 33.3% → 43.4%
and "no legal form published" 34.3% → 24.9% as the corpus gets more corporate.
The claim the multi-page crawl rests on — and it holds on both frames
| frame A | frame B | |
|---|---|---|
| disclosure needed a page past the homepage | 69.0% | 72.3% |
3.3 points apart across two corpora built on opposite principles. This is the one population number stable enough to state flatly: roughly 70% of Polish website disclosures are not on the homepage. The rungs that actually paid, both frames:
A: /=67 · /kontakt/=27 · /kontakt=20 · /polityka-prywatnosci/=16 · /regulamin/=9 · /regulamin=6B: /=90 · /kontakt=30 · /kontakt/=27 · /regulamin=18 · /kontakt.html=10 · /polityka-prywatnosci=9
An Actor that fetched only the homepage would have returned 67 rows instead of 216 on frame A, and 90 instead of 325 on frame B.
Checksums on the two runs: NIP 97.7% / 98.4% pass · REGON 98.1% / 96.2%. Conditional emit, measured by re-fetching delivered pages and re-scanning them with independently written detectors: NIP 100% · KRS 98.9% · REGON 99.3% · registry court 100% · share capital 100%, with 0 non-deterministic re-parses.
We re-probed the misses, rather than just counting them
The 217 rows marked "publishes no disclosure" are the largest unbilled bucket, so a sample of 40 was re-probed from scratch — homepage plus every guessed path, parsed again:
35 genuinely publish nothing4 unreachable on the re-probe1 found after all -> a real ladder gap, on /polityka-prywatnosci/
87.5% of that bucket is genuine. The one miss ran out of its 10-request per-domain budget before
reaching the page; raising maxDiscoveryRequestsPerDomain recovers some of it, and you pay nothing
extra for the ones that still return nothing.
🧰 Every input
| Input | Default | What it does |
|---|---|---|
domains | demo batch | The list. Bare domains, homepage URLs, direct legal-page URLs or email addresses |
startUrls | — | The same list in Apify request-list format |
domainsFileUrl | — | A public CSV / TSV / TXT / JSON / JSONL URL |
sourceDatasetId + domainFieldName | — | Read the domains out of another Actor's dataset |
skipDomains | — | Never fetched, never charged |
previousDatasetId | — | Suppress everything an earlier run already delivered |
maxItems | 1000 | Billing cap. This × the row price is your maximum spend |
maxDomains | 0 | Stop after reading this many input rows |
maxPagesParsed | 4 | Pages of one site that may be parsed |
maxDiscoveryRequestsPerDomain | 10 | Request budget per domain |
perDomainTimeoutSecs | 60 | Hard per-domain deadline |
requestConcurrency | 10 | Auto-clamped to the run's memory (~80 MB per parallel domain) |
requestTimeoutSecs / maxRequestRetries | 25 / 2 | Each retry on a fresh proxy IP |
discoveryChannels | all five | homepage, anchor, pathGuess, sitemap, wpJson |
deepJsDiscovery | false | Read inline JSON and JS bundles for a client-rendered footer. No browser |
followWwwAndRootVariants | true | Retry with/without www. |
proxyConfiguration / proxyCountry | Apify datacenter / none | Country pinning forces residential — see the input hint |
escalateToResidentialOnBlock | true | One escalation on a genuine block |
escalateToUnblockerOnBlock | false | Second escalation for the anti-bot tail |
customUserAgent / extraHttpHeaders | — | |
respectRobotsTxt | false | See Legal & fair use |
validateTaxId | true | Run the NIP and REGON checksums |
extractRegon / extractCourt / extractShareCapital / extractOfficer / extractSocials / extractPolicyUrls | true | Per-field switches |
emailPolicy | all | all · exclude-role (drop biuro@, kontakt@, …) · role-only |
requireRegistryId / requireKrs / requireValidNip / requireContact | false | Quality gates — filtered rows are never charged |
legalFormFilter | — | e.g. only sp. z o.o. and S.A. |
minShareCapital | 0 | Filter on the published kapitał zakładowy |
minFieldsRequired | 0 | Minimum populated value fields |
tldFilterMode / tldFilter | none | Opt-in suffix filter |
dedupeBy | krs-then-domain | Deduplication happens before billing |
includeMissRows | false | Deliver an unbilled row for every domain that produced nothing |
includeFieldSources | false | Which page each field came from |
flattenOutput | true | Flat row (CSV-friendly) vs nested objects |
💷 Pricing
Pay per event: one charge per company disclosure returned. No monthly fee. No start fee.
You are charged only for a delivered row carrying at least one registry identifier (KRS, NIP or
REGON) or a named registry court. Delivery and billing are atomic — pushData(item, event) — so
you can never be charged for a row you did not receive, and never receive one you were not charged
for.
Never charged:
- a host that does not resolve or refuses the connection
- a host that answers with a 403 or a challenge page
- a domain that publishes no disclosure — including every sole trader, who has no duty to
- a duplicate, collapsed before billing
- a row removed by your own quality filters
- an unbilled coverage row (
includeMissRows)
RUN_SUMMARY in the key-value store keeps those apart — unreachable_host,
blocked_or_challenged, publishes_no_disclosure, filtered_out_by_your_settings, duplicate,
suppressed, skipped_by_robots, error — so you can reconcile your whole input list against
your invoice.
Set maxItems to cap a run's cost absolutely.
❓ FAQ
Do I need a KRS number to use this? No — and that is the point. Every other Polish company-data product on the store is keyed by a NIP or a KRS you must already have. This one starts from a domain, which is what a lead list actually contains.
How is this different from scraping the KRS register?
The register tells you about a company you have already identified. This tells you which company
a website belongs to, and adds the email and phone the register does not hold. Use both: the
registerUrl on every row is a direct link into the Ministry of Justice's own KRS API for the full
extract.
Why is the NIP fill so much higher than the KRS fill?
Because a NIP is what a Polish business puts in its footer for invoicing, while the KRS is what the
statute requires — and only from companies the statute binds. Sole traders have a NIP and no KRS
at all. statutoryScope tells you which case each row is.
Is nipValid: false a bug?
No. It means the company published a NIP that fails the official checksum — a typo in their own
footer, most often. You are seeing the site as it is. Use requireValidNip to drop those rows.
Why is there no KRS checksum? Because the KRS does not have one. See Honest limits.
Can I run this monthly without paying twice for the same companies?
Yes. Pass the previous run's dataset id as previousDatasetId and every KRS, NIP and domain it
delivered is suppressed before a single request is made.
Does it use a browser?
No. got-scraping + cheerio, which is what keeps a run cheap. deepJsDiscovery reads inline
JSON and JavaScript bundles for a client-rendered footer without launching one.
What if a site blocks the Apify proxy?
A genuine block escalates to a residential IP automatically, and escalateToUnblockerOnBlock adds
a second escalation. You can also supply your own proxyConfiguration. Blocked domains are
reported in RUN_SUMMARY and are never charged.
⚠️ Honest limits
- The KRS carries no checksum.
krsFormatValidis a format and range check — ten digits, not all zeros — and nothing more. Claiming a checksum here would be a lie you could test in five minutes. The NIP and REGON checks are real checksums, and they are the ones to filter on. - The duty binds only some legal forms. A sole trader, a sp.j., a sp.k. and a s.c. have
no KSH website-disclosure obligation. On a mixed domain list, a large share of the misses are
correct output.
statutoryScopeis how you tell. Feed it limited-company domains. - A town read out of a
z siedzibą w …clause is in the Polish locative case, exactly as the company printed it —Bytowiefor Bytów,Krakowiefor Kraków. Reversing a Polish declension would be a guess, and a wrong guess ships as a populated wrong field, which is worse than a published one. The postcode is canonical either way; filter on that. registeredOfficeis sometimes postcode + town only, when the company publishes no street in its statutory clause. That is what the page says.officerNamefill is low (16-17%, stable across both frames). Polish companies publish the zarząd in the KRS, not usually on the website. Where a site does name one, it is captured; where it does not, the field is null rather than inferred.shareCapitalfill is 13-21% depending on the list. The art. 206 §1 pt 4 duty is widely honoured on the regulamin page and widely ignored in footers. This is compliance reality, not a parser gap — the conditional emit rate where the label is present is 95% offline and 100% on the live re-fetch. Note the denominator there is only 19 pages, which is too small to be evidence on its own - the validator prints it as not-enforced rather than reporting a flattering percentage.- The REGON negative-direction sweep flags ~10% of rows. Every one inspected was a phone number passing a nine-digit checksum by luck. Stated so you can reproduce the check rather than take the 0-misses claim on trust.
- A site can be wrong about itself. Everything here is what the company published. Where it
publishes a stale address or a mistyped NIP, that is what you get — with
nipValid: falsetelling you so. - Rating and last-updated of the source pages are not knowable from the bytes; if a company has not touched its regulamin since 2019, its disclosure is 2019's.
🧾 Legal & fair use
These are statutory public disclosures that Polish companies are required by the Kodeks spółek handlowych to publish on their own websites. The Actor reads only pages a company publishes openly; it never logs in, never bypasses a paywall, and never touches personal data behind an account.
respectRobotsTxt is off by default and available as an input, because some buyers' own
compliance policy asks for it. Skipped pages are reported and never charged.
You are responsible for complying with each site's Terms of Service and with GDPR / RODO for
any personal data in the output — a named prezes zarządu, a personal email address. Fields that
can carry personal data (officerName, officerRole) can be switched off with extractOfficer.
Part of a family of statutory website-disclosure scrapers: 🇩🇪 Impressum · 🇬🇧 UK trading disclosures · 🇳🇱 KvK · 🇧🇪 ondernemingsnummer · 🇮🇹 note legali · 🇪🇸 aviso legal · 🇫🇷 mentions légales · 🇳🇴 Brønnøysund.