# Czech Company Scraper - IČO, DIČ & Website Contacts (`scrapersdelight/cz-ico-website-contact-scraper`) Actor

Turn Czech company domains into B2B leads from each site's statutory disclosure (občanský zákoník § 435): obchodní firma, IČO with its mod-11 check digit, DIČ, sídlo, the rejstřík entry (soud, oddíl, vložka), datová schránka, email and phone — optionally checked against ARES. $0.007 per disclosure.

- **URL**: https://apify.com/scrapersdelight/cz-ico-website-contact-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$7.00 / 1,000 per statutory disclosure returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 🇨🇿 Czech IČO Website Scraper — IČO, DIČ, Sídlo, Rejstřík & Contacts

Give it a list of Czech company domains. Get back the **statutory disclosure each site is legally
obliged to publish about itself**: registered name, **IČO with its mod-11 check digit verified**,
DIČ, registered seat, the obchodní rejstřík entry (court, oddíl, vložka), datová schránka, email
and phone — and, if you want it, the same company **checked against ARES**, the Ministry of
Finance's official register, in the same run.

A domain that is dead, blocked, or publishes no disclosure is **never delivered and never charged**.

**$0.007 per disclosure returned.** No subscription. No login. No browser.

***

### ⚖️ Why this data exists, and why it is clean

**Zákon č. 89/2012 Sb., občanský zákoník, § 435 odst. 1:**

> *"Každý podnikatel musí uvádět na obchodních listinách a v rámci informací zpřístupňovaných
> veřejnosti prostřednictvím dálkového přístupu své jméno a sídlo. Podnikatel zapsaný v obchodním
> rejstříku uvede na obchodní listině též údaj o tomto zápisu včetně oddílu a vložky; podnikatel
> zapsaný v jiném veřejném rejstříku uvede údaj o svém zápisu do tohoto rejstříku… Údaj
> o identifikujícím čísle se uvede, bylo-li přiděleno."*

"Information made available to the public by remote access" is the website. So **every Czech
entrepreneur** — company, družstvo, spolek and **sole trader (OSVČ) alike** — has to publish, on
its own site, its name and its registered seat, its register entry including the *oddíl* and
*vložka* where it has one, and its IČO where one has been assigned.

Three things follow, and they are the whole point of this Actor:

1. **The corpus is the entire Czech commercial web**, not one directory with a member list. If a
   Czech business has a website, the disclosure is supposed to be on it.
2. **The business wrote it about itself.** This is not a scraped guess or a third-party
   aggregator's stale copy — it is the company's own statutory statement of who it is.
3. **It is public by design.** § 435 exists precisely so that anyone dealing with the business can
   read it.

This is the same product shape as our DACH Impressum, UK Trading Disclosure, Dutch KvK, Italian
Note Legali, Spanish Aviso Legal, French Mentions Légales, Belgian Ondernemingsnummer and
Norwegian Enhetsregisteret scrapers — in the jurisdiction none of them cover.

***

### 🔍 What it does, per domain

Plain HTTP. **No browser**, which is what keeps a run cheap and fast.

1. **GET the homepage and read its FOOTER.** On the Czech corpus the § 435 disclosure is in the
   site-wide footer surprisingly often. Free — the page is already fetched.
2. **If fields are still missing, follow the footer's own links**, ranked: *obchodní podmínky* >
   *kontakt / kontakty* > *o nás* > *ochrana osobních údajů / GDPR*. Then the XML sitemap, then the
   WordPress page index, then the conventional Czech paths (`/kontakt`, `/obchodni-podminky`,
   `/vop`, `/o-nas`, `/gdpr`, …).
3. **Merge field by field.** A Czech site routinely puts the IČO in the footer, the email on
   `/kontakt` and the full register entry on `/obchodni-podminky`. One page is often not the whole
   disclosure, so up to three pages are parsed and merged, and `fieldSources` can tell you which
   URL every value came from.
4. **Optionally look the IČO up in ARES** and return the register's own answer beside the site's.

**Every statutory field is read inside a statement window** anchored on a disclosure phrase —
never by a page-wide scan. A page-wide 8-digit grab on a Czech site returns an order number, a
bank-account prefix or a product code far more often than an IČO.

***

### 📦 What you get — every field

#### Identity

| Field | What it is |
|---|---|
| `registeredName` | The obchodní firma as the site states it. By Czech law the legal form is part of the firm name, which is what makes this readable. |
| `registeredNameSource` | `label` · `suffix` · `before-statutory-clause` · `register-clause` — how it was found. |
| `tradingName` | A trading name where the site names one separately. |
| `companyName` | Best available name: the registered one, else JSON-LD `legalName`, else `og:site_name`. |
| `legalForm` | `Společnost s ručením omezeným`, `Akciová společnost`, `Spolek`, `Družstvo`, … |
| `legalFormAbbrev` | The short form: `s.r.o.`, `a.s.`, `v.o.s.`, `z.s.`, `o.p.s.`, `OSVČ`. |

#### Registry — the wedge

| Field | What it is |
|---|---|
| `ico` | The IČO, canonical 8 digits, zero-padded. |
| `icoRaw` | Exactly as the page printed it (`26 168 685`, `CZ26168685`, …). |
| `icoSource` | `labelled` · `labelled-loose` · `register-clause` · `bare` · `derived-from-dic`. Everything except `labelled` is **only** emitted when the check digit passes. |
| **`icoChecksumValid`** | **The IČO run through its weighted mod-11 check digit** (weights 8,7,6,5,4,3,2). A number that fails is still delivered, flagged `false`, because that is a fact about the website. |
| `icoFormatValid` | Format check, reported separately from the checksum. |
| `dic` | The DIČ (`CZ` + 8–10 digits). |
| `dicCountry` / `dicKind` | `CZ` or `SK`; `legal-person` or `natural-person`. |
| **`dicMatchesIco`** | **Do the tax id and the registry id agree?** For a Czech legal person the DIČ is exactly `CZ` + the IČO, so the two check each other. A sole trader's DIČ is built on a rodné číslo and legitimately does not match — that is why `dicKind` is published beside it. |
| `registerKind` | `obchodní rejstřík` · `živnostenský rejstřík` · `spolkový rejstřík` · `nadační rejstřík` · … |
| `registerCourt` | `Městský soud v Praze`, `Krajský soud v Brně`, … — always in the nominative, whichever case the site used. |
| `registerSection` / `registerInsert` | The *oddíl* (`A`,`B`,`C`,`Dr`,`H`,`L`,`N`,`Pr`,`R`,`S`,`Zs`) and the *vložka*. |
| `registerFileRef` | The two assembled: `C 34244, Krajský soud v Hradci Králové`. |
| `registerUrl` | Deep link into **or.justice.cz** for this exact IČO. |
| `aresUrl` | Deep link into **ARES** for this exact IČO. |
| `registeredOffice` | The sídlo. |
| `registeredOfficePostcode` / `registeredOfficeCity` | PSČ and obec, parsed out of it. |
| `registeredOfficeSource` | `label` (an explicit "se sídlem") or `disclosure-window` (an address sitting inside the disclosure block). Published so you can filter to the strictly-labelled seat if you need to. |

#### Contact

| Field | What it is |
|---|---|
| `email` / `emails` | Primary plus up to ten. Cloudflare-obfuscated addresses are decoded; `(zavináč)` / `(tečka)` spellings are reassembled. |
| `emailConflict` | Set when the primary email is on a domain that is neither the site's nor a freemail provider — usually a web agency's. |
| `phone` / `phones` | E.164 (`+420…`). |
| `addressLine` / `postcode` / `city` | The contact or provozovna address, which is frequently a different place from the sídlo. |
| **`dataBoxId`** | **The datová schránka ID** — the company's statutory electronic mailbox, a legally-served delivery channel. No other jurisdiction in this family has an analogue. |
| `bankAccount` / `iban` | The invoicing account where the site publishes one. The IBAN is checked against its own ISO 13616 mod-97; the domestic account number is returned **as published**, with no ČNB checksum claimed. |
| `officerName` / `officerRole` | A named jednatel / ředitel / kontaktní osoba. Personal data — handle under GDPR. |
| `socialLinks`, `termsUrl`, `privacyPolicyUrl` | From the page's own links. |

#### ARES verification — optional, and it never overwrites the site

Turn on **"Verify the IČO against ARES"** and each row also carries `aresName`,
`aresRegisteredOffice`, `aresPostcode`, `aresCity`, `aresRegion` (kraj), `aresDic`,
`aresVatRegistered`, `aresLegalForm` (+ its code), `aresIncorporatedOn`, `aresDissolvedOn`,
`aresActive`, `aresNace`, `aresPrimarySource`.

And three flags that are the actual point:

- `aresNameAgrees`, `aresAddressAgrees`, `aresDicAgrees` — `true`, `false`, or `null` when there is
  nothing to compare.
- `aresDisagreement` — a sentence naming the difference, e.g.
  *"name: site says "Galileo Corporation s.r.o.", register says "Obec Klentnice""*.
- `aresStatus` — `found` · `not-in-register` · `lookup-error` · `not-looked-up`.

**The register never silently overwrites what the website published.** The site's own fields stay
exactly as parsed and ARES lands in its own `ares*` columns. Where the two disagree the row says
so, and you decide. `not-in-register` is a *verified* negative (ARES answers a non-existent IČO
with a documented not-found body, which is checked for) — a transport failure is reported as
`lookup-error` and never dressed up as "no such company".

#### Provenance & accounting

`domain`, `inputUrl`, `resolvedUrl`, `disclosureUrl`, `disclosureSource`, `discoveryChannel`,
`pagesParsed`, `pagesChecked`, `fieldCount`, `stableId`, `status`, `missReason`, `country`,
`statute`, `fetchedAt`, `elapsedMs`, and `fieldSources` on request.

***

### 📊 Measured results — real numbers, not claims

Everything below was measured. Nothing is an estimate.

#### The corpus

Real Czech businesses sourced from **OpenStreetMap** (`office`, `shop`, `craft`, `amenity`,
`healthcare`, `tourism` elements carrying a `website` tag), clipped to the Czech state boundary and
picked **without reference to whether they publish a disclosure**. Numbers for the sweep used here
are in `SIGNOFF.md`.

#### The headline runs: 1,191 Czech domains across TWO unrelated sampling frames

One corpus measures a *parser*; it does not measure a *population*. So everything below that
describes the Czech web — how much of it is reachable, how much publishes a disclosure, where the
disclosure sits — was measured **twice, on two frames that share no sampling basis**:

- **Frame 1 — OpenStreetMap.** Point-of-interest based: shops, offices, crafts and amenities with
  a `website` tag, clipped to the Czech border. Skews to businesses with physical premises.
- **Frame 2 — Tranco.** Popularity/traffic based: `.cz` hosts from an aggregated ranking list. No
  geography, no POI tags. Skews to larger and more online-native businesses.

The two host lists overlap by only 10.8%, and the frame-2 sample was drawn from the **disjoint
remainder**, so no domain appears in both. Both runs used the same build.

| | frame 1 · OSM | frame 2 · Tranco |
|---|---|---|
| Domains attempted | 591 | 600 |
| Dead — did not resolve | 91 (15.4%) | 66 (11.0%) |
| Blocked / challenged | 24 (4.1%) | 38 (6.3%) |
| **Dead-or-walled** | **19.5%** | **17.3%** |
| Reachable | 476 | 496 |
| **Published a § 435 disclosure** | **209 — 43.9%** | **191 — 38.5%** |
| Rows delivered · rows charged | 209 · **209** | 191 · **191** |

**Delivered == charged, exactly, on both.** Delivery and billing happen in one call.

**Plan on 38–44% of reachable**, and on roughly a fifth of any OSM- or crawl-derived Czech host
list being dead or walled before you start. Both figures sit in the **Belgian** band the statute
predicts rather than the UK's: reg. 25 binds UK *companies*, while § 435 binds *every* Czech
entrepreneur, sole traders included.

#### Where the disclosure sits depends on your list — so we give you the range

This is the number that moved most between frames, and it is worth knowing before you set
`maxPagesParsed` or turn channels off:

| Found on | frame 1 · OSM | frame 2 · Tranco |
|---|---|---|
| The homepage footer | 102 (48.8%) | 59 (30.9%) |
| A dedicated page | 95 (45.5%) | 111 (58.1%) |
| Merged across both | 12 (5.7%) | 21 (11.0%) |

An **18-point swing**, and the two frames disagree about which is more common. A high-street
business puts its IČO in the site-wide footer; a larger online business puts it on
`/obchodni-podminky` and links to it. That is why the cascade runs the homepage *and* the ranked
links *and* the guessed paths by default — on either frame, turning one off would cost you
roughly a third of the rows.

#### Offline validation, against the raw bytes

The shipping parser was re-run over **172 pages captured live through the shipping proxy**, with
**7,233 assertions** comparing every advertised field to independent evidence in the source.
**0 failures.**

- **IČO check digit: 72 valid, 0 invalid, across every row that published one.**
- The checksum is **re-implemented a second, different way** in the validator (closed form
  `(11 − r) % 10`, no branches, versus the shipping code's branching rule), so a bug in one cannot
  validate itself. Both are anchored on four real published IČOs — Seznam.cz `26168685`, Alza.cz
  `27082440`, Škoda Auto `00177041`, ČEZ `45274649` — and on four deliberate one-digit corruptions.

**Conditional emit rate — the check that matters more than fill.** Of the 72 pages that carried an
IČO **label** at all, **71 emitted a value: 98.6%** — and the 591-domain platform run reproduced
the same figure independently, **205 of 208, 98.6%**. A fill rate tells you nothing on its own,
because a dead regex produces plausible nulls rather than errors. A near-100% *conditional* rate
says the parser is sound and the field is genuinely absent on the rest of the corpus.

**Negative sweep — the check in the other direction.** Every page that emitted *no* IČO was swept
with the independent checksum for a valid one the parser had missed. 3 of 100 turned up a
checksum-valid 8-digit run. Eyeballed one by one and looked up in ARES:

| Page | Candidate | Verdict |
|---|---|---|
| mesto-votice.cz | 26112019 | a date, 26.11.2019 |
| novypoddvorov.cz | 14092026, 29072026 | dates in a news list |
| zshstropnice.cz | 75000776 | **a genuine miss** — a real IČO, printed bare with no label and no register clause anywhere near it |

A random 8-digit run passes a mod-11 check digit about one time in ten, so the sweep is an **upper
bound**, not a defect count. **Measured genuine miss rate: 1 in 100.**

#### Field fill — and why it depends on the list you bring

Measured on both frames, 400 billed rows. **Fill is not a fixed property of this Actor; it is a
property of what your businesses publish**, and the two frames differ enough to matter:

| Field | frame 1 · OSM | frame 2 · Tranco |
|---|---|---|
| `ico` | **99.0%** | **97.4%** |
| `email` | 91.4% | 89.0% |
| `phone` | 94.7% | **80.1%** |
| `addressLine` | 86.6% | 81.7% |
| `registeredOffice` | 85.2% | 87.4% |
| `registeredOfficePostcode` | 82.8% | 79.6% |
| `registeredOfficeCity` | 68.9% | 57.6% |
| `registeredName` | 57.9% | **73.3%** |
| `legalForm` | 55.0% | **69.1%** |
| `dic` | 45.5% | **57.1%** |
| `registerCourt` | 17.7% | **29.8%** |
| `registerFileRef` | 13.4% | **25.1%** |
| `dataBoxId` | 18.2% | 14.1% |
| `bankAccount` | 8.1% | 11.5% |
| `officerName` | 10.0% | 8.9% |

Median **11** populated value fields per row, p90 **14**, max **18** of 22.

The pattern is consistent and useful: a **premises-based list** (Maps, a directory, OSM) gives you
more **phones and addresses**; a **domain-based list** of larger online businesses gives you more
**registry depth** — registered name, legal form, DIČ, and roughly double the register entries.
`ico` and `email` are near-flat on both, which is what you are mostly buying.

**The IČO check digit held on every row of both runs: 393 valid, 0 invalid.** Where each came
from, across all 400 rows: a tight label 378 · the checksum-gated loose pass 6 · derived from a
legal-person DIČ 4.

**`dicMatchesIco` agreed on 77 of 78** frame-1 rows publishing both a Czech legal-person DIČ and
an IČO. The single mismatch is a real finding: that site publishes its own IČO beside a different
entity's DIČ.

**The IČO check digit held on the whole run: 207 valid, 0 invalid.** The two rows without an IČO
qualified on a DIČ and a register entry instead. Where each IČO came from: a tight label 199, the
checksum-gated loose pass 5, derived from a legal-person DIČ 3 — the two fallback passes earned
their place on 8 rows a tight-label-only parser would have dropped.

**`dicMatchesIco` agreed on 77 of 78** rows where both a Czech legal-person DIČ and an IČO were
published. The single mismatch is a real finding, not a parse error: that site publishes its own
IČO next to a different entity's DIČ.

#### ARES verification, measured over 393 lookups

**393 looked up · 383 found · 10 not in the register · 0 lookup errors.** On frame 1, name agrees
88 / disagrees 32; seat agrees 125 / disagrees 46; DIČ agrees 86 / disagrees **0**.

The disagreements are the point of the feature, and they fall into three kinds.

**The website being loose.** A municipal site whose footer names the web agency that built it
("Galileo Corporation s.r.o.") rather than the obec whose IČO it publishes.

**A genuine firmographic.** A Lipno holiday-let whose operator's registered seat is in Prague, not
at the property.

**A Slovak company on a `.cz` domain — and this is the one you cannot catch any other way.** All
ten "not in the register" rows came from frame 2, and **six of them are Slovak businesses selling
into Czechia** (kondela.cz, dedoles.cz, nejzlato.cz, noezon.cz, naureus.cz, enerso.cz). A Slovak
IČO is *also* eight digits and uses the *same* mod-11 check digit, so the checksum cannot tell
them apart and neither can any format test. Two independent signals do: `dicCountry: "SK"` on the
published tax id, and `aresStatus: "not-in-register"`. If you are buying Czech leads, those ten
rows are the ones to look at first.

**ARES never overwrites or completes what the website published.** It is an exact-key lookup, not
a search: a corrupted digit, a transposition and a zero-padded short number all return "not
found" rather than a plausible near-match, and a malformed key is reported as `lookup-error`
rather than dressed up as "no such company". The site's own fields stay exactly as parsed.

**Read the low numbers honestly.** The register entry is at ~20% because most small Czech traders
publish an IČO and stop there — § 435 requires the *oddíl* and *vložka* only of those entered in
the obchodní rejstřík, and a large share of this corpus are OSVČ in the živnostenský rejstřík who
have neither. The same is true of the datová schránka: 15% is how many Czech SMEs publish theirs,
not how many the parser can find. Where a **label** was present, the parser emitted a value on
**100%** of pages for `registerKind` and for the labelled seat, and **90%** for `dataBoxId`.

#### A real row, from a real run

```json
{
  "domain": "charon-eu.cz",
  "disclosureUrl": "https://www.charon-eu.cz/kontakty",
  "disclosureSource": "legal-page",
  "discoveryChannel": "guess",
  "registeredName": "Jitka Filipová s.r.o.",
  "registeredNameSource": "suffix",
  "legalForm": "Společnost s ručením omezeným",
  "legalFormAbbrev": "s.r.o.",
  "ico": "03515630",
  "icoRaw": "03515630",
  "icoSource": "labelled",
  "icoChecksumValid": true,
  "dic": "CZ03515630",
  "dicKind": "legal-person",
  "dicMatchesIco": true,
  "registerKind": "obchodní rejstřík",
  "registerCourt": "Krajský soud v Hradci Králové",
  "registerSection": "C",
  "registerInsert": "34244",
  "registerFileRef": "C 34244, Krajský soud v Hradci Králové",
  "registerUrl": "https://or.justice.cz/ias/ui/rejstrik-$firma?ico=03515630",
  "aresUrl": "https://ares.gov.cz/ekonomicke-subjekty?ico=03515630",
  "registeredOffice": "Kyjevská 39, Pardubičky, 530 03 Pardubice",
  "registeredOfficePostcode": "530 03",
  "registeredOfficeCity": "Pardubice",
  "registeredOfficeSource": "label",
  "email": "benesov@charon-eu.cz",
  "phone": "+420603229333",
  "officerName": "Jindřich Filip",
  "officerRole": "Jednatel",
  "fieldCount": 17,
  "status": "ok"
}
```

***

### 🧰 Every input

**Which websites to read** — `domains` (bare domain, homepage URL, or a direct link to the
disclosure page), `startUrls` (Apify's URL-list format, so Make / Clay / Sheets can hand a list
over natively), `sourceDatasetId` + `domainFieldName` (enrich another Actor's output),
`domainsFileUrl` (a CSV/TSV/TXT/JSON/JSONL at a URL), `skipDomains`, `previousDatasetId` (never
re-buy a row a previous run already delivered).

**Limits and cost** — `maxItems` (a cap on **billed** rows), `maxDomains`,
`maxDiscoveryRequestsPerDomain`, `maxPagesParsed`, `perDomainTimeoutSecs`, `requestConcurrency`,
`requestTimeoutSecs`, `maxRequestRetries`.

**How to find the disclosure** — `discoveryChannels` (homepage · footer links · sitemap ·
WordPress index · guessed Czech paths), `followWwwAndRootVariants`, `deepJsDiscovery` (reads the
page's inline JSON and JS bundles for a client-rendered footer — still no browser),
`respectRobotsTxt`.

**Network** — `proxyConfiguration`, `proxyCountry`, `escalateToResidentialOnBlock`,
`escalateToUnblockerOnBlock`, `customUserAgent`, `extraHttpHeaders`.

**Verification** — `verifyAgainstAres`, `requireAresMatch`, `validateChecksum`.

**Which rows to keep** — `requireIco`, `requireValidIco`, `requireContact`, `minFieldsRequired`,
`tldFilterMode` + `tldFilter`, `emailPolicy` (all / role-only / exclude-role), `dedupeBy`.

**Which fields to extract** — `extractRegisterEntry`, `extractDataBox`, `extractBankAccount`,
`extractOfficer`, `extractSocials`, `extractPolicyUrls`.

**Output** — `includeMissRows`, `includeFieldSources`, `flattenOutput`.

***

### 💰 Pricing

**$0.007 per statutory disclosure returned.** Pay-per-event, one event, no start fee, no
subscription.

**You are charged for a row only when it is delivered to you**, and delivery and billing happen in
the same call, so the two can never drift apart. These are **never** charged:

- a host that does not resolve or refuses the connection
- a host that answers with a challenge page
- a domain that publishes **no** disclosure
- a duplicate (collapsed *before* delivery)
- a row your own filters removed
- an unbilled coverage row

`maxItems` is a hard ceiling on billed rows, so you always know the worst case before you start.
The run's `RUN_SUMMARY` record keeps the three kinds of nothing apart — unreachable, blocked,
publishes-no-disclosure — so you can see exactly what your list did.

***

### ❓ FAQ

**How is this different from an ARES scraper?**
An ARES actor is handed an IČO or a company name and returns the register record. This Actor is
handed a **website** and its whole job is to find the identifier — the case where you have a domain
list (from Maps, a directory, an ad platform, a CRM export) and no identifiers at all. The two are
complements: turn the domain into an IČO here, then the IČO opens every Czech register there is.
You can also do both in one pass with **Verify the IČO against ARES**.

**Do I need an ARES key or a login?** No. Nothing here needs an account anywhere.

**Does it use a browser?** No. `got-scraping` + `cheerio`. That is what keeps it cheap. For the
minority of sites whose footer only exists after a client-side render, `deepJsDiscovery` reads the
inline JSON payloads and JS bundles a render would have read from — still without launching one.

**What if a site publishes an IČO that fails the check digit?** You get it, with
`icoChecksumValid: false`. Dropping it would hide a fact about that website. Use
`requireValidIco` if you only want numbers that pass.

**Why is `registeredName` only filled about half the time?** Because a Czech footer frequently
prints the IČO and the seat without repeating the firm name (it is already in the site header or
the logo). Where the name *is* in the disclosure block, it is read. If you need a name on every
row, turn on **Verify the IČO against ARES** — `aresName` is the register's own, and it is
populated for every IČO the register knows.

**Can two domains of the same company produce two billed rows?** Not by default. `dedupeBy` is
`ico-then-domain`, and the collapse happens **before** delivery, so a duplicate can never be
charged.

**Can I re-run monthly without re-buying rows?** Yes — pass the previous run's dataset id as
`previousDatasetId` and every IČO and domain it delivered is suppressed before any request is made.

***

### ⚠️ Honest limits

- **An unlabelled number is refused, however valid it looks — this is deliberate, and it cost us
  rows.** An earlier version read any checksum-valid 8-digit run inside a disclosure block even
  with no IČO label. Measured across 1,191 domains on two frames, every route with a label behind
  it ran at 98–100% precision against the official register; the unlabelled one ran at **40%** —
  2 right, 3 wrong, and its wrong answers were `20230401` and `01062026`, which are dates. A
  mod-11 check digit is passed by luck about one time in ten, so a checksum is not evidence that
  an unlabelled number is an identifier; only the label is. The pass was removed. It costs about
  2 rows per 1,200 domains, and buys you a field you can trust.
- **So a bare IČO with no label and no register clause anywhere near it is not read.** Measured at
  **1 page in 100** that emitted no IČO.
- **`tradeOffice` was cut.** The plan was to publish the municipal trade-licensing office named in
  an OSVČ disclosure. Across 172 live pages the phrase appears on **zero** of them: Czech sole
  traders name the register, not the office that keeps it. A column of nulls is worse than no
  column.
- **`registeredName` can be the web agency's on a municipal or small-business site.** A footer
  reading "© 2026 Galileo Corporation s.r.o." next to the municipality's own IČO is common.
  **Verify the IČO against ARES** catches exactly this: `aresNameAgrees: false` plus an
  `aresDisagreement` naming both.
- **A Slovak business on a `.cz` domain looks Czech, and it is not rare.** A Slovak IČO is also 8
  digits and uses the *same* mod-11 check digit, so neither the checksum nor any format test can
  tell them apart. **Measured: 6 of 600 domains on the popularity frame** (kondela.cz, dedoles.cz,
  nejzlato.cz, noezon.cz, naureus.cz, enerso.cz). Two tells are published rather than guessed at:
  `dicCountry: "SK"` when the page's tax id is Slovak, and `aresStatus: "not-in-register"` when
  ARES has never heard of the number. Turn on the ARES check if this matters to you.
- **No ČNB bank-account checksum.** `bankAccount` is returned as published. Only the IBAN is
  checked (ISO 13616 mod-97).
- **Some hosts refuse the Apify datacenter proxy.** Residential escalation is on by default and
  Unblocker is one switch away; both are reported in the run's transport stats so you can see what
  happened rather than guess.
- **Personal data.** `officerName`, and any personal email, are personal data under GDPR. § 435
  makes the *disclosure* public; it does not make your processing lawful. That part is yours.

***

### 🧾 Legal & fair use

This Actor reads pages that Czech law requires those businesses to publish for public reading, and
reads them the way a browser does: plain HTTP, ordinary headers, modest concurrency, a hard
per-domain budget. It solves no CAPTCHA, defeats no paywall and uses no login.

`robots.txt` is **not** obeyed by default, because a § 435 disclosure is published under a legal
duty of publicity; the switch exists for buyers whose own compliance policy asks for it. You remain
responsible for complying with each site's Terms of Service and for handling any personal data in
the output lawfully.

Data source: the businesses' own websites. Optional verification: **ARES**, provided by the Czech
Ministry of Finance as an open public register.

# Actor input Schema

## `domains` (type: `array`):

The company websites to read. A bare domain (alza.cz), a homepage URL (https://www.alza.cz/) or a direct link to the page carrying the disclosure (https://www.alza.cz/obchodni-podminky) all work. Leave every input empty and the Actor runs a small built-in Czech demo batch so the run still returns rows.

## `startUrls` (type: `array`):

The same list in Apify's standard URL-list format, so an upstream Actor or an integration can hand it over natively. Merged with Domains.

## `sourceDatasetId` (type: `string`):

Dataset ID of an earlier run (yours or another Actor's). Every row's website field becomes an input domain.

## `domainFieldName` (type: `string`):

Field name to read the domain from in the source dataset or file, e.g. "website" or "company.url". Left empty, the Actor tries domain, website, websiteUrl, url, site, homepage.

## `domainsFileUrl` (type: `string`):

A public CSV, TSV, TXT, JSON or JSONL file of domains, one per line or one per row.

## `skipDomains` (type: `array`):

Domains you already have. They are dropped before any request is made, so they cost nothing and are never charged.

## `previousDatasetId` (type: `string`):

Dataset ID of an earlier run of this Actor. Every IČO and domain it delivered is suppressed, so a monthly re-run never re-buys a row you already paid for.

## `maxItems` (type: `integer`):

Stop after this many BILLED rows. A row is billed only when a statutory disclosure was found and delivered, so this is a hard ceiling on what the run can cost. 0 means no cap.

## `maxDomains` (type: `integer`):

Stop queueing after this many input domains, whether or not they yield a disclosure. 0 means no limit.

## `maxDiscoveryRequestsPerDomain` (type: `integer`):

How many pages the Actor may fetch while hunting for the disclosure on one site, after the homepage.

## `maxPagesParsed` (type: `integer`):

How many fetched pages are actually parsed and merged. Czech sites routinely split the disclosure between the footer, /kontakt and /obchodni-podminky, so more than one is normal.

## `perDomainTimeoutSecs` (type: `integer`):

Give up on one domain after this long and move on. The domain is reported as unreachable and never charged.

## `requestConcurrency` (type: `integer`):

How many domains to work on at once. Each one needs roughly 80 MB (a fetched page plus its DOM), so the Actor lowers this automatically if the memory you allocated cannot support it.

## `requestTimeoutSecs` (type: `integer`):

Timeout for a single HTTP request.

## `maxRequestRetries` (type: `integer`):

How many times to retry a failed request, each time on a fresh proxy IP.

## `discoveryChannels` (type: `array`):

The cascade, in order. homepage reads the site-wide footer (where most Czech disclosures live), anchor follows the footer's own links to obchodní podmínky / kontakt / o nás, sitemap and wpJson index the site, pathGuess tries the conventional Czech paths.

## `followWwwAndRootVariants` (type: `boolean`):

If the host does not answer, try the other form once before giving up.

## `deepJsDiscovery` (type: `boolean`):

For sites whose footer only exists after a client-side render. No browser is launched: the Actor reads the inline JSON payloads and JS bundles a render would have read from. Slower, off by default.

## `respectRobotsTxt` (type: `boolean`):

Off by default. A statutory disclosure is published under § 435 precisely so that the public can read it; this switch exists because some buyers' own compliance policy asks for it.

## `proxyConfiguration` (type: `object`):

Apify DATACENTER proxy by default, which is the cheapest rung that works on this corpus. Supply your own proxies here if you prefer.

## `proxyCountry` (type: `string`):

Pinning a country means RESIDENTIAL exit nodes, because Apify datacenter proxies cannot be pinned to a country. Leave at None for the cheaper, measured-faster datacenter path.

## `escalateToResidentialOnBlock` (type: `boolean`):

When a host refuses the datacenter IP with a 403 or a challenge page, retry once through a residential IP.

## `escalateToUnblockerOnBlock` (type: `boolean`):

A second escalation for hosts that refuse both datacenter and residential. Costs more per request, so it is off by default.

## `customUserAgent` (type: `string`):

Override the browser User-Agent sent with every request.

## `extraHttpHeaders` (type: `object`):

Additional headers sent with every request, as a JSON object.

## `verifyAgainstAres` (type: `boolean`):

After the disclosure is parsed, look the published IČO up in ARES, the Ministry of Finance's official register, and return the register's own name, seat, legal form, DIČ, VAT-payer status and incorporation date alongside the site's. The register NEVER overwrites what the website published: where the two disagree the row says so in aresDisagreement. Adds one request per row.

## `requireAresMatch` (type: `boolean`):

Drop any row whose IČO is not in ARES, or whose registered name the register disagrees with. Dropped rows are never charged. Turns the ARES lookup on automatically.

## `validateChecksum` (type: `boolean`):

Run the published IČO through its weighted mod-11 check digit and report the result in icoChecksumValid. A number that fails is still delivered, flagged false, because that is a fact about the website.

## `requireIco` (type: `boolean`):

Drop rows that publish a disclosure but no IČO (a foreign trader, or an entity that has not been assigned one). Dropped rows are never charged.

## `requireValidIco` (type: `boolean`):

Stricter than the above: the IČO must also pass its mod-11 check digit. Dropped rows are never charged.

## `requireContact` (type: `boolean`):

Drop rows with no way to reach the business. Dropped rows are never charged.

## `minFieldsRequired` (type: `integer`):

Drop any row with fewer than this many populated value fields. 0 keeps everything.

## `tldFilterMode` (type: `string`):

Restrict the input by domain ending. Off by default: a large share of real Czech businesses use .com or .eu, so filtering to .cz would silently discard them.

## `tldFilter` (type: `array`):

The endings the filter applies to, e.g. cz, eu, com.

## `emailPolicy` (type: `string`):

Role mailboxes are info@, obchod@, fakturace@ and the like. Agencies split hard on whether they count as a lead, so it is your call.

## `dedupeBy` (type: `string`):

Collapse duplicates BEFORE delivery, so a duplicate can never be charged twice. Two domains of the same company share one IČO.

## `extractRegisterEntry` (type: `boolean`):

The obchodní rejstřík entry § 435 requires by name: which register, which court, the oddíl and the vložka, assembled into one file reference.

## `extractDataBox` (type: `boolean`):

The seven-character ID of the company's statutory electronic mailbox, where it publishes one.

## `extractBankAccount` (type: `boolean`):

The account number and IBAN a Czech site publishes for invoicing. The IBAN is checked against its own ISO 13616 mod-97; the domestic account number is returned as published.

## `extractOfficer` (type: `boolean`):

A jednatel, ředitel or kontaktní osoba named next to their role. Personal data: handle it under GDPR.

## `extractSocials` (type: `boolean`):

LinkedIn, Facebook, Instagram and the rest, from the page's own links.

## `extractPolicyUrls` (type: `boolean`):

Links to the site's obchodní podmínky and ochrana osobních údajů pages.

## `includeMissRows` (type: `boolean`):

Also deliver a row for every domain that produced nothing, saying which kind of nothing it was: unreachable, blocked, publishes no disclosure, filtered out by your own settings, duplicate or suppressed. None of these is ever charged.

## `includeFieldSources` (type: `boolean`):

Adds fieldSources, mapping every populated field to the URL it was read from.

## `flattenOutput` (type: `boolean`):

One flat object per row, which is what a CSV export and most CRMs want. Turn off for a nested company / registry / contact / banking / people / ares structure.

## Actor input object example

```json
{
  "domains": [
    "karasek.cz",
    "markpjetri.cz",
    "rimoto.cz",
    "1a-trade.cz",
    "envisan.cz",
    "carpfood.cz",
    "nekvinda-obchod.cz",
    "charon-eu.cz",
    "luggi.cz",
    "vinarstvihulata.cz"
  ],
  "maxItems": 1000,
  "maxDomains": 0,
  "maxDiscoveryRequestsPerDomain": 8,
  "maxPagesParsed": 3,
  "perDomainTimeoutSecs": 60,
  "requestConcurrency": 10,
  "requestTimeoutSecs": 25,
  "maxRequestRetries": 2,
  "discoveryChannels": [
    "homepage",
    "anchor",
    "sitemap",
    "wpJson",
    "pathGuess"
  ],
  "followWwwAndRootVariants": true,
  "deepJsDiscovery": false,
  "respectRobotsTxt": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "proxyCountry": "none",
  "escalateToResidentialOnBlock": true,
  "escalateToUnblockerOnBlock": false,
  "verifyAgainstAres": false,
  "requireAresMatch": false,
  "validateChecksum": true,
  "requireIco": false,
  "requireValidIco": false,
  "requireContact": false,
  "minFieldsRequired": 0,
  "tldFilterMode": "none",
  "emailPolicy": "all",
  "dedupeBy": "ico-then-domain",
  "extractRegisterEntry": true,
  "extractDataBox": true,
  "extractBankAccount": true,
  "extractOfficer": true,
  "extractSocials": true,
  "extractPolicyUrls": true,
  "includeMissRows": false,
  "includeFieldSources": false,
  "flattenOutput": true
}
```

# Actor output Schema

## `records` (type: `string`):

The dataset of Czech company leads (one item per domain that published a statutory § 435 website disclosure).

## `runSummary` (type: `string`):

Coverage accounting for the run: domains attempted, rows billed, the IČO check-digit tally, the conditional emit rate, the ARES agreement tally, and the three kinds of nothing (unreachable host, blocked host, publishes no disclosure) kept apart - none of them charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "karasek.cz",
        "markpjetri.cz",
        "rimoto.cz",
        "1a-trade.cz",
        "envisan.cz",
        "carpfood.cz",
        "nekvinda-obchod.cz",
        "charon-eu.cz",
        "luggi.cz",
        "vinarstvihulata.cz"
    ],
    "maxItems": 1000,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/cz-ico-website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "karasek.cz",
        "markpjetri.cz",
        "rimoto.cz",
        "1a-trade.cz",
        "envisan.cz",
        "carpfood.cz",
        "nekvinda-obchod.cz",
        "charon-eu.cz",
        "luggi.cz",
        "vinarstvihulata.cz",
    ],
    "maxItems": 1000,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/cz-ico-website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "karasek.cz",
    "markpjetri.cz",
    "rimoto.cz",
    "1a-trade.cz",
    "envisan.cz",
    "carpfood.cz",
    "nekvinda-obchod.cz",
    "charon-eu.cz",
    "luggi.cz",
    "vinarstvihulata.cz"
  ],
  "maxItems": 1000,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scrapersdelight/cz-ico-website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/cz-ico-website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vZ7yjBI9nRtJbAVUo/builds/bbUxcRbX3PGcEL2qQ/openapi.json
