# Irish Company Website Scraper - CRO No, VAT & Contacts (`scrapersdelight/ie-company-website-disclosure-scraper`) Actor

Turn Irish company domains into registry-grade B2B leads from each site's statutory website disclosure (Companies Act 2014 s.151): registered name, CRO number, legal form, registered office, Eircode, county, mod-23-checked VAT number, charity and regulator IDs, email, phone. $0.006 per disclosure.

- **URL**: https://apify.com/scrapersdelight/ie-company-website-disclosure-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$6.00 / 1,000 per company disclosure returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Irish Company Website Scraper — CRO Number, VAT, Registered Office & Contacts

Turn a list of **Irish company domains** into registry-grade B2B leads, read from the one thing every
Irish company is legally obliged to publish on its own website: its **Companies Act 2014 s.151
disclosure**.

For each domain you get the **registered name, legal form, CRO number, which CRO register that
number is on, registered office (with Eircode and county), mod-23-checked VAT number, charity and
regulator identifiers, email, phone and postal address** — plus the URL of the page each fact came
from.

**$0.006 per disclosure returned.** A domain that is dead, blocked, or publishes no disclosure is
never delivered and **never charged**.

***

### The statute this reads

**Companies Act 2014, section 151 ("Particulars to be published")** requires every Irish company to
specify — *on its websites* — its **name and legal form**, the **place of registration and the number
with which it is registered**, and the address of its **registered office**. s.151 implements what is
now **article 26 of Directive (EU) 2017/1132**, the same European provision behind the German
*Impressum*, the Belgian *art. III.74 CDE* disclosure and the UK's trading-disclosure regulations.

So every field below is a statement the company made about itself **because the law told it to**.
That is what makes this corpus the cleanest Irish B2B source there is: it is not a directory with a
member list, it is the whole Irish commercial web.

**Ireland has no named legal-notice page.** Unlike Germany's `/impressum` or France's
`/mentions-legales`, s.151 says only that the particulars must be *on the website*. Measured on the
platform across two independent 400-domain samples, the winning page was found by the **homepage
footer** 30 and 55 times and by a **ranked footer link** 25 and 69 times, with guessed paths winning
at most once — and **44% to 55% of billed rows** were only complete once two pages were merged.
This Actor reads the footer first and then follows the site's own legal links, rather than guessing
a page name.

***

### What you get — output fields

Every field below was **seen populated in real captured bytes** before it was advertised. Fill
percentages are measured, not estimated, and the honest-limits section says which ones are sparse
and why.

#### Identity

| Field | What it is |
|---|---|
| `registeredName` | The legal name as the disclosure states it ("Inkplus Limited", "Aran Sweater Market ULC") |
| `tradingName` | The "trading as" / "T/A" name where the page gives one |
| `companyName` | Best available name — the registered name, else JSON-LD `legalName`, else `og:site_name` |
| `legalForm` | The Companies Act 2014 type: `LTD`, `DAC`, `CLG`, `ULC`, `PLC`, `LLP`, `LP`, `ILP`, `SE`, `Co-operative society`, `Sole trader`, `Registered business name`. Irish-language statutory forms (`Teoranta`, `Teo`, `Cuideachta Phoiblí Theoranta`) are recognised as the legal name endings they are |

#### Register

| Field | What it is |
|---|---|
| `companyNumber` | The CRO number, canonicalised to bare digits with leading zeros stripped, exactly as the CRO prints it |
| `companyNumberRaw` | Exactly as the page printed it, before normalisation |
| `companyNumberSource` | Which pattern found it: `label-alt`, `jurisdiction`, `jurisdiction-label`, `label`, `bare` |
| `companyNumberFormatValid` | Format + issued-range check. **A CRO number carries no check digit**, so this is honestly named — it is not a checksum |
| `companyNumberScheme` | `IE-CRO` |
| `croRegistrationType` | **`company` or `business-name`** — see "The CRO number trap" below |
| `jurisdiction` | The register the page names: `Ireland`, or `England and Wales` / `Scotland` / `Northern Ireland` when an Irish-facing site is run by a British company |
| `country` | The **register's** country, never the corpus label |
| `registeredOffice`, `registeredOfficePostcode`, `registeredOfficePostcodeValid`, `registeredOfficeCity`, `registeredOfficeCounty` | The registered office, split |
| `vatNumber`, `vatNumberValid`, `vatChecksumVariant`, `vatCountry`, `vatBranch` | VAT id, **validated with the published Irish mod-23 check character** |
| `charityNumber` | The Charities Regulator's **RCN** |
| `charityTaxExemptionCHY` | Revenue's **CHY** charitable-tax-exemption number — a *different register*, so a different field |
| `regulator`, `regulatorNumber` | Central Bank of Ireland, PSRA, Law Society, Medical/Dental Councils, CORU, NMBI, HIQA, Charities Regulator, CRU, RGII, Safe Electric, the Private Security Authority, the Commission for Aviation Regulation, Fáilte Ireland, SEAI, Chartered Accountants Ireland, the Data Protection Commission — plus its licence or reference number where published |

#### Contact

`email` · `emails[]` · `emailConflict` · `phone` (E.164) · `phones[]` · `addressLine` · `postcode` ·
`postcodeValid` · `city` · `county` · `socialLinks[]` · `officerName` · `officerRole` ·
`termsUrl` · `privacyPolicyUrl`

#### Provenance and accounting

`domain` · `resolvedUrl` · `disclosureUrl` · `disclosureSource` · `discoveryChannel` ·
`pagesParsed` · `pagesChecked` · `fieldCount` · `stableId` · `status` · `missReason` ·
`elapsedMs` · `fetchedAt` · optional `fieldSources` (the exact URL every single field came from)

***

### The CRO number trap — and the field that fixes it

**The CRO runs two registers on one numeric space.** Verified against the register's own public
search API, not assumed:

> CRO number **368047** returns **GOOGLE IRELAND LIMITED** (*LTD – Private Company Limited by
> Shares*) **and THE RETREAT BEAUTY SALON** (*Business name – Body Corporate*).

So a bare CRO number on a website is **ambiguous** unless the page says which register it is on. Some
pages do say — creative-it.ie publishes *"Business Name Registered in Ireland · CRO No. 390462"*,
which is a registered business name, i.e. a sole trader or partnership trading under a name that is
not the owner's own. **s.151 does not even bind a business name.**

`croRegistrationType` reports what the page actually said — `company`, `business-name`, or `null`
when the page does not say. **It is never guessed.** Turn on *"Only numbers on the COMPANY register"*
to keep only the rows the page confirms are companies.

No other Actor in the store ships this distinction.

***

### Accuracy — checked against the register itself

Field-fill percentages cannot tell you whether a number is the *right* number. So every CRO number
this parser emitted from a real Irish corpus was looked up in the **CRO's own public search API** and
compared with the registered name read off the same website:

```
offline, 73 captured domains:  41 of 41 CRO numbers are live register entries   100.0%
live platform run, 400 domains: 37 of 38 are live register entries              97.4%
   the one exception was a HONG KONG company number printed on a global contact
   page, and it is refused by the build that shipped
   name agreement across both: 47 exact or partial, 0 resolving to an unrelated company
```

An earlier build scored 46 of 47 (97.9%) and the seven disagreements were each a *different kind* of
wrong number. Every one is now refused:

| What the page said | What it really was | How it is refused now |
|---|---|---|
| `EPA Registration Number: 72372-1-86703` | a **US EPA** pesticide registration | the digits may not be followed by a hyphen and more digits |
| `registration number 556231-7825` | a **Swedish** organisationsnummer | same guard |
| `registration number 0110-01-140960` | a **Japanese** corporate number | same guard |
| `Register of Architects in Ireland, registration no. 10179` | an **architect's personal** registration | professional-register window veto |
| `Richemont UK Limited … London W1J 5QT … Company Registration Number: 348692` | a **UK** company number | foreign-postal-address window veto |
| `Name of Hosting Company: … Company Number: IE312796` | the hotel's **booking engine** | third-party-disclosure window veto |
| `registered in England and Wales under company number 472968` | an **English** company | foreign-register window veto |

***

### Honest limits

- **Enrichment, not discovery.** It does not find Irish companies; it enriches the domain list you
  supply.
- **No CRO deep link.** The UK sibling emits a one-click register URL; Ireland cannot. Verified
  against the register's own API: `core.cro.ie` keys its company route on an **internal entityId**
  (401323 for CRO number 368047), not on the published number, so a link built from the number would
  open the wrong company or nothing. Shipping one would be a fabricated fact, so the field is cut.
- **No `entityKind`.** A UK Companies House prefix encodes LLP / society / CIO; a bare Irish CRO
  number encodes nothing at all.
- **No data-protection registration number.** Ireland's DPC registration scheme was abolished when
  GDPR applied in May 2018, so there is no Irish equivalent of the UK ICO number to publish.
- **CRO numbers carry no check digit.** `companyNumberFormatValid` is a format and issued-range
  check and is named accordingly. The **VAT** number is genuinely checksum-validated (mod-23, both
  the 8-character and the 9-character published forms).
- **A UK company number in the 1–7 digit range is indistinguishable from a CRO number** when the page
  names no register at all. Measured: one row in the accuracy check (samaritans.ie) was exactly this
  and is now refused by the foreign-address guard — but the general case has no clean signal. Use
  *"Only companies registered in Ireland"* when that matters.
- **A CRO number in the 1900–2099 band is rejected.** A copyright year is the most common false
  positive on an Irish footer. Roughly 200 of the ~780,000 issued numbers fall in that band and are
  lost; that trade was deliberate.
- **`officerName` and `regulatorNumber` are very sparse.** `officerName` was 0.0% on the
  high-street frame and **3.2%** on the larger-sites frame — bigger companies name a director more
  often. `regulatorNumber` was **0.0% on both** live frames, though it is populated on 4 of 73
  offline domains (michaelbarryauctioneers.ie publishes `PSRA LICENCE NO: 004097`,
  smartmoveproperty.ie `PSRA LICENSE NO: 003964`). Both are kept rather than cut because the bytes
  prove they populate — expect them at a few percent at best. Unlike a German Impressum, s.151 does
  not require naming a director at all.
- **Eircode fill is low — 21.4% and 19.8% of billed rows across the two frames.** Eircode only arrived in 2015 and a large share of
  Irish business footers still print an address without one. The address itself is still returned;
  `postcodeValid` tells you whether a published code uses the official Eircode character set (one
  sampled site publishes `R95 C2HO` — the letter `O` is not in the Eircode alphabet, so it cannot be
  a real code, and it is reported rather than silently "corrected").
- **`registeredOffice` is frequently an accountant's address** — that is what was filed.
- **Yield is UK-shaped, not Belgian, and it is a property of your list.** s.151 binds *companies*.
  Like UK reg. 25 and unlike Belgium's art. III.74 CDE, it does not bind sole traders or
  partnerships trading under the owner's own name — they have no CRO number to publish. Measured
  **18.8% and 40.3% of reachable** on two independent frames; see the table below before planning a
  job around either figure.

***

### Which lists convert — measured on TWO independent samples

One corpus measures a parser, not a population. So the coverage numbers below come from **two
unrelated sampling frames**, each 400 Irish domains run on the platform, because a single frame
cannot tell you which of its numbers are facts about Ireland and which are facts about the list:

| | **Frame 1 — OpenStreetMap** | **Frame 2 — Tranco `.ie`** |
|---|---|---|
| selected on | a mapped business with a `website` tag | a domain with measurable traffic rank |
| skews to | high-street retail and hospitality | larger, online-first businesses |
| billed rows | 56 of 400 | **126 of 400** |
| of all attempted | 14.0% | **31.5%** |
| **of reachable** | **18.8%** | **40.3%** |
| host did not resolve | 19.3% | 8.3% |
| refused / challenged | 6.3% | 13.5% |
| published no disclosure | 60.0% | 44.8% |
| needed a second page | 55.4% of rows | 44.4% of rows |

**Your hit rate will land between about 1 in 5 and 2 in 5 of the domains that answer, and which end
depends on your list, not on this Actor.**

Frame 1's 18.8% is the floor, and it is a **composition** figure rather than a limit of the parser.
That frame is **55% retail and 26% hotels and attractions** — 81% of it is exactly the part of the
Irish economy that s.151 does not reach, because a shop or a guesthouse run by a sole trader has no
CRO number to publish at all. Only **15% of it is `office`**, and that subset returned **37.8% of
reachable**, against `tourism` at 7.9%. The same five-fold spread is visible inside one frame, and
frame 2 reaching 40.3% on a traffic-ranked list confirms it from the other direction.

So: a mapped-high-street list sits at the bottom of that range, a list of Irish companies with real
web traffic or a B2B-filtered list sits at the top, and the number to plan around is the one for the
list you actually have.

Two things worth knowing before you plan a job around this:

- **OpenStreetMap `website` tags rot.** Frame 1 lost **19.3%** of its hosts to DNS failures and
  refused connections, against **8.3%** on Tranco. If your list came from a mapping or directory
  export, budget for a fifth of it being stale — those domains are attempted and **never charged**.
- **Bigger sites block more.** Frame 2 was refused on 13.5% against frame 1's 6.3%, and needed
  170 residential escalations against 56. Also never charged.

### What is frame-independent, and what is not

Field-level fill — how often a field is populated **on a row we deliver** — held steady across both
frames, within about five points on almost every field:

```
                     frame 1   frame 2
registeredName         73.2%    75.4%
legalForm              75.0%    76.2%
companyNumber          62.5%    67.5%
jurisdiction           73.2%    78.6%
registeredOfficeEircode 21.4%   19.8%
vatNumber              42.9%    38.1%
email                  94.6%    89.7%
county                 60.7%    54.8%
Irish VAT mod-23 pass  95.2%   100.0%
```

Two fields moved more, and both are composition rather than parsing:

- **`registeredOffice` 51.8% → 69.0%.** Larger companies publish a registered office; a
  high-street shop often publishes only a trading address.
- **`phone` 82.1% → 65.9%.** The reverse — a shop puts its phone in the footer, a software company
  puts a contact form there.

So treat the field percentages as reliable and the **coverage** percentages as a range.

### Where each field actually lives

Measured with `includeFieldSources` on, counting which page each field was read from:

```
registeredOffice / its Eircode, city and county   sub-page  8 of 8   homepage 0
companyNumber and everything derived from it      sub-page  6        homepage 4
registeredName                                    sub-page  6        homepage 2
addressLine / city / county                       sub-page  7        homepage 3
email                                             sub-page  5        homepage 5
phone                                             sub-page  1        homepage 8
```

**The registered office was never once in the homepage footer.** That is why the Actor merges
across pages field by field instead of stopping at the first page that mentions a company number —
a homepage-only run would return the number and the phone and lose the statutory address entirely.

### Billing — you pay for rows, not for attempts

One event, `disclosure-scraped`, **$0.006**, charged through `Actor.pushData(record, event)` so
delivery and billing are atomic. There is **no start fee** and both Apify auto-events are removed.

**Never charged:**

- a host that does not resolve or refuses the connection
- a host that answers with a Cloudflare / challenge page (an HTTP 200 carrying a block page is
  treated as a transport failure, never as "this company publishes nothing")
- a domain that genuinely publishes no s.151 disclosure
- a duplicate (collapsed on the CRO number, else the registrable domain, **before** anything is
  pushed)
- a row your own quality filters removed
- an unbilled coverage row
- a domain abandoned at the per-domain deadline

`RUN_SUMMARY` in the key-value store itemises **the three kinds of nothing** separately — unreachable
host, blocked host, publishes no disclosure — so you can tell a transport problem from an empty
market.

***

### How it works

Two to four hops per domain, plain HTTP (`got-scraping` + `cheerio`, **no browser**), Apify
**DATACENTER** proxy by default with a residential escalation that fires only on a real refusal.

1. **Homepage → footer.** The s.151 particulars sit in the site-wide footer as often as anywhere.
2. **Ranked legal links** off that footer (terms > legal > company information > privacy >
   about/contact), then the XML sitemap, the WordPress page index and finally guessed paths.
3. **Field-by-field merge** across the pages read — Irish sites routinely split the disclosure
   between a terms page and a contact page — recording which URL each field came from.

Every statutory field is read inside a **statement window** anchored on a disclosure phrase, never by
a page-wide scan. That matters far more in Ireland than in the UK: a CRO number is **bare digits**,
so a page-wide numeric grab would return an order number, a price or a phone fragment far more often
than a company number. There is deliberately **no bare-digit fallback of any kind** — every pattern
requires an explicit label (`company number`, `registration no`, `CRO`) or an explicit Irish
jurisdiction clause (`registered in Ireland`).

***

### Inputs worth knowing about

- **Four ways to supply the list** — paste domains, `startUrls` (so Make / Zapier / Clay can hand a
  URL list over natively), a CSV/TXT/JSON file URL, or the dataset ID of a previous Actor run.
- **Two suppression inputs** — `skipDomains`, and `previousDatasetId`, which loads an earlier run of
  this Actor as a do-not-redeliver list so a monthly refresh never re-charges you for a company you
  already bought.
- **Quality filters that run before billing** — require a registry ID, require a checksum-valid VAT,
  require a contact, require the COMPANY register, Irish-registered only, a minimum populated-field
  count, a jurisdiction filter, and a role-email policy. **A filtered-out row is never charged.**
- **`includeMissRows`** turns on full coverage accounting: one unbilled row per attempted domain with
  its `status` and `missReason`.
- **ANY TLD is accepted.** Measured on 6,686 real Irish business domains from OpenStreetMap: **59.2%
  `.ie`, 37.9% `.com`**. Filtering to `.ie` would silently discard a third of the corpus. The TLD
  filter is an opt-in input, never a built-in assumption.

***

### Source and legality

The pages read are the company's own public website: its footer, its terms page, its legal or
company-information page. These are pages published **because a statute requires them to be
published**, and they exist to be read and indexed.

`robots.txt` is **not** fetched by default (it costs an extra request per domain and these pages are
meant to be indexed); a *"Respect robots.txt"* switch is provided for buyers whose own compliance
policy asks for it.

**You are responsible for complying with each site's Terms of Service.** Statutory disclosures are
public by law, but any personal data in the output — a named director, a personal email address — is
yours to handle lawfully under the **GDPR** and the **Data Protection Act 2018**. Ireland's
supervisory authority is the Data Protection Commission.

***

### Not what you are looking for?

- **The CRO register itself** (search by name, filings, officers) — several Actors in the store do
  that. This one reads company *websites*, which is where the contact details are.
- **Emails and phones only, from any website** — `scrapersdelight/website-contact-scraper` does that
  for a fraction of the price.
- **Another jurisdiction** — the same product exists for the DACH *Impressum*, the Dutch KvK
  disclosure, the Italian *note legali*, the UK trading disclosure, the Spanish *aviso legal*, the
  French *mentions légales* and the Belgian *ondernemingsnummer*.

# Actor input Schema

## `domains` (type: `array`):

One entry per company. Accepts all three shapes: a bare domain ("inkplus.ie"), a homepage URL ("https://www.chmarine.com") or a direct legal/terms page URL ("https://x.ie/terms-and-conditions"). ANY TLD is accepted - measured on 6,686 real Irish business domains from OpenStreetMap, only 59.2% are .ie and 37.9% are .com, so filtering to .ie would silently drop a third of the corpus. The Actor finds each site's Companies Act 2014 s.151 disclosure and parses it into one lead per domain. Leave empty to run the built-in Irish demo batch.

## `startUrls` (type: `array`):

The same domain list handed over as URLs, so Make, Zapier, Clay or a Google Sheet can pass it natively (a link to a text/CSV file of URLs also works). Merged with "Irish domains".

## `sourceDatasetId` (type: `string`):

Dataset ID of a previous Actor run. Each item's domain/website column is read and enriched - the real agency workflow: run a directory, Maps or Golden Pages scraper first, then pipe its output here.

## `domainFieldName` (type: `string`):

Which field of the source dataset (or which CSV column of the list file) holds the domain. Leave empty to auto-detect domain / website / websiteUrl / url / site / homepage. Dotted paths like "company.website" work.

## `domainsFileUrl` (type: `string`):

URL of a CSV, TXT, JSON or JSONL file holding the domains - for lists too big to paste into the editor. The column is picked with "Domain field / CSV column".

## `skipDomains` (type: `array`):

Domains to skip outright - accounts you already own, competitors, do-not-contact entries. Matched on the registrable domain, so "www.x.ie/page" and "x.ie" are the same entry.

## `previousDatasetId` (type: `string`):

Dataset ID of an earlier run of THIS Actor. Its stableId / companyNumber / domain values are loaded as a suppression list, so a monthly re-run never re-delivers - and never re-charges you for - a company you already bought.

## `maxItems` (type: `integer`):

Hard cap on rows DELIVERED AND BILLED this run. Distinct from "Max domains": coverage depends on YOUR LIST, so it was measured on two independent 400-domain samples. An OpenStreetMap frame (mapped high-street businesses) billed 56 - 14.0% of all attempted, 18.8% of those that answered. A Tranco .ie frame (domains with real traffic) billed 126 - 31.5% and 40.3%. So 1,000 cold domains yield roughly 140 to 315 billed rows depending on where the list came from. 0 = unlimited.

## `maxDomains` (type: `integer`):

Cap on domains ATTEMPTED, applied before any request. Use it to sample a big list cheaply. 0 = attempt every domain supplied.

## `maxDiscoveryRequestsPerDomain` (type: `integer`):

Hard cap on discovery + confirmation requests per domain, counted AFTER the homepage. This is the knob that stops one slow host burning a minute of a run. Measured average across the whole cascade: 6.6 requests per domain.

## `maxPagesParsed` (type: `integer`):

How many fetched pages may be parsed and merged for one domain. The homepage counts as one. Raising it finds more fields on sites that split the disclosure across terms/contact/privacy pages; lowering it to 1 makes the run homepage-only and very cheap.

## `perDomainTimeoutSecs` (type: `integer`):

Wall-clock deadline for one domain, discovery included.

## `requestConcurrency` (type: `integer`):

How many domains are worked in parallel. Higher is faster; keep it modest to stay polite to small business sites. Memory scales with it, so the Actor lowers it automatically (and says so in the log) if the run does not have the memory for the value you pick.

## `requestTimeoutSecs` (type: `integer`):

Timeout for a single HTTP request.

## `maxRequestRetries` (type: `integer`):

Retries per request, each on a FRESH proxy IP (got-scraping's own retry reuses the flagged IP, which is useless against a soft block). Measured over 400 domains: 2,790 requests, 6.6 per domain.

## `discoveryChannels` (type: `array`):

Which channels may be used to locate the disclosure, tried in this order. MEASURED on two independent 400-domain samples, counting which channel found the winning page: homepage footer 30 and 55, ranked footer links 25 and 69, XML sitemap 1 and 1, guessed paths 0 and 1. Between 44% and 55% of billed rows were only complete once two pages were merged, and the registered office was never once in the homepage footer - so lowering "Max pages parsed" to 1 makes a run cheap but costs you the statutory address. Unlike the German Impressum, Ireland has no conventional page name for the s.151 disclosure. The homepage footer is free (that page is already fetched).

## `followWwwAndRootVariants` (type: `boolean`):

If the homepage fails, retry the other host form (www.x.ie to x.ie and back) before declaring the domain unreachable.

## `deepJsDiscovery` (type: `boolean`):

Last-resort channel for a footer that only exists after client-side render. No Chromium is launched - this build is deliberately browser-free, because a browser image would multiply the compute cost of every run. Instead the Actor reads what a render would have read FROM: the page's inline JSON payloads and its external JavaScript bundles, searching them for a legal-page URL. Costs up to 6 extra requests per domain and is never charged separately.

## `respectRobotsTxt` (type: `boolean`):

Fetch and honour each site's robots.txt before requesting anything. Off by default because it costs one extra request per domain and these pages exist to be indexed. Turn it on if your own compliance policy requires it.

## `proxyConfiguration` (type: `object`):

Proxy settings. Apify DATACENTER is the default and is the measured winner on this corpus: usable transport on the large majority of reachable Irish domains, with a residential retry only when a host actually challenges us.

## `proxyCountry` (type: `string`):

Pin the exit country for sites that geo-tailor their content. MEASURED: Apify datacenter proxies cannot be pinned to a country (they answer HTTP 407), so choosing a country here switches the run to RESIDENTIAL exit nodes in that country. Leave on None for the cheaper, measured-faster datacenter path.

## `escalateToResidentialOnBlock` (type: `boolean`):

On a 403 / challenge (never on a dead host and never on a 404), retry the request once on RESIDENTIAL. Measured on two 400-domain samples: 56 escalations / 16 recovered on a high-street list, and 170 / 63 on a list of larger Irish sites - bigger sites block far more. It only fires on an actual refusal, and a refused domain is never charged.

## `escalateToUnblockerOnBlock` (type: `boolean`):

A SECOND escalation, after residential, for the Cloudflare-managed-challenge tail. Off by default because Unblocker requests are billed to your Apify account on top of the row price.

## `customUserAgent` (type: `string`):

Override the browser User-Agent sent on every request. Leave empty for the built-in Chrome 124 fingerprint.

## `extraHttpHeaders` (type: `object`):

Additional request headers, merged over the defaults (e.g. a From: header identifying your crawler).

## `tldFilterMode` (type: `string`):

No TLD filter is applied by default, deliberately: of 6,686 real Irish business domains sourced from OpenStreetMap, 59.2% are .ie and 37.9% are .com. Filtering to .ie would discard a third of the corpus, including many of the larger companies. Use this only when your own list really is single-TLD.

## `tldFilter` (type: `array`):

The TLD list the mode above applies to, without the dot: ie, com, eu, org.

## `irishRegisteredOnly` (type: `boolean`):

Drop rows whose disclosure names a register OUTSIDE the Republic. Measured on 73 Irish-corpus domains: 8 of them (cotswoldoutdoor.ie, quick-step.ie, boodles.com and others) are .ie or Irish-facing storefronts operated by companies registered in England, Wales, Scotland or Northern Ireland. Their UK company number is never emitted in the CRO field, and this switch removes the row entirely.

## `requireRegistryId` (type: `boolean`):

Deliver (and bill) only companies whose disclosure carries a CRO number, a VAT number or a charity number.

## `requireCompanyRegister` (type: `boolean`):

The CRO runs two registers on ONE numeric space - companies and registered business names - and this is verified, not assumed: CRO number 368047 returns both GOOGLE IRELAND LIMITED (a company) and THE RETREAT BEAUTY SALON (a business name). Turn this on to keep only rows where the page itself says the number is a company registration. A business name is a sole trader or partnership trading under a name that is not the owner's own; s.151 does not bind it.

## `requireValidVat` (type: `boolean`):

Deliver (and bill) only rows whose Irish VAT number passes the published mod-23 check character. A row with no VAT number, or one that fails the check, is filtered out here and never billed.

## `requireContact` (type: `boolean`):

Deliver (and bill) only companies with an email or a phone number.

## `jurisdictionFilter` (type: `array`):

Keep only companies whose disclosure names one of these registers. Ireland is inferred from a CRO number or from the wording ("registered in Ireland"); the UK jurisdictions appear when an Irish-facing site is operated by a British company. Leave empty for all.

## `minFieldsRequired` (type: `integer`):

Quality floor: a row must carry at least this many of the 22 value fields before it is delivered and billed. 0 = no floor. Measured median on real Irish domains: 10.

## `emailPolicy` (type: `string`):

Agencies split hard on whether info@ / enquiries@ counts as a lead. "Role only" keeps just those; "Exclude role" keeps only named mailboxes.

## `dedupeBy` (type: `string`):

Which key collapses duplicates BEFORE anything is pushed or charged. The default uses a well-formed CRO number when the page carries one and the registrable domain otherwise, so two domains owned by the same company collapse to one billed row - and a repeated input line is dropped at queue time, before it costs a request.

## `validateTaxId` (type: `boolean`):

Run the published Irish mod-23 check character on every VAT number found, covering both the 8-character form (IE6388047V) and the newer 9-character form (IE3206488LH), and emit vatNumberValid plus which variant it passed. Note that CRO numbers carry NO check digit at all, so companyNumberFormatValid is a format and issued-range check, never a checksum.

## `extractRegulator` (type: `boolean`):

The Irish regulatory body a professional-services site must name (Central Bank of Ireland, PSRA for estate agents, Law Society, Medical and Dental Councils, CORU, HIQA, the Charities Regulator, RGII, Safe Electric, the Private Security Authority and others) plus its licence or reference number where published.

## `extractCharity` (type: `boolean`):

Ireland keeps TWO charity identifiers on two different registers and they get two fields: the Charities Regulator's RCN, and Revenue's older CHY charitable-tax-exemption number.

## `extractOfficer` (type: `boolean`):

A director, proprietor or company secretary named next to a label on the page. Honest expectation: Irish sites rarely publish one - 0.0% of billed rows on a high-street sample and 3.2% on a sample of larger Irish sites. Unlike a German Impressum, s.151 does not require it at all.

## `extractSocials` (type: `boolean`):

LinkedIn, Facebook, Instagram, X, YouTube and TikTok links present on the pages read.

## `extractPolicyUrls` (type: `boolean`):

The site's terms and privacy-policy URLs, linked from the pages read.

## `includeMissRows` (type: `boolean`):

Emit a row for every domain that produced no lead - dead host, blocked, publishes no disclosure, filtered out, a duplicate, or suppressed - with its status and missReason, so you can do coverage accounting. These rows are NEVER charged.

## `includeFieldSources` (type: `boolean`):

Add a fieldSources object naming the exact URL each field was read from. Useful when a disclosure is split across a terms page and a contact page and you need to audit which said what.

## `flattenOutput` (type: `boolean`):

On: one flat row, ready for Google Sheets, Clay or a CSV export. Off: fields grouped into company {}, registry {}, contact {}, people {} and policies {} objects.

## Actor input object example

```json
{
  "domains": [
    "patrickbourkemenswear.ie",
    "pifs.ie",
    "enableireland.ie",
    "inkplus.ie",
    "smartmoveproperty.ie",
    "chmarine.com",
    "carrollsirishgifts.com",
    "lombre.ie",
    "creative-it.ie",
    "castle-hotel.ie"
  ],
  "maxItems": 1000,
  "maxDomains": 0,
  "maxDiscoveryRequestsPerDomain": 8,
  "maxPagesParsed": 3,
  "perDomainTimeoutSecs": 60,
  "requestConcurrency": 10,
  "requestTimeoutSecs": 25,
  "maxRequestRetries": 2,
  "discoveryChannels": [
    "homepage",
    "anchor",
    "sitemap",
    "wpJson",
    "pathGuess"
  ],
  "followWwwAndRootVariants": true,
  "deepJsDiscovery": false,
  "respectRobotsTxt": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "proxyCountry": "none",
  "escalateToResidentialOnBlock": true,
  "escalateToUnblockerOnBlock": false,
  "tldFilterMode": "none",
  "irishRegisteredOnly": false,
  "requireRegistryId": false,
  "requireCompanyRegister": false,
  "requireValidVat": false,
  "requireContact": false,
  "jurisdictionFilter": [],
  "minFieldsRequired": 0,
  "emailPolicy": "all",
  "dedupeBy": "company-number-then-domain",
  "validateTaxId": true,
  "extractRegulator": true,
  "extractCharity": true,
  "extractOfficer": true,
  "extractSocials": true,
  "extractPolicyUrls": true,
  "includeMissRows": false,
  "includeFieldSources": false,
  "flattenOutput": true
}
```

# Actor output Schema

## `records` (type: `string`):

The dataset of Irish company leads (one item per domain that published a Companies Act 2014 s.151 website disclosure).

## `runSummary` (type: `string`):

Coverage accounting for the run: domains attempted, rows billed, and the three kinds of nothing (unreachable host, blocked host, publishes no disclosure) kept apart - none of them charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "patrickbourkemenswear.ie",
        "pifs.ie",
        "enableireland.ie",
        "inkplus.ie",
        "smartmoveproperty.ie",
        "chmarine.com",
        "carrollsirishgifts.com",
        "lombre.ie",
        "creative-it.ie",
        "castle-hotel.ie"
    ],
    "maxItems": 1000,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/ie-company-website-disclosure-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "patrickbourkemenswear.ie",
        "pifs.ie",
        "enableireland.ie",
        "inkplus.ie",
        "smartmoveproperty.ie",
        "chmarine.com",
        "carrollsirishgifts.com",
        "lombre.ie",
        "creative-it.ie",
        "castle-hotel.ie",
    ],
    "maxItems": 1000,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/ie-company-website-disclosure-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "patrickbourkemenswear.ie",
    "pifs.ie",
    "enableireland.ie",
    "inkplus.ie",
    "smartmoveproperty.ie",
    "chmarine.com",
    "carrollsirishgifts.com",
    "lombre.ie",
    "creative-it.ie",
    "castle-hotel.ie"
  ],
  "maxItems": 1000,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scrapersdelight/ie-company-website-disclosure-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/ie-company-website-disclosure-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6menJEDWtoLDPbfaN/builds/QNDagzZ9EdxtHnJgz/openapi.json
