# Italian Note Legali Scraper - P.IVA, REA, PEC & Contacts (`scrapersdelight/it-note-legali-contact-scraper`) Actor

From $8 per 1,000 leads, no start fee. Turn Italian company domains into registry-grade B2B leads from each site's statutory D.Lgs. 70/2003 disclosure: checksum-validated P.IVA, codice fiscale, REA + CCIAA, capitale sociale, PEC, email, phone and sede legale. No login, no API key.

- **URL**: https://apify.com/scrapersdelight/it-note-legali-contact-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$8.00 / 1,000 per italian company lead returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 🇮🇹 Italian Note Legali Scraper — company leads with a validated P.IVA

Paste a list of **Italian company domains** (or point this at another Actor's dataset) and get back **one B2B/KYB lead per domain**, read off the statutory disclosure every Italian commercial website is legally required to publish: the **ragione sociale**, the **partita IVA checked against the official DPR 633/72 check digit**, the **codice fiscale**, the **REA entry and its Camera di Commercio**, the **capitale sociale**, the **sede legale** split into street / CAP / comune / provincia, and a contact **email, PEC and phone** — for **$8 per 1,000 leads, no start fee**.

### The one thing to know before you buy: what you feed it decides what you get

Measured on **1,550 real Italian domains**, drawn four different ways on purpose:

| What you feed it | Domains tested | **Comes back with a validated P.IVA** |
|---|---:|---:|
| A business-directory list (PagineGialle trades, restaurants, clinics, shops) | 253 | **90.5 %** |
| Mid-cap Italian brands | 148 | 57.4 % |
| An OpenStreetMap SME export (31.5 % of it was already dead or walled) | 699 | 46.5 % — but **68.6 %** of the hosts that answered |
| An undifferentiated `.it` domain list | 450 | **32.7 %** |

Feed it trading companies and roughly **9 domains in 10** return a checksum-validated partita IVA.
Feed it a raw `.it` list and it is closer to **1 in 3**, because much of that tail is blogs,
associations, public bodies, media and parked domains with no art. 7 obligation at all. Everything
that does not resolve comes back **free** — see the full breakdown below.

***

**Why a statutory disclosure and not a directory?** Because **D.Lgs. 9 aprile 2003 n. 70, art. 7** — Italy's transposition of the E-Commerce Directive — obliges every provider of information-society services to make accessible *"in modo facile, diretto e permanente"* its name or denominazione, the sede where it is established, an email address allowing rapid contact, its **Registro delle Imprese / REA** entry and its **partita IVA**. **Art. 2250 Codice Civile** adds the registro office and number, the capitale sociale with whether it is *interamente versato*, and single-shareholder status. The corpus is therefore **the whole Italian commercial web**, not one directory's member list — and every field is something the company itself published about itself, by law.

***

### ✅ What it does

Two things, in this order, and the order is the product:

1. **Reads the site-wide footer first.** Italian practice is not German practice. A German site links an *Impressum* and puts everything on it; an Italian site prints `P.IVA 01234567890` in the footer of **every page** and very often has **no dedicated legal page at all**. Measured over 1,550 real Italian domains: of the **786 validated partite IVA** this Actor extracted, **747 — 95.0 % — came off the homepage**. A crawler that only follows a "Note legali" link finds the other 39.
2. **Then merges the dedicated page, when there is one.** The richer art. 2250 fields — REA, CCIAA, capitale sociale — and the **PEC** live on *Note legali*, *Dati societari*, or, very commonly, inside the **GDPR privacy policy**, which must name the *titolare del trattamento*. Pages the site actually links are tried **before** any guessed path, and the page carrying the validated P.IVA owns the identity fields.

#### The five things a generic imprint scraper gets wrong here

- 🧮 **It validates the P.IVA, offline.** The full DPR 633/72 check digit — 11 digits, odd positions summed, even positions doubled and reduced by 9, total a multiple of 10. **Zero external calls.** An analytics id, an order number or a stray 11-digit run will not validate, so an unvalidated string never ships as a P.IVA. **All 786 numbers** extracted in the benchmark below passed, and on every captured page the suite also asserts the reverse — that no labelled, checksum-valid P.IVA present in the bytes was missed. The 16-character **codice fiscale** is separately checked against its own DM 23/12/1976 check character, and one that fails is **returned flagged**, never silently dropped.
- 🏷 **It ties the company name to the id it actually took.** A single Italian legal block routinely names several entities — the operator, its parent, the hosting provider, a separate data-protection contractor. Taking the first company-form match on the page produced visibly wrong owners in testing (`illy.com` returned a cleaning-services supplier from further down the page; `campari.com` returned `SOUTH AFRICA Spa` out of a subsidiary list). The name is the form-bearing phrase **nearest the fiscal id**, and the address is the one nearest it too — never a hosting provider's street.
- 📛 **It tells you where the name came from.** `ragioneSocialeSource` is `legal-form` (a real denominazione ending in S.r.l. / S.p.A. / S.n.c. …), `copyright` (the footer signature), `og-site-name` or `page-title`. A 99.6 % "company name" fill rate that quietly includes page titles is a misleading number; **397 of 995 billable rows carry a true legal-form name**, and the column says so on every row.
- 🗺 **It resolves the provincia, it does not guess it.** A printed `(BS)` **and** a spelled-out `(Brescia)` both resolve, against the closed set of 107 province sigle. A CAP is **never** used to infer a province — that would be a guess wearing the costume of data. Where the page does not print one, `provincia` is `null`.
- 💸 **It charges only for a lead it delivers.** A dead host, an anti-bot wall, a site that publishes no disclosure, a row under your quality bar, a filtered row, a duplicate, a domain the run never reached — all are returned to you (turn on **Include unbilled miss rows**) and **none is ever charged**. On the shipped build's 300-domain live run below, **211 rows were delivered and exactly 211 events were charged**; the other 89 domains cost nothing.

***

### 📊 Measured, from real runs

Everything below is from actual runs on real Italian domains, through the shipping transport. Nothing is estimated.

#### The corpus

**1,550 real Italian domains**, drawn four different ways on purpose, so no number is flattered by a hand-picked list:

| Corpus tier | Domains | Reachable | Resolved a disclosure | **Validated P.IVA** |
|---|---:|---:|---:|---:|
| **PagineGialle SMB** — plumbers, restaurants, clinics, body shops, bakeries, across 20 cities | 253 | 250 (98.8 %) | 250 (98.8 %) | **229 (90.5 %)** |
| **OpenStreetMap SMEs** — shops, offices and trades carrying a `website` tag | 699 | 479 (68.5 %) | 474 (67.8 %) | **325 (46.5 %)** |
| **Mid-cap Italian brands** | 148 | 102 (68.9 %) | 100 (67.6 %) | **85 (57.4 %)** |
| **Tranco `.it` tail**, ranks 60k–960k — mixed, **including non-commercial sites** | 450 | 396 (88.0 %) | 392 (87.1 %) | **147 (32.7 %)** |
| **All** | **1,550** | **1,227 (79.2 %)** | **1,216 (78.5 %)** | **786 (50.7 %)** |

**Read that table before you buy. What you feed this Actor decides what you get.** On a directory list of *actual trading Italian businesses* it returns a validated P.IVA on **9 domains in 10**. On an undifferentiated `.it` list it returns one on **1 in 3**, because much of the `.it` tail is blogs, associations, public bodies, media and parked domains with no art. 7 obligation.

The OpenStreetMap row is lower than the PagineGialle row for a reason worth knowing if you build lists that way: **OSM `website` tags rot.** 31.5 % of those hosts were dead or walled on the day of the run — but of the ones that *answered*, **68.6 % carried a validated P.IVA**, right in line with the other business tiers.

A separate 260-domain check on **northern** OSM entries (Lombardia, Veneto, Liguria, Trentino-Alto Adige, which the main sample under-represents) resolved **199 of 260** and found a validated P.IVA on **153 — 76.9 % of the resolved rows**, versus 68.6 % for the centre-and-south sample. So the headline table is the **conservative** read, not the flattering one.

#### Coverage over all 1,550 domains

| Outcome | Domains | Share | Charged? |
|---|---:|---:|---|
| ✅ resolved a disclosure | 1,216 | 78.5 % | only if it meets your quality bar |
| 🔴 `dead` — host does not resolve (NXDOMAIN / TLS) | 189 | 12.2 % | never |
| 🟠 `blocked` — anti-bot wall on every attempt | 87 | 5.6 % | never |
| ⚪ `unreachable` — transport failed after retries | 47 | 3.0 % | never |
| ⚫ `no_disclosure_found` — reachable, nothing readable at all | 11 | 0.7 % | never |

With the **shipped default quality bar** (`minFieldsRequired: 2`), **995 of the 1,550 domains (64.2 %) produced a billable lead**. Everything else came back free.

**The number that matters most for expectation-setting:** of the **1,227 reachable** domains, **441 (35.9 %) published no partita IVA this Actor could read.** Only **11** of those 441 published *nothing at all*; the rest published something — a name, an email, a phone, an address — just not the fiscal id. They split three ways: **209** still cleared your quality bar on their other fields and are delivered and charged as ordinary leads, **221** fell under it and come back `partial` and free, and **11** come back `no_disclosure_found` and free. If you only want rows carrying a validated P.IVA, turn on **Only rows with a checksum-valid P.IVA** and the other 430 stop being billable too.

#### Per-field fill over the 995 billable leads

| Field | Fill | | Field | Fill |
|---|---:|---|---|---:|
| `ragioneSociale` | **99.6 %** (991/995) | | `provincia` | 31.0 % (308/995) |
| `email` | **89.0 %** (886/995) | | `codiceFiscale` | 21.1 % (210/995) |
| `partitaIva` + `partitaIvaValid` | **79.0 %** (786/995) | | `pec` | 15.4 % (153/995) |
| `phone` + `phoneE164` | 76.2 % (758/995) | | `reaNumber` | 15.2 % (151/995) |
| `socialProfiles` | 75.0 % (746/995) | | `cciaa` | 11.8 % (117/995) |
| `sedeLegale` / `street` / `cap` / `comune` | 66.4 % (661/995) | | `reaProvince` | 11.1 % (110/995) |
| `formaGiuridica` | 49.2 % (490/995) | | `capitaleSociale` | 9.4 % (94/995) |
| `vatNumber` (IT + P.IVA) | 79.0 % (786/995) | | `capitaleSocialeVersato` | 6.5 % (65/995) |
| `emails` (all addresses) | 89.0 % | | `direttoreResponsabile` | 1.2 % (12/995) |
| `ragioneSocialeSource` | 99.6 % | | `legaleRappresentante` | **1.2 %** (12/995) |
| `stableId` · `status` · `fetchedAt` · `domain` · `website` · `sourceUrl` · `sourceType` · `discoveryMethod` · `pagesRead` · `fieldCount` | 100 % | | `registrazioneTribunale` | **0.1 %** (1/995) |

`privacyPolicyUrl`, `noteLegaliUrl` and `missReason` fill only when the site links one / the row is a miss.

#### Is a low fill rate our parser, or the page? — measured, not assumed

A raw fill rate mixes two completely different things: a field the page never published, and a field our parser walked past. A regex that silently matches nothing produces plausible nulls, not errors, so the two are indistinguishable in the output. They are separated here by measuring **label-present → value-emitted**: of the pages whose raw bytes carry a label for a field, how often did a value actually come out? Measured over **160 pages captured through the shipping proxy**, with label detectors written independently of the extractors they audit:

| Field | Pages carrying the label | Value emitted | **Conversion** |
|---|---:|---:|---:|
| `partitaIva` | 112 | 112 | **100.0 %** |
| `capitaleSociale` | 22 | 22 | **100.0 %** |
| `pec` | 20 | 19 | 95.0 % |
| `phone` | 99 | 94 | 94.9 % |
| `formaGiuridica` | 77 | 72 | 93.5 % |
| `email` | 105 | 98 | 93.3 % |
| `cciaa` | 25 | 23 | 92.0 % |
| `reaNumber` | 32 | 28 | 87.5 % |
| `codiceFiscale` | 41 | 34 | 82.9 % |
| `sedeLegale` | 122 | 94 | 77.0 % |
| `legaleRappresentante` | 3 | 0 | see below |

**Every P.IVA that is on the page comes off the page — 112 of 112.** The 79 % fill rate above is therefore a fact about Italian websites, not a limit of this Actor. The same holds for the capitale sociale. The suite also asserts the **negative direction** on every captured page: if no P.IVA was emitted, the raw bytes must contain no labelled, checksum-valid one that was missed — **0 pages failed that**. The checksum used in that test is a second, independently written implementation of the DPR 633/72 algorithm, cross-checked against the shipping one over 20,000 numbers, because a validator that imports the code it is checking proves only that the code equals itself.

The three `legaleRappresentante` pages were read by hand: two are the boilerplate *"in persona del suo legale rappresentante pro tempore"* and one is a news article about a bank's CEO. None names the site operator's representative, so 0 of 3 is the correct answer, not a miss.

#### The registry fields are a function of company size, not of effort

| Field | OSM SMEs | PagineGialle SMB | Mid-cap | Tranco tail |
|---|---:|---:|---:|---:|
| `partitaIva` | 76 % | **92 %** | 90 % | 65 % |
| `sedeLegale` | 66 % | 76 % | 73 % | 53 % |
| `email` | 89 % | **97 %** | 88 % | 81 % |
| `phone` | 78 % | **100 %** | 76 % | 47 % |
| `formaGiuridica` | 56 % | 15 % | **86 %** | 58 % |
| `reaNumber` | 15 % | 4 % | **38 %** | 19 % |
| `cciaa` | 12 % | 1 % | **36 %** | 14 % |
| `codiceFiscale` | 24 % | 2 % | **37 %** | 29 % |
| `capitaleSociale` | 8 % | 2 % | **27 %** | 14 % |
| `pec` | 18 % | 5 % | **24 %** | 18 % |

A one-person *ditta individuale* has no società form, no share capital and usually no REA line in its footer — it prints its P.IVA and its phone number, and this Actor returns both at 92 % and 100 %. An S.p.A. prints the whole art. 2250 block. **If you are buying this for REA, CCIAA and capitale sociale, feed it companies, not sole traders.**

**`legaleRappresentante` at 1.2 % and `registrazioneTribunale` at 0.1 % are the statute, not the parser.** Art. 7 D.Lgs. 70/2003 requires the *name or denominazione of the provider*; it does **not** require naming the person who represents it — which is why a German Impressum carries a Geschäftsführer (§5 TMG requires the *Vertretungsberechtigter*) and an Italian footer usually does not. **If you need a director, this is the wrong source; you want a Registro Imprese lookup.**

#### Where the data came from

| | |
|---|---:|
| P.IVA found on the **homepage / site-wide footer** | **747 of 786 (95.0 %)** |
| P.IVA found only on a **dedicated page** | 39 (5.0 %) — privacy ×16, contatti ×8, legal-notice ×7, note-legali ×5, termini ×2 |
| Disclosure resolved from the homepage alone | 797 domains |
| Enriched by a page the site **linked** | 404 domains |
| Enriched by a **guessed path** | 15 domains |
| Pages read | 4,762 total, **3.07 per domain** |
| Requests issued, misses included | 8,869 total, **5.72 per domain** |
| Unique P.IVA values | 776 of 786 (10 domains shared a company) |
| Name provenance | `legal-form` 397 · `page-title` 314 · `copyright` 158 · `og-site-name` 122 |

#### Transport

Datacenter proxy, fresh exit IP on every retry, one escalation to residential-IT on a genuine block. Measured on the escalation ladder against 8 mid-cap domains:

| Rung | Domains resolving a P.IVA |
|---|---|
| Home IP (direct) | 8 / 8 |
| Apify **DATACENTER** | 7 / 8 — `barilla.com` answered HTTP 403 with a 5.9 KB body on four consecutive fresh IPs |
| Apify **RESIDENTIAL + country-IT** | **8 / 8** — the same `barilla.com` returned 200 / 853,552 bytes and parsed cleanly |

So datacenter is the default and residential is the **escalation**, not a blanket upgrade you pay for. Over the 1,550-domain run: 9,300 requests, 224 escalated to residential-IT, **54 recovered**. Average usable page: 163 KB.

#### Live run on the platform — delivered vs charged, and cost

Runs of the published build at its default 1,024 MB, both counters polled to stability before either number was read:

| | Small run | 300-domain run | **300-domain run, shipped build** |
|---|---:|---:|---:|
| Domains in | 5 | 300 | 300 |
| Rows delivered | 5 | 300 | 300 |
| Rows with `status: ok` | 3 | 220 | **211** |
| **`chargedEventCounts`** | **3** | **220** | **211** |
| Delivered `ok` == charged | ✅ | ✅ | **✅** |
| Duration | 25.2 s | 585.1 s | 605.1 s |
| Peak memory | 175 MB | 696 MB | **703 MB** of 1,024 |
| Platform usage | $0.002998 | $0.069221 | $0.072864 |

**Marginal cost, from the slope between the small run and the shipped build's run:**
($0.072864 − $0.002998) ÷ (211 − 3) = **$0.000336 per delivered lead**, i.e. **$0.34 per 1,000**.
At **$0.008 per lead** that is a **95.8 % margin**, and the price is a pricing decision rather than
an efficiency one — there is no browser anywhere in the default path.

On the shipped build's run the other 89 domains came back free: 26 blocked, 22 dead, 15 timed out,
14 partial, 8 unreachable, 2 with no disclosure, 2 filtered.

***

### 🧾 Output

One row per input domain.

| Field | Description |
|---|---|
| `domain` | Registrable domain (eTLD+1) — one domain = one row = at most one charge. |
| `website` | The company homepage. |
| `sourceUrl` | The page the identity fields came from. |
| `sourceType` | `homepage` · `note-legali` · `legal-notice` · `privacy` · `cookie-policy` · `contatti` · `termini` · `chi-siamo`. |
| `discoveryMethod` | `homepage` · `footer-link` · `path-guess`. |
| `ragioneSociale` | The company name. |
| `ragioneSocialeSource` | **Where that name came from**: `legal-form` · `copyright` · `og-site-name` · `page-title`. Only `legal-form` is a true denominazione. |
| `formaGiuridica` | `S.p.A.` · `S.r.l.` · `S.r.l.s.` · `S.a.s.` · `S.n.c.` · `S.a.p.a.` · `S.c.a r.l.` · `Soc. Coop.` · `S.S.D.` · `A.S.D.` · `S.s.` |
| `partitaIva` | 11-digit partita IVA, normalised. |
| `partitaIvaValid` | **Official DPR 633/72 check digit.** `true` on every number this Actor returns — an id that fails is not returned at all. |
| `vatNumber` | The intra-EU form, `IT` + the P.IVA, ready for a VIES call. |
| `codiceFiscale` | Company CF (11 digits, usually identical to the P.IVA) or a ditta individuale's 16-character CF. |
| `codiceFiscaleType` | `company` or `individual` — so you never have to guess which you got. |
| `codiceFiscaleValid` | For the 16-character form, the DM 23/12/1976 check character. A failing id is **flagged, not dropped**. |
| `reaNumber` | Repertorio Economico Amministrativo number. |
| `reaProvince` | The CCIAA province named on the REA entry. |
| `cciaa` | The Camera di Commercio / Registro Imprese office holding the company file. |
| `capitaleSociale` | Share capital as a **number**, parsed with Italian separators (`1.000.000,00` → `1000000`). Never a served zero. |
| `capitaleSocialeVersato` | `true` when the page says *interamente versato*, `false` when it says only partly, `null` when it does not say. |
| `sedeLegale` | The registered-office line. |
| `street`, `cap`, `comune`, `provincia` | The same address split up. `provincia` is a real sigla from the closed set of 107, or `null`. |
| `email` | Contact address, role mailboxes preferred (`info@`, `contatti@`, `amministrazione@`, `commerciale@`…). |
| `emails` | Every address found on the pages read, deduplicated. |
| `pec` | **Posta elettronica certificata** — the company's *domicilio digitale*, a legally binding address for service. Recognised by host (`legalmail.it`, `pec.it`, `arubapec.it`, `postecert.it`, `@pec.…`) and never mixed into `email`. |
| `phone` | Phone as the page printed it. |
| `phoneE164` | The same number normalised to `+39…`, ready for a CRM import. |
| `legaleRappresentante` | Legale rappresentante / amministratore unico or delegato. **2.0 % fill — read the note above.** |
| `direttoreResponsabile` | Direttore responsabile, for a site registered as a testata giornalistica. 2.4 %. |
| `registrazioneTribunale` | That site's tribunal registration under L. 47/1948 art. 5 / L. 62/2001. 0.2 %. |
| `socialProfiles` | LinkedIn, Facebook, Instagram, X, YouTube, TikTok, Pinterest **profile** links. Content URLs (a `/watch?v=`, an `/embed/`, a share intent) are excluded, so three links to one marketing video do not import as three profiles. |
| `privacyPolicyUrl`, `noteLegaliUrl` | Those pages, when the site links them. |
| `pagesRead` | How many pages this row cost. |
| `fieldCount` | How many of the 19 value fields are populated — this is what your quality bar is measured against. |
| `stableId` | `IT` + the validated P.IVA when there is one, else the domain. The key to deduplicate and re-key against. |
| `status` | `ok` · `partial` · `no_disclosure_found` · `blocked` · `dead` · `unreachable` · `error` · `filtered` · `not_attempted`. **Only `ok` is ever charged.** |
| `missReason` | Plain-language reason a non-`ok` row is not a lead. |
| `fetchedAt` | ISO timestamp. |
| `sourceText` | Only with **Include the cleaned disclosure text** on — the page text, capped at 40,000 characters, so you can audit a field or re-parse it yourself. |

Two dataset views ship with it: **Italian company leads** (the flat lead) and **Registry / KYB view** (the Registro Imprese identity). A `RUN_SUMMARY` record in the key-value store reconciles every domain you submitted against what was delivered and what was charged.

***

### ⚙️ Input

Bring your domains any of five ways — they are merged, normalised to one entry per registrable domain, and deduplicated before a single request goes out:

| Input | What it is |
|---|---|
| **Italian company domains** | `["bialetti.com", "https://www.lavazza.it", "esempio.it/note-legali"]`. Bare domains, homepage URLs and direct Note legali URLs all work. |
| **Start URLs** | Apify's Start-URLs shape, so the Console link-list UI, a file upload, Make, Zapier or a Google Sheet can pass the list natively. |
| **Enrich an existing dataset** | Dataset ID of a previous run — the real agency workflow: run PagineGialle or Google Maps Italy first, then pipe its `website` column here. |
| **Domain list file URL** | A public CSV / TXT / JSON / JSONL for lists too large to paste. |
| **Suppression list** | Applied **before** any fetch, so a suppressed domain never costs a request and can never be billed. |

Then the knobs worth knowing:

- **Minimum populated fields** *(default 2)* — the quality floor **and** the billing bar. A row under it comes back `partial`, free. You set your own standard for what counts as a lead.
- **Only rows with a checksum-valid P.IVA** — the strictest bar there is: no valid P.IVA, no charge.
- **Max extra pages per domain** *(default 6)* — set it to `0` for homepage-only, which is the cheapest setting and still yields 95.0 % of all the P.IVA finds. Raise it when the registry fields matter more than request count.
- **Only these provinces / CAP prefixes / legal forms** — `provinciaFilter: ["MI","RM"]`, `capFilter: ["20","00"]`, `formaGiuridicaFilter: ["S.p.A."]`. Filtered rows are never charged.
- **Deduplicate by** — the default collapses on the validated P.IVA when there is one and the domain otherwise, so two domains owned by the same company become one billed row.
- **Escalate to residential-IT on a block** *(on)* — see the transport table.
- **Include unbilled miss rows** — coverage accounting for every domain you submitted.

```json
{
  "domains": ["bialetti.com", "lavazza.it", "moleskine.com"],
  "minFieldsRequired": 2,
  "provinciaFilter": ["MI", "TO"],
  "includeMissRows": true
}
```

***

### 💵 Pricing

**Pay per event — $0.008 per delivered lead ($8 per 1,000). No start fee.**

- You are charged **only** for a row that resolved a disclosure **and** met your `minFieldsRequired` bar (`status: ok`).
- `dead`, `blocked`, `unreachable`, `no_disclosure_found`, `partial`, `filtered`, duplicate and `not_attempted` rows are still returned so you can see what happened to every domain you submitted — and **none of them is ever charged**.
- Rows are **delivered and billed atomically** (`pushData(row, event)`), so a charge cap truncates delivery too. You can never be billed for a row you did not receive, and never receive a row you were not billed for. Verified on both live runs above: 3 = 3 and 220 = 220.
- There is **one** charge event. No run-start fee, no per-dataset-item fee, no browser surcharge.

***

### 🚧 Honest limits

**This is an enrichment tool, not a discovery tool.** It does not find Italian companies for you; it turns a domain list you already have into legal-entity records.

**Feed it businesses.** The corpus table above is the whole story: 90.5 % P.IVA on a real business-directory list, 32.7 % on an undifferentiated `.it` list. A blog, an association, a comune's website or a parked domain has nothing to give. And 35.9 % of reachable domains publish no P.IVA at all — those come back free.

**Four classes of domain do not resolve, and all four are free.** Measured on the 1,550-domain benchmark:

| Class | Share | What happens |
|---|---:|---|
| Dead host — NXDOMAIN, TLS failure, connection refused | 12.2 % | `dead`. Identical from every exit, so this is a stale list entry, never reported as a block. Directory and OSM lists rot fast: 31.5 % of the OpenStreetMap tier was dead or walled. |
| Enterprise CDN / WAF | 5.6 % | `blocked`. 403 or a challenge interstitial on repeated fresh datacenter IPs; residential-IT recovered 54 of 224 escalations. Re-running a blocked list later costs nothing, because blocked rows are free. |
| Transport failure after retries | 3.0 % | `unreachable`. |
| Reachable but publishes nothing readable at all | 0.7 % | `no_disclosure_found`. A further 35.2 % of reachable domains publish something but no P.IVA — those return `partial`, also free. |

**The mid-cap tier is the hard one.** Small Italian business sites are plain HTML and resolve at 98.8 %; large brands sit behind Akamai, Cloudflare and Incapsula, which is why that tier reached only 68.9 % — and why residential escalation exists. An Apify datacenter IP being refused by a brand's CDN is measured and reported per row; it is never presented as "this company publishes nothing". You can also point the **Proxy** input at your own pool.

**The registry fields are sparse because of who publishes them, not because they are hard to read.** REA 15.2 %, CCIAA 11.8 %, capitale sociale 9.4 %, PEC 15.4 % over the billable rows — but see the per-tier table above: on the mid-cap tier those are 38 %, 36 %, 27 % and 24 %, and on sole traders they are 4 %, 1 %, 2 % and 5 %. Where the label IS on the page, the parser converts it 87.5 %, 92.0 %, 100 % and 95.0 % of the time, so this is a property of Italian footers, not of the extraction. **If you need the full Registro Imprese file for every company, you want a registry lookup, not a website read.**

**`partitaIvaValid` proves the number is well-formed, not that the company is trading.** The DPR 633/72 check digit is arithmetic, not a registry call. For live registration status, VIES or a Registro Imprese source is a different data path and a different Actor.

**A 99.8 % company-name fill rate is not 99.8 % legal names.** `ragioneSocialeSource` splits it honestly: `legal-form` 397, `page-title` 314, `copyright` 158, `og-site-name` 122 of the 995 billable rows. A `page-title` name can be a tagline ("Rivendita accessori bagno") rather than a denominazione. Filter on `ragioneSocialeSource == "legal-form"` when you need the real thing.

**No SMTP.** Nothing here says a mailbox exists. A PEC address is by construction a real, legally registered mailbox; an ordinary `info@` is whatever the site published.

**A slow domain cannot eat your run.** Each domain gets its own wall-clock deadline (`perDomainTimeoutSecs`, default 60 s) with a hard outer cut at 75 s, and the run stops starting new domains at 80 % of its own timeout, returning the rest as `not_attempted` — free — with a warning naming the count. On the shipped build's 300-domain live run, 15 domains hit that cut and were returned free as `error`.

**A walled run fails loudly.** If more than 70 % of attempted domains are blocked on a run of 10 or more, the Actor calls `Actor.fail()` with the count rather than reporting a proxy problem as "these sites publish nothing". Nothing extra is charged.

***

### ⚖️ Legal & fair use

The data this Actor reads is a **statutory public disclosure**, mandated by **D.Lgs. 9 aprile 2003 n. 70, art. 7** (attuazione della direttiva 2000/31/CE sul commercio elettronico) and, for company acts, **art. 2250 del Codice Civile**. The company is legally required to publish exactly these identifying details, in a permanently and directly accessible form. That makes it the most defensible starting point there is for Italian B2B prospecting: no login, no private profile, no paywall.

- **robots.txt is per-domain.** Check the specific domains you care about; legal-notice and footer content exists to be indexed.
- **Personal data.** `legaleRappresentante` and `direttoreResponsabile` are names of natural persons where a site publishes them (1.2 % each). Handling them makes you the data controller under the **GDPR / D.Lgs. 196/2003 as amended**.
- **Outreach compliance is yours.** Italian commercial email is governed by the GDPR and the Codice Privacy; a PEC address in particular is a legal service address and is not an invitation to market to it. You are responsible for lawful basis, opt-out and record-keeping.
- You are responsible for complying with each site's Terms of Service.

# Actor input Schema

## `domains` (type: `array`):

One entry per company. Accepts a bare domain ("bialetti.com"), a homepage URL ("https://www.lavazza.it") or a direct Note legali / Dati societari URL. The Actor reads the site-wide footer first — that is where most Italian companies print their P.IVA — and then merges any Note legali, Dati societari, privacy or contatti page it can reach. One lead per domain. Leave empty to run a 5-domain Italian demo batch.

## `startUrls` (type: `array`):

The same domain list handed over in Apify's Start-URLs shape, so the Console link-list UI, a file upload, Make, Zapier or a Google Sheet can pass it natively. Merged with "Italian company domains".

## `sourceDatasetId` (type: `string`):

Dataset ID of a previous Actor run. Each item's website/domain column is read and enriched — the real agency workflow: run PagineGialle, Google Maps Italy or your own list first, then pipe its output here.

## `domainFieldName` (type: `string`):

Which field of the source dataset (or which CSV column of the list file) holds the domain. Leave empty to auto-detect website / websiteUrl / domain / url / site / homepage. Dotted paths like "company.website" work.

## `domainsFileUrl` (type: `string`):

URL of a public CSV, TXT, JSON or JSONL file holding the domains — for lists too large to paste into the editor. The column is picked with "Domain field / CSV column".

## `skipDomains` (type: `array`):

Domains to skip outright — existing customers, competitors, do-not-contact entries. Matched on the registrable domain, so "www.x.it/page" and "x.it" are the same entry. Applied BEFORE any fetch, so a suppressed domain never costs a request and can never be billed.

## `maxItems` (type: `integer`):

Hard cap on rows DELIVERED AND BILLED this run. Distinct from "Max domains attempted": measured on 740 real Italian domains, roughly 3 of 4 reachable sites yield a billable lead, so 1,000 domains return far fewer than 1,000 billed rows. 0 = unlimited.

## `maxDomains` (type: `integer`):

Cap on domains ATTEMPTED, applied before any request. Use it to sample a large list cheaply. 0 = attempt every domain supplied.

## `maxDiscoveryRequestsPerDomain` (type: `integer`):

How many pages beyond the homepage may be read to enrich a lead. 0 = homepage only (cheapest; measured to still yield the P.IVA on the large majority of Italian sites). Pages the site links are always tried before guessed paths.

## `perDomainTimeoutSecs` (type: `integer`):

Wall-clock deadline for one domain, discovery included. Stops a single slow host stalling the pool.

## `requestConcurrency` (type: `integer`):

How many domains are worked in parallel. Higher is faster; keep it modest to stay polite to small business sites.

## `requestTimeoutSecs` (type: `integer`):

Timeout for a single HTTP request.

## `maxRequestRetries` (type: `integer`):

Retries per request, each on a FRESH proxy exit IP — got-scraping's own retry reuses the flagged IP, which is useless against a soft block.

## `escalateToResidentialOnBlock` (type: `boolean`):

On a 403 / 429 / challenge (never on a dead host or a 404), retry on RESIDENTIAL + country-IT. MEASURED: barilla.com answered 403 on four consecutive fresh datacenter IPs and returned 853,552 usable bytes on residential-IT. Small SME lists rarely trigger it.

## `customUserAgent` (type: `string`):

Override the browser User-Agent sent on every request. Leave empty for the built-in Chrome 124 fingerprint.

## `extraHttpHeaders` (type: `object`):

Additional request headers merged over the defaults — for example a From: header identifying your crawler.

## `proxyConfiguration` (type: `object`):

Proxy settings. Apify DATACENTER is the default and is the measured winner on this corpus; residential is used only as the block escalation above, or when you pin it here. Your own proxy is never overridden.

## `minFieldsRequired` (type: `integer`):

Quality floor AND billing bar: a row must carry at least this many of the 19 value fields (ragione sociale, forma giuridica, P.IVA, codice fiscale, REA number + province, CCIAA, capitale sociale, sede legale, street, CAP, comune, provincia, email, PEC, phone, legale rappresentante, direttore responsabile, registrazione tribunale) before it is delivered and charged. Rows under the bar come back as `partial`, free. 0 = no floor.

## `requireValidPartitaIva` (type: `boolean`):

Deliver (and bill) only companies whose disclosure carries a partita IVA that passes the official DPR 633/72 check digit. A page with no P.IVA, or one whose number fails the checksum, comes back `partial` and is never charged.

## `requireContact` (type: `boolean`):

Deliver (and bill) only companies with an email, a PEC or a phone number.

## `provinciaFilter` (type: `array`):

Two-letter province sigle — MI, RM, TO, NA. A row whose sede legale falls outside them is dropped before delivery and is never charged. Leave empty for the whole country.

## `capFilter` (type: `array`):

Postal-code prefixes, matched from the left: \["20", "00"] keeps Milan and Rome. Combine with the province filter or use it alone.

## `formaGiuridicaFilter` (type: `array`):

Legal forms to keep, punctuation-insensitive: \["S.p.A."] for società per azioni only, \["S.r.l.", "S.r.l.s."] for limited companies, \["Soc. Coop."] for cooperatives.

## `dedupeBy` (type: `string`):

Which key collapses duplicates BEFORE anything is pushed or charged. The default uses the validated P.IVA when the page carries one and the registrable domain otherwise, so two domains owned by the same company collapse to one billed row. "No deduplication" genuinely means none: you get, and pay for, one row per input line.

## `includeMissRows` (type: `boolean`):

Emit a row for every domain that produced no lead — dead host, blocked, publishes no disclosure, under your quality bar, filtered out, duplicate, suppressed or not attempted — with its status and missReason, so you can do coverage accounting. These rows are NEVER charged.

## `includeRawText` (type: `boolean`):

Attach the text of the page the identity fields came from (capped at 40,000 characters) so you can audit a value or re-parse it yourself. Makes the dataset much larger.

## Actor input object example

```json
{
  "domains": [
    "bialetti.com",
    "lavazza.it",
    "granarolo.it",
    "moleskine.com",
    "technogym.com"
  ],
  "maxItems": 1000,
  "maxDomains": 0,
  "maxDiscoveryRequestsPerDomain": 6,
  "perDomainTimeoutSecs": 60,
  "requestConcurrency": 10,
  "requestTimeoutSecs": 30,
  "maxRequestRetries": 2,
  "escalateToResidentialOnBlock": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "minFieldsRequired": 2,
  "requireValidPartitaIva": false,
  "requireContact": false,
  "dedupeBy": "piva-then-domain",
  "includeMissRows": false,
  "includeRawText": false
}
```

# Actor output Schema

## `records` (type: `string`):

The dataset of parsed Italian statutory-disclosure leads — one item per domain.

## `runSummary` (type: `string`):

Coverage accounting for the run: how many domains resolved, how many were blocked, dead, or publish no disclosure, and how many events were charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "bialetti.com",
        "lavazza.it",
        "granarolo.it",
        "moleskine.com",
        "technogym.com"
    ],
    "maxItems": 1000,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/it-note-legali-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "bialetti.com",
        "lavazza.it",
        "granarolo.it",
        "moleskine.com",
        "technogym.com",
    ],
    "maxItems": 1000,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/it-note-legali-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "bialetti.com",
    "lavazza.it",
    "granarolo.it",
    "moleskine.com",
    "technogym.com"
  ],
  "maxItems": 1000,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scrapersdelight/it-note-legali-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/it-note-legali-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/z4NXQoRpgypazS9L4/builds/8E8J9tvKwagmzTy3V/openapi.json
