# Portuguese NIPC Scraper - Company Data & Contacts (`scrapersdelight/pt-csc171-website-contact-scraper`) Actor

Turn Portuguese company domains into B2B leads from each site's statutory art. 171 CSC block: firma, legal form, sede, conservatória, mod-11-checked NIPC, capital social, VAT, email and phone. Charged only for rows with a statutory identifier. $1.70 per 1,000 companies.

- **URL**: https://apify.com/scrapersdelight/pt-csc171-website-contact-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.70 / 1,000 per portuguese company disclosure returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 🇵🇹 Portuguese NIPC Scraper — Firma, NIPC, Sede, Capital Social & Contacts

Paste a list of Portuguese company websites and get back **one registry-grade B2B lead per company**,
read off the statutory block every Portuguese company must print on its own website: the **firma**
(registered name), the **legal form**, the **sede** (registered office), the **conservatória**, the
**NIPC checked with the official mod-11 check digit**, the **capital social**, the VAT number, and the
company's **email and phone** — each with the URL it was found on.

- **Contact fields first:** on held-out Portuguese business lists, **89.9–94.8 %** of billed rows carry an
  email and **77.8–88.3 %** a phone number.
- **The identifier is real:** **89.9–93.5 %** of billed rows carry an NIPC, every one mod-11-valid and
  printed on the page. Cross-checked against the European Commission's VIES service, **148 of 156** were
  active VAT registrations, and the registered name VIES returned agreed with the firma we read on
  **111 of 114** comparable rows.
- **You pay only for a statutory row.** $1.70 per 1,000 companies. A dead site, a walled site, a site that
  prints no statutory block, a contact-only page and a sole trader's personal NIF are **never charged**.

***

### 🏛️ Where the data comes from

Portuguese companies are required to print their firma, legal form, sede, conservatória, NIPC and
capital social on their own websites. The Actor finds that block on each site and reads it.

***

### 🔍 What it does, per domain

1. **Fetches the homepage** — directly from Apify's network — and reads its **footer** and its **JSON-LD**
   (`taxID` / `vatID` / `legalName`). Some companies print `SEDE: … | NIF: 500 674 205` on every page.
2. **Ranks the site's own same-site links with Portuguese labels** taken from real Portuguese footers —
   *Ficha Técnica, Informação Legal, Aviso Legal* > *Termos e Condições, Condições Gerais* > *Quem Somos,
   Sobre Nós* > *Política de Privacidade, Proteção de Dados* > *Resolução Alternativa de Litígios* >
   *Contactos* — and merges up to three of those pages field by field. **This is where the block lives:**
   on the held-out lists, **74 of 77** and **94 of 99** billed rows were found on a subpage, not the
   homepage.
3. **XML sitemap and the WordPress page index** — only when the footer offered no candidate link at all.
4. **Parses every field inside a statement window** anchored on a statutory phrase (NIPC, NIF, pessoa
   colectiva, matrícula, conservatória, capital social, com sede na …), never by a page-wide scan — a
   Portuguese NIPC and a Portuguese phone number are **both nine digits**, so a page-wide scan would read
   one as the other.
5. **Validates and refuses**: a number that fails mod 11 is refused; a personal NIF is withheld; a
   foreign register number is refused; nothing is ever repaired, completed or guessed.

A site that blocks a direct request (403 / challenge / connection reset) is retried on **Apify
RESIDENTIAL exit nodes in Portugal**, with a fresh session per attempt. Dead hosts and 404s are never
retried.

***

### 📦 What you get — every field

#### Identity

| Field | What it is |
|---|---|
| `registeredName` | The **firma** as printed — "Sublinhar, Lda.", "Modelo Continente Hipermercados, S.A." — read from the statement that carries the NIPC, preferring the name **nearest** the number. |
| `legalForm` | Sociedade por quotas (Lda.) · Unipessoal Lda. · S.A. · SGPS, S.A. · Cooperativa (CRL) · EIRL · ACE · Sucursal — from the firma's suffix or the statement's own words. |
| `entityKindIndicated` | The entity class the NIPC's **prefix** indicates (5 = pessoa colectiva, 6 = public administration, 9x = condomínio / sociedade irregular / não residente …). An indication from the number's range, not a registry fact. |
| `companyName` | `registeredName`, else the JSON-LD name, else `og:site_name`. |
| `tradingName` | A brand stated as operated by the firma, when it differs from the firma. |

#### Registry — the statutory core

| Field | What it is |
|---|---|
| `companyNumber` | The **NIPC** — nine digits, canonicalised (separators removed), exactly as printed. |
| `companyNumberFormatted` | `508 225 531`. |
| `companyNumberRaw` | The digits as they appeared on the page. |
| `companyNumberChecksumValid` | Always `true` on a delivered NIPC — an invalid one is never emitted. |
| `companyNumberSource` | `label` (NIF / NIPC / pessoa colectiva / contribuinte …), `matricula`, `label-en`, `label-loose` (checksum-gated), `json-ld`, or `bare` (an unlabelled 3-3-3 run in a statement, prefix 5 only). |
| `registeredOffice` · `registeredOfficePostcode` · `registeredOfficeCity` | The **sede**. Postcode `NNNN-NNN` and localidade when printed; a sede printed without a postcode is kept as printed (`registeredOfficeSource: clause-no-postcode`) with the postcode left empty. |
| `conservatoria` · `conservatoriaLocality` | "Conservatória do Registo Comercial de Coimbra" / "C.R.C. de Lisboa", as printed. |
| `matriculaNumber` · `matriculaIsNipc` | The matrícula number. Since 2006 it **is** the NIPC ("número único de matrícula e de pessoa colectiva") — `matriculaIsNipc` says so. |
| `capitalSocial` · `capitalSocialEur` · `capitalFullyPaidStated` | "€66.000,00" as printed, 66000 as a number ("150 mil euros" → 150000), and whether "integralmente realizado" is stated. |
| `inLiquidation` · `liquidationStatement` | art. 171's "em liquidação" (and insolvência / PER) when the site says so. |
| `vatNumber` · `vatNumberValid` · `vatMatchesNipc` | A published VAT number (PT + NIF) and its mod-11 result; whether it decomposes to the same NIPC. |
| `derivedVatNumber` | `PT` + NIPC — the form the company's VAT number takes **if** it is VAT-registered. |
| `registerUrl` | The European Commission **VIES** REST lookup for PT + NIPC (JSON: `isValid`, registered name). There is no free, stable, public per-company page on the Portuguese registry side. |
| `viesStatus` · `viesName` · `viesAddress` | Only with **Cross-check with VIES** on: `VALID` / `INVALID` / `UNAVAILABLE` and the name and address VIES holds. |

#### Contact

| Field | What it is |
|---|---|
| `email` · `emails` | On-site addresses first. Cloudflare-protected and entity-encoded `mailto:` addresses are decoded; "(at)", "\[at]" and "arroba … ponto" forms too. Regulator and ADR-centre mailboxes (CNPD, CNIACC, CICAP, TRIAVE, Livro de Reclamações, gov.pt …) are **dropped**, even on free-mail hosts. |
| `phone` · `phones` | E.164 (`+351…`), fixed, mobile and 707/800/808 numbers. A nine-digit run after a fiscal label is never read as a phone; a regulator's phone printed under its name is dropped. |
| `addressLine` · `postcode` · `city` | A postal address from the pages read (often the shop or trading address, which may differ from the sede). |
| `livroReclamacoesLinked` | Whether the site links the mandatory electronic **Livro de Reclamações** — a "real, trading, consumer-facing Portuguese business" flag. |
| `socialLinks` | Company profiles (LinkedIn, Facebook, Instagram, YouTube, TikTok …); platform help and privacy pages are filtered out. |
| `termsUrl` · `privacyPolicyUrl` | The site's **own** terms and privacy pages. |

#### Provenance & accounting

`domain` · `inputUrl` · `resolvedUrl` · `disclosureUrl` (where the block was found) · `disclosureSource`
(`homepage` / `legal-page` / `merged` / `homepage-json-ld`) · `discoveryChannel` · `pagesParsed` ·
`pagesChecked` · `fieldCount` · `fieldSources` (opt-in: the exact URL of every field) · `stableId`
(the NIPC, else the domain) · `status` · `missReason` · `fetchedAt` · `elapsedMs`.

***

### 👤 Sole traders

A sole trader (*empresário em nome individual*) prints an individual NIF (prefix 1, 2, 3, 45 or the
retired 8 range) rather than an NIPC. Such domains are reported as `status: sole_trader` and **never
billed**; the number itself is returned only when **Include a sole trader's personal NIF** is on.

***

### 📊 Measured results — real numbers, not claims

#### The test corpus, two opposite-biased frames

- **Frame A — OpenStreetMap:** every `office` / `shop` / `craft` element carrying a `website` tag inside
  Portugal's national boundary (mainland, Madeira, Açores): **4,240 hosts** — `.pt` 61.9 %, `.com` 31.8 %
  (78 `.com.pt`), others 6.3 %. Over-represents small high-street traders, many of them sole traders.
- **Frame B — Tranco top-1M, `.pt` tail (rank ≥ 200,000):** traffic-ranked, over-represents larger
  companies with modern sites.
- A stable hash split (`md5(domain) % 5`) held out **300 OSM** and **283 Tranco** domains that were never
  looked at while the parser was written.

#### What the held-out domains returned (build 0.1.6, shipping defaults)

| | A — OSM businesses | B — Tranco `.pt` |
|---|---|---|
| domains | 300 | 283 |
| **billed rows** | **77 — 25.7 %** (32.1 % of reachable) | **99 — 35.0 %** (39.6 % of reachable) |
| contact details only, no statutory block (not billed) | 103 | 58 |
| no statutory block at all (not billed) | 58 | 92 |
| sole trader, personal NIF (not billed) | 2 | 0 |
| host dead (not billed) | 41 | 11 |
| blocked on every rung (not billed) | 19 | 22 |
| duplicate company (not billed) | 0 | 1 |
| **delivered == charged** | 77 == 77 | 99 == 99 |

**Quote coverage as a range: 25.7–35.0 % of a Portuguese domain list, 32.1–39.6 % of the reachable ones**,
depending on how many of the list are companies rather than sole traders.

#### Per-field fill on the billed rows (A / B)

| field | A | B |
|---|---|---|
| **email** | **94.8 %** | **89.9 %** |
| **phone** | **88.3 %** | **77.8 %** |
| **companyNumber (NIPC)** | **93.5 %** | **89.9 %** |
| registeredName (firma) | 74.0 % | 80.8 % |
| legalForm | 75.3 % | 79.8 % |
| registeredOffice (sede) | 75.3 % | 77.8 % |
| registeredOffice postcode | 64.9 % | 65.7 % |
| livroReclamacoesLinked | 62.3 % | 67.7 % |
| addressLine | 51.9 % | 49.5 % |
| conservatória | 26.0 % | 33.3 % |
| capital social | 20.8 % | 27.3 % |
| matrícula | 14.3 % | 13.1 % |
| VAT number printed | 9.1 % | 14.1 % |

Fields move between frames; that is why both are printed. The conservatória and capital social are
statutory, but most Portuguese sites simply do not print them — see below for how we know it is the sites
and not the parser.

#### How you can tell it is not a parser gap

An **independently written** detector (a fiscal label, then a nine-digit run that passes an independently
written mod 11, in the pessoa-colectiva range) was run over every page the Actor read:

- development pages (251 domains, captured on the platform): **88 of 88** domains carrying such a label got
  their NIPC emitted.
- held-out pages (482 domains, captured on the platform): **150 of 150** — after one fix: the first pass
  was 149 of 150, because "registada com o número fiscal 505825686" opened no statement window. That anchor
  was added, the miss re-tested, and the number is quoted as found.

And in the other direction, 50 random **unbilled** held-out domains were re-crawled far deeper (every
legal-, contact- or about-looking same-site link plus ten guessed paths — 723 pages) with the independent
detector: **46 of 49 reachable really publish no NIPC anywhere**. Of the other three, one listed its
*member* companies' NIPCs (not its own — correctly not attributed), one blocked our subpage requests, and
one printed it on a Termos page its homepage never links to.

#### We never repair a number

Output is a subset of input: every emitted NIPC, VAT number, phone, postcode, conservatória, capital and
firma is asserted present in the bytes it came from. Every single-digit mutation of every real NIPC in the
test pages (**12,555** mutants) was fed back in: a mutant is emitted only when it is itself a valid number
(then it *is* the input); **0 were "repaired"** into a different number. A newline between digit groups
never welds them into one number; a labelled number that fails mod 11 is refused; placeholders
(`500000000`, `123456789`) are refused; a French SIREN and an Australian ACN printed on a Portuguese page
are refused.

#### Cross-checked against the EU's own register of VAT numbers

All 156 NIPCs read on the held-out pages were checked live against the European Commission's **VIES**:
**148 VALID**; **8 INVALID** — each one mod-11-valid and printed on the site, belonging to entities that are
not (or no longer) VAT-active: a political party, dissolved companies. Where both names were present, the
VIES name matched the firma we read on **111 of 114**; two of the three "mismatches" are abbreviations of
the same company (C\&P Lda. = Catarina & Pedro Lda.), and one is a real pairing error on a page naming two
insurers. Turn on **Cross-check with VIES** to get this answer on every row (measured on the platform:
23 of 23 checked VALID, 0 unavailable).

#### Transport

Direct from Apify's network, with a residential retry only on a refusal. Held-out A: **1,115 direct
requests, 66 residential** (20 recovered); B: **1,218 direct, 154 residential** (74 recovered).

#### A real row (build 0.1.6, held-out frame B)

```json
{
  "domain": "bybebe.pt",
  "resolvedUrl": "https://www.bybebe.pt/",
  "disclosureUrl": "https://www.bybebe.pt/pt/termos-e-condicoes",
  "disclosureSource": "legal-page",
  "registeredName": "Sublinhar, Lda.",
  "legalForm": "Sociedade por quotas (Lda.)",
  "entityKindIndicated": "Pessoa colectiva (sociedade, associação ou outra entidade registada no RNPC)",
  "companyNumber": "508225531",
  "companyNumberFormatted": "508 225 531",
  "companyNumberSource": "label",
  "companyNumberChecksumValid": true,
  "registeredOffice": "bybebé, Rua Brotero, 28 – 3030-317 Coimbra",
  "registeredOfficePostcode": "3030-317",
  "registeredOfficeCity": "Coimbra",
  "conservatoria": "Conservatória do Registo Comercial de Coimbra",
  "matriculaNumber": "508225531",
  "matriculaIsNipc": true,
  "capitalSocial": "€66.000,00",
  "capitalSocialEur": 66000,
  "inLiquidation": false,
  "derivedVatNumber": "PT508225531",
  "email": "info@bybebe.com",
  "phone": "+351239724592",
  "livroReclamacoesLinked": true,
  "socialLinks": ["https://www.facebook.com/bybebePT", "https://www.instagram.com/bybebe.pt/", "https://www.tiktok.com/@bybebe.pt"],
  "termsUrl": "https://www.bybebe.pt/pt/termos-e-condicoes",
  "privacyPolicyUrl": "https://www.bybebe.pt/pt/politica-de-privacidade",
  "fieldCount": 17,
  "status": "ok"
}
```

(Abridged for width — the real row also carries the provenance fields and `registerUrl`; `vies*` appear only with the VIES option on. VIES, asked the same
day: `508225531` → VALID, "SUBLINHAR LDA".)

#### What to expect from your own list

- A list of **companies** (Lda., S.A.) converts best; a list of high-street shops converts at the low end,
  because many are sole traders with no art. 171 duty.
- Contact-only sites are common in Portugal (103 of 300 in the OSM frame): the Actor tells you they exist
  (`includeMissRows`) but never bills them.

***

### 🧰 Every input

| Input | Default | What it does |
|---|---|---|
| `domains` | demo batch | Domains, homepage URLs or direct legal-page URLs. Any TLD. |
| `startUrls` | — | The same list as URLs (Make / Zapier / Clay / Sheets). |
| `sourceDatasetId` + `domainFieldName` | — | Enrich a previous Actor's dataset (e.g. a Google Maps or Páginas Amarelas run). |
| `domainsFileUrl` | — | A CSV / TXT / JSON / JSONL list at a URL. |
| `skipDomains` · `previousDatasetId` | — | Suppression: never re-deliver or re-charge a company you already have. |
| `maxItems` | 1000 | Hard cap on billed rows. |
| `maxDomains` | 0 (all) | Cap on domains attempted. |
| `maxDiscoveryRequestsPerDomain` | 8 | Requests per domain after the homepage. |
| `maxPagesParsed` | 3 | Subpages parsed and merged per domain. |
| `perDomainTimeoutSecs` · `requestTimeoutSecs` · `maxRequestRetries` | 60 · 25 · 2 | Budgets; every request and every domain also has a hard deadline. |
| `requestConcurrency` | 10 | Parallel domains (lowered automatically to what the memory can hold). |
| `discoveryChannels` | homepage, anchor, sitemap, wpJson | `pathGuess` is opt-in (measured: it found 1 extra of 49 unbilled domains). |
| `followWwwAndRootVariants` · `deepJsDiscovery` · `respectRobotsTxt` | on · off · off | Host variants; JS-bundle link search; robots.txt. |
| `proxyConfiguration` | direct | Your own proxy for the FIRST attempt, if you want one. |
| `proxyCountry` · `escalateToResidentialOnBlock` · `escalateToUnblockerOnBlock` | PT · on · off | The retry ladder on a refusal. |
| `customUserAgent` · `extraHttpHeaders` | — | Request identity. |
| `tldFilterMode` · `tldFilter` | none | Include / exclude TLDs (no filter by default — 38 % of the OSM corpus is not `.pt`). |
| `requireNipc` · `requireCapitalSocial` · `requireContact` · `excludeLiquidation` · `legalFormFilter` · `minFieldsRequired` | off | Quality filters — a filtered row is never billed. |
| `emailPolicy` | all | all · role-only (info@, geral@) · exclude-role. |
| `dedupeBy` | NIPC, else domain | Collapse duplicates before anything is pushed or charged. |
| `validateTaxId` | on | Run mod 11 on every NIPC / VAT number. |
| `includeSoleTraderNif` · `extractOfficer` | off · off | Return a sole trader's NIF; return a named gerente / administrador. |
| `verifyWithVies` | off | Live VIES status + registered name per NIPC. |
| `extractSocials` · `extractPolicyUrls` · `includeMissRows` · `includeFieldSources` · `flattenOutput` · `saveRawPages` | on · on · off · off · on · off | Output shape and audit trail (`saveRawPages` stores every page read, gzipped, in the run's key-value store). |

***

### 💰 Pricing

**$1.70 per 1,000 companies** — one `disclosure-scraped` event ($0.0017) per billed row, no start fee.
Delivery and billing are atomic, so a spending cap can never leave you with an unpaid row or charge you
for one you did not get.

**Never charged:** a dead host · a host that refused every rung · a site with no statutory block · a
contact-only site · a sole trader's personal NIF · a duplicate · a row your own filters removed · any
coverage row.

Compared with the generic EU legal-notice extractors on the store: they list Portugal as a supported
format, but their fields are an untyped `vatId` / `registrationNumber` modelled on the German Impressum
("e.g. HRB 133604"), their own READMEs say register fields are "richest in German-speaking markets", and
they bill contact-page fallback rows. None reads the NIPC with its check digit, the conservatória or the
capital social.

***

### ❓ FAQ

**Does it find Portuguese companies for me?** No — it turns a list of domains into company records.
Feed it a Google Maps, Páginas Amarelas or directory export via `sourceDatasetId`.

**Is the NIPC the same as the VAT number?** The VAT number is `PT` + the NIF. `derivedVatNumber` gives
that form; only VIES (`verifyWithVies`) can tell you the company is actually VAT-active.

**Why is the conservatória empty on most rows?** Because most sites do not print it. The label-present → value-emitted check above (150 of 150) is how we know.

**Big retailers?** Measured on the platform on 2026-09-23: worten.pt, fnac.pt and decathlon.pt answer 403
on every rung, including residential Portugal; they are reported as `blocked` and never charged (try your
own proxy or the Unblocker toggle). bertrand.pt, which refused our recon from a home IP, delivered a row.

**Why no browser?** The Swedish sibling measured a real Chromium recovering 3.6–5.0 % more domains at ~4×
the cost per domain; this Actor ships plain HTTP only. JavaScript-only shops (e.g. an SPA with no server
HTML) are reported as `no_disclosure`.

***

### ⚠️ Data limits

- **Coverage is 25.7–35.0 % of a domain list**; most Portuguese sites print either nothing statutory or only
  contact details.
- **A mod-11-valid NIPC is well formed, not necessarily live:** 8 of 156 were not VAT-active in VIES.
- **Name ↔ NIPC pairing** is read from the same statement, nearest first; on a page naming two companies it
  can still pair the wrong ones — measured 1 in 114 against VIES.
- **Capital social 20.8–27.3 %, conservatória 26.0–33.3 %, VAT printed 9.1–14.1 %** — what sites print.
- **The sede may be printed without a postcode**; it is then kept as printed with no postcode/city.
- **7.3–8.1 % of live sites refuse even the residential retry** (19 of 259 and 22 of 272 on the held-out frames).
- **Guessed paths are off** — an unlinked Termos page is missed (1 of 49 unbilled domains in the re-probe).
- **Foreign companies** operating in Portugal are delivered with the NIPC of their Portuguese branch when
  printed; a foreign register number is never emitted as an NIPC.

# Actor input Schema

## `domains` (type: `array`):

One entry per company. Accepts a bare domain ("pingodoce.pt"), a homepage URL ("https://www.radiopopular.pt") or a direct legal-page URL ("https://x.pt/termos-e-condicoes"). ANY TLD is accepted — Portuguese companies trade on .com and .com.pt as well as .pt. The Actor finds each site's statutory art. 171 CSC block (firma, tipo, sede, conservatória, NIPC, capital social) and parses it into one lead per domain. Leave empty to run the built-in Portuguese demo batch.

## `startUrls` (type: `array`):

The same domain list handed over as URLs, so Make, Zapier, Clay or a Google Sheet can pass it natively (a link to a text/CSV file of URLs also works). Merged with "Portuguese company domains".

## `sourceDatasetId` (type: `string`):

Dataset ID of a previous Actor run. Each item's domain / website column is read and enriched — the real agency workflow: run a Google Maps, Páginas Amarelas or directory scraper first, then pipe its output here.

## `domainFieldName` (type: `string`):

Which field of the source dataset (or which CSV column of the list file) holds the domain. Leave empty to auto-detect domain / website / site / sitio / url. Dotted paths like "company.website" work.

## `domainsFileUrl` (type: `string`):

URL of a CSV, TXT, JSON or JSONL file holding the domains — for lists too big to paste into the editor. The column is picked with "Domain field / CSV column".

## `skipDomains` (type: `array`):

Domains to skip outright — accounts you already own, competitors, do-not-contact entries. Matched on the registrable domain, so "www.x.pt/pagina" and "x.pt" are the same entry.

## `previousDatasetId` (type: `string`):

Dataset ID of an earlier run of THIS Actor. Its stableId / NIPC / domain values are loaded as a suppression list, so a monthly re-run never re-delivers — and never re-charges you for — a company you already have.

## `maxItems` (type: `integer`):

Hard cap on rows DELIVERED AND BILLED this run. Distinct from "Max domains": art. 171 CSC binds sociedades, not sole traders, and many sites never publish the block, so a mixed Portuguese list yields a fraction of its domains as billed rows. 0 = no cap.

## `maxDomains` (type: `integer`):

Cap on domains ATTEMPTED, applied before any request. Use it to sample a big list cheaply. 0 = attempt every domain supplied.

## `maxDiscoveryRequestsPerDomain` (type: `integer`):

Hard cap on discovery + confirmation requests per domain, counted AFTER the homepage. This is the knob that stops one slow host burning a minute of a run.

## `maxPagesParsed` (type: `integer`):

How many SUBPAGES may be parsed and merged for one domain, on top of the homepage. Raising it finds more fields on sites that split the block across Termos e Condições, Política de Privacidade and Contactos; lowering it makes each domain cheaper.

## `perDomainTimeoutSecs` (type: `integer`):

Wall-clock deadline for one domain, all discovery included.

## `requestConcurrency` (type: `integer`):

How many domains are worked in parallel. Higher is faster; keep it modest to stay polite to small business sites. The Actor lowers it automatically, with a warning, when the memory allocated to the run cannot support it — roughly 70 MB per parallel domain.

## `requestTimeoutSecs` (type: `integer`):

Timeout for a single HTTP request.

## `maxRequestRetries` (type: `integer`):

Retries after the first attempt. The first attempt is DIRECT; a transient error gets one more direct try, and a block (403 / challenge) is retried on RESIDENTIAL exit nodes with a FRESH session each time.

## `discoveryChannels` (type: `array`):

Which channels may be used to locate the statutory block. homepage = the homepage itself (footer + JSON-LD); anchor = the site's own same-site links ranked by Portuguese labels (Ficha Técnica, Informação Legal, Aviso Legal, Termos e Condições, Quem Somos, Política de Privacidade, Contactos); sitemap and wpJson run ONLY when the footer offered no candidate link; pathGuess (off by default) tries ten conventional Portuguese paths.

## `followWwwAndRootVariants` (type: `boolean`):

If the homepage fails, retry the other host form (www.x.pt to x.pt and back) before declaring the domain unreachable.

## `deepJsDiscovery` (type: `boolean`):

A cheap, browser-free channel for a footer that only exists after client-side render: the Actor reads the page's inline JSON payloads and its external JavaScript bundles, searching them for a legal-page URL. Costs up to 6 extra requests per domain and is never charged separately.

## `respectRobotsTxt` (type: `boolean`):

Skip pages the site's robots.txt disallows. Costs one extra request per domain. Off by default.

## `proxyConfiguration` (type: `object`):

Leave as is for the measured default: DIRECT requests first, then Apify RESIDENTIAL (Portugal) only for a host that actually refuses us. Supply your own proxy URLs or an Apify proxy group here to route the FIRST attempt through them instead.

## `proxyCountry` (type: `string`):

ISO country code for the RESIDENTIAL retry. PT by default, because some Portuguese retailers geo-tailor or geo-block. Only used when a host refuses the direct request.

## `escalateToResidentialOnBlock` (type: `boolean`):

On a 403 / challenge / connection reset (never on a dead host and never on a 404), retry on RESIDENTIAL exit nodes pinned to the country above, a fresh session per attempt. It only fires on an actual refusal, so it costs nothing on a clean list.

## `escalateToUnblockerOnBlock` (type: `boolean`):

A SECOND escalation, after residential, for the Cloudflare-managed-challenge tail. Off by default because Unblocker requests are billed to your Apify account on top of the row price.

## `customUserAgent` (type: `string`):

Override the browser User-Agent sent on every request. Leave empty for the built-in Chrome 124 fingerprint.

## `extraHttpHeaders` (type: `object`):

Additional request headers, merged over the defaults (e.g. a From: header identifying your crawler).

## `tldFilterMode` (type: `string`):

No TLD filter is applied by default, deliberately: a real Portuguese business corpus is only part .pt — the rest is .com, .com.pt, .eu and others — so filtering to .pt would discard real companies. Use "include" or "exclude" with the list below.

## `tldFilter` (type: `array`):

The TLD list the mode above applies to, without the dot: pt, com, com.pt, eu.

## `requireNipc` (type: `boolean`):

Deliver (and bill) only companies whose block carries a mod-11-valid collective NIPC. Rows identified only by a PT VAT number, or by firma plus conservatória / capital social, are filtered out here and never billed.

## `requireCapitalSocial` (type: `boolean`):

Deliver (and bill) only companies that print their capital social, which art. 171 n.º 2 requires of every Lda., S.A. and sociedade em comandita por acções.

## `requireContact` (type: `boolean`):

Deliver (and bill) only companies with an email or a phone number.

## `excludeLiquidation` (type: `boolean`):

Art. 171 n.º 1 obliges a company in liquidation to say so ("sendo caso disso, a menção de que a sociedade se encontra em liquidação"). Turn this on to drop those rows (and insolvência / PER statements with them) so they are never billed.

## `legalFormFilter` (type: `array`):

Keep only rows whose legal form contains one of these words, e.g. "Lda." , "Unipessoal", "S.A.", "SGPS", "Cooperativa". Empty = every form.

## `minFieldsRequired` (type: `integer`):

Quality floor: a row must carry at least this many of the 18 value fields before it is delivered and billed. 0 = no floor.

## `emailPolicy` (type: `string`):

Agencies split hard on whether info@ / kontakt@ counts as a lead. "Role only" keeps just those; "Exclude role" keeps only named mailboxes.

## `dedupeBy` (type: `string`):

Which key collapses duplicates BEFORE anything is pushed or charged. The default uses the NIPC when the block carries a mod-11-valid one and the registrable domain otherwise, so two domains owned by the same company collapse into one billed row.

## `validateTaxId` (type: `boolean`):

Run the Portuguese mod-11 check digit on every NIPC and PT VAT number found and emit the result. A number that fails is never promoted to the NIPC column, and no number is ever repaired or completed.

## `includeSoleTraderNif` (type: `boolean`):

Off by default. A sole trader (empresário em nome individual) prints an individual NIF (prefix 1, 2, 3, 45 or the retired 8 range) instead of an NIPC. These domains are reported as status sole\_trader and never billed; turn this on to also return the number in soleTraderNif.

## `extractOfficer` (type: `boolean`):

Return a gerente, administrador or encarregado de proteção de dados named next to a label on the pages read (officerName, officerRole). Off by default.

## `extractSocials` (type: `boolean`):

LinkedIn, Facebook, Instagram, X, YouTube and TikTok links present on the pages read.

## `extractPolicyUrls` (type: `boolean`):

The site's terms and privacy-policy URLs, linked from the pages read.

## `includeMissRows` (type: `boolean`):

Emit a row for every domain that produced no lead — dead host, blocked, publishes no statutory block, contact details only, a sole trader, filtered out, a duplicate, or suppressed — with its status and missReason, so you can do coverage accounting. These rows are never billed.

## `includeFieldSources` (type: `boolean`):

Add a fieldSources object naming the exact URL each field was read from. Useful when the block is split across a Termos e Condições page and a Contactos page and you need to audit which said what.

## `flattenOutput` (type: `boolean`):

On: one flat row, ready for Google Sheets, Clay or a CSV export. Off: fields grouped into company {}, registry {}, contact {}, people {} and policies {} objects.

## `verifyWithVies` (type: `boolean`):

For every billed row with an NIPC, ask the European Commission VIES service whether PT+NIPC is an ACTIVE VAT registration today and under what registered name, and add viesStatus / viesName / viesAddress. A checksum proves a number is well formed; VIES says it is live. Measured on 156 held-out NIPCs: 148 VALID, 8 INVALID, and the VIES name agreed with the firma read off the site on 111 of 114 comparable rows. Adds roughly 0.5-1 s per row (the EU endpoint is rate-limited, so calls are serialised). Never changes what is billed.

## `saveRawPages` (type: `boolean`):

Store every page the parser actually read, gzipped, in this run's key-value store (key <domain>\_\_<n>), so any row can be re-checked against the exact bytes it came from. Adds storage cost; off by default.

## Actor input object example

```json
{
  "domains": [
    "pingodoce.pt",
    "leroymerlin.pt",
    "kuantokusta.pt",
    "cgd.pt",
    "darty.pt",
    "zippy.pt",
    "lactogal.pt",
    "topatlantico.pt",
    "sanitop.pt",
    "b-online.pt"
  ],
  "maxItems": 1000,
  "maxDomains": 0,
  "maxDiscoveryRequestsPerDomain": 8,
  "maxPagesParsed": 3,
  "perDomainTimeoutSecs": 60,
  "requestConcurrency": 10,
  "requestTimeoutSecs": 25,
  "maxRequestRetries": 2,
  "discoveryChannels": [
    "homepage",
    "anchor",
    "sitemap",
    "wpJson"
  ],
  "followWwwAndRootVariants": true,
  "deepJsDiscovery": false,
  "respectRobotsTxt": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "proxyCountry": "PT",
  "escalateToResidentialOnBlock": true,
  "escalateToUnblockerOnBlock": false,
  "tldFilterMode": "none",
  "requireNipc": false,
  "requireCapitalSocial": false,
  "requireContact": false,
  "excludeLiquidation": false,
  "minFieldsRequired": 0,
  "emailPolicy": "all",
  "dedupeBy": "nipc-then-domain",
  "validateTaxId": true,
  "includeSoleTraderNif": false,
  "extractOfficer": false,
  "extractSocials": true,
  "extractPolicyUrls": true,
  "includeMissRows": false,
  "includeFieldSources": false,
  "flattenOutput": true,
  "verifyWithVies": false,
  "saveRawPages": false
}
```

# Actor output Schema

## `records` (type: `string`):

One item per domain that published a statutory art. 171 CSC block (firma, tipo, sede, conservatória, NIPC, capital social) plus its contacts.

## `runSummary` (type: `string`):

Coverage accounting: domains attempted, rows billed, and the kinds of nothing (unreachable host, blocked host, no statutory block, contact details only, sole trader with a personal NIF) kept apart - none of them charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "pingodoce.pt",
        "leroymerlin.pt",
        "kuantokusta.pt",
        "cgd.pt",
        "darty.pt",
        "zippy.pt",
        "lactogal.pt",
        "topatlantico.pt",
        "sanitop.pt",
        "b-online.pt"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/pt-csc171-website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "pingodoce.pt",
        "leroymerlin.pt",
        "kuantokusta.pt",
        "cgd.pt",
        "darty.pt",
        "zippy.pt",
        "lactogal.pt",
        "topatlantico.pt",
        "sanitop.pt",
        "b-online.pt",
    ],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/pt-csc171-website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "pingodoce.pt",
    "leroymerlin.pt",
    "kuantokusta.pt",
    "cgd.pt",
    "darty.pt",
    "zippy.pt",
    "lactogal.pt",
    "topatlantico.pt",
    "sanitop.pt",
    "b-online.pt"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call scrapersdelight/pt-csc171-website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/pt-csc171-website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kxYrE78XZ6ERtwRH3/builds/ZORsmxAoWkbmDDqt6/openapi.json
