# Swedish Org.nr Scraper - Organisationsnummer & Contacts (`scrapersdelight/se-orgnr-website-contact-scraper`) Actor

Turn Swedish company domains into registry-grade B2B leads from each site's statutory 28 kap. 5 § disclosure: företagsnamn, Luhn-validated organisationsnummer, säte, momsregistreringsnummer, F-skatt, bankgiro, email and phone. No login, no API key. $6 per 1,000 disclosures.

- **URL**: https://apify.com/scrapersdelight/se-orgnr-website-contact-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$6.00 / 1,000 per swedish company disclosure returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 🇸🇪 Swedish Org.nr Scraper — Organisationsnummer, Säte, Moms & Contacts

**Give it a list of Swedish company domains. Get back one registry-grade B2B lead per company,
read off the disclosure Swedish law already obliges that company to publish on its own website.**

Not a directory scrape. Not a guess. Every identifying field in the output is a statement the
business is legally required to make about itself, in public, on its own site — and the
`organisationsnummer` is checked against its real Luhn check digit before it is delivered.

***

### ⚖️ Why this data exists, and why it is clean

**Aktiebolagslagen (2005:551) 28 kap. 5 §**, quoted verbatim from
[riksdagen.se](https://www.riksdagen.se/sv/dokument-och-lagar/dokument/svensk-forfattningssamling/aktiebolagslag-2005551_sfs-2005-551/):

> *"Ett aktiebolags brev, fakturor, orderblanketter och webbplatser ska ange bolagets företagsnamn,
> den ort där styrelsen har sitt säte samt bolagets organisationsnummer enligt lagen (1974:174) om
> identitetsbeteckning för juridiska personer m.fl. Om bolaget har gått i likvidation, ska också
> detta anges."*

In English: **every Swedish aktiebolag's websites must state its company name, the municipality
where the board has its registered seat, and its organisationsnummer — and must say so if the
company has gone into liquidation.**

Three things follow, and all three shape what this Actor does:

1. **The corpus is the whole Swedish commercial web.** There is no member list to scrape and no
   index of disclosures. Any Swedish company domain is a candidate.
2. **The `säte` is a statutory field in its own right.** It is the *ort* — the municipality — where
   the board sits, and it is frequently a different place from the postal address on the contact
   page. It gets its own column here.
3. **A company in liquidation must say so.** That sentence, published by the company itself, is a
   B2B signal worth a column too — and an input to filter on.

**The duty binds aktiebolag.** An *enskild näringsidkare* (sole trader) is not covered by 28 kap.
5 § and has no organisationsnummer at all — Bolagsverket's own words: *"Enskilda näringsidkare har
sitt personnummer som identitetsbeteckning i stället för ett organisationsnummer."* Feed this Actor
a mixed Swedish list and a large share of it simply is not in scope. Measured on two independent
400-domain samples: **22.9% of live domains on a high-street sample, 33.9% on a traffic-ranked
one** — so the more of your list is `AB`, the closer you sit to the top of that range. Both numbers,
and how they were taken, are below.

***

### 🔍 What it does, per domain

Plain HTTP (`got-scraping` + `cheerio`) over the Apify proxy, with a real browser only where the
measurement said a browser was the *only* route:

1. **Homepage + its footer.** On a Swedish site the 28 kap. 5 § statement sits in the site-wide
   footer far more often than on a dedicated page.
2. **JSON-LD on the same page** — `taxID` / `vatID` / `legalName` inside
   `application/ld+json`. Free, exact, and on the seed probe it was the *only* route to the
   identifier on 2 of 17 reachable sites, both of which render their footer client-side.
3. **Ranked Swedish links.** *Köpvillkor* → *om oss* → *juridisk information* → *allmänna villkor*
   → *integritetspolicy* → *kontakt*. This is the single biggest lever in the build. The English
   ladder used by the UK sibling (*terms and conditions*, *legal notice*, *privacy policy*) scores
   **zero** on every one of those Swedish link texts.
4. **XML sitemap → WordPress REST page index.**
5. Two further rungs that are built, working, and **off by default because they were measured and
   did not pay**: guessed Swedish paths (`/kopvillkor`, `/om-oss`, `/kontakt` …) and a **real
   Chromium**. The numbers for both are below — switch them on deliberately, not hopefully.

Fields are **merged across pages**: a Swedish site routinely carries the organisationsnummer in the
footer, the email on `/kontakt` and the säte on `/kopvillkor`. Turn on **per-field provenance** and
every field names the exact URL it came from.

**A domain that is dead, walled, publishes no disclosure, or turns out to be a sole trader is never
delivered as a lead and never billed.**

***

### 📦 What you get — every field

One flat row per company (or grouped objects — your choice, one switch).

#### Identity

| Field | What it is |
|---|---|
| `registeredName` | The företagsnamn as the disclosure states it |
| `registeredNameSource` | `label` / `suffix` / `prefix` / `is-a-company` — how it was read |
| `tradingName` | A varumärke or trading name, when the site names one separately |
| `companyName` | Best available name: registered name, else JSON-LD `legalName`, else `og:site_name` |
| `legalForm` | Aktiebolag (AB) · Publikt aktiebolag (AB publ) · Handelsbolag (HB) · Kommanditbolag (KB) · Ekonomisk förening · BRF · Ideell förening · Stiftelse · Filial |
| `entityKindIndicated` | What the leading digit group of the organisationsnummer **indicates** |

#### Registry — the statutory core

| Field | What it is |
|---|---|
| `companyNumber` | The organisationsnummer, canonical ten digits |
| `companyNumberFormatted` | The published form, `NNNNNN-NNNN` |
| `companyNumberRaw` | Exactly as printed on the page |
| `companyNumberSource` | `label` / `label-en` / `legal-form` / `label-loose` / `bare` / `json-ld` |
| `companyNumberChecksumValid` | **A real checksum.** The Luhn (mod-10) check digit, not a format guess |
| `companyNumberScheme` | `SE-ORGNR` |
| `sate` | **The statutory säte** — the ort where the board has its seat |
| `jurisdiction` | `Sweden`, or `Sweden (foreign-registered entity)` for a 3-group number |
| `registerUrl` | A free public lookup keyed on the same number |
| `registeredOffice` / `…Postcode` / `…City` | The postal address stated in the disclosure |
| `vatNumber` | The momsregistreringsnummer **as published on the page** |
| `vatNumberValid` | Luhn on its ten-digit core |
| `vatCore` / `vatBranch` / `vatCountry` / `vatChecksumVariant` | The decomposition |
| `derivedVatNumber` | The `SE……01` form **derived** from the organisationsnummer |
| `vatMatchesOrgNr` | Does the published momsnummer decompose to the published organisationsnummer? |
| `inLiquidation` / `liquidationStatement` | The 28 kap. 5 § second-sentence statement, verbatim |
| `fSkattStated` | The site states it is *godkänd för F-skatt* |
| `bankgiro` / `bankgiroValid` · `plusgiro` / `plusgiroValid` | Sweden's business payment identifiers, each with its own mod-10 check |

**About `derivedVatNumber`.** A Swedish momsregistreringsnummer is `SE` + the ten-digit
organisationsnummer + a two-digit sequence. That is verified, not assumed: three real
organisationsnummer read off these companies' own pages were put through the European Commission's
VIES API on 2026-09-20 and all three came back `isValid: true` with the right company name
(`SE556035867201` → *Clas Ohlson Aktiebolag*, `SE556628159701` → *Nordic Nest AB*,
`SE202100548901` → *BOLAGSVERKET*). **But having an organisationsnummer does not make a company
VAT-registered.** `derivedVatNumber` is the *form* the number takes if the company is registered —
it is never presented as proof that it is, which is why it lives in its own column and never in
`vatNumber`.

**About `entityKindIndicated`.** Bolagsverket publishes the mapping (5 = aktiebolag, filial, bank,
försäkringsbolag, europabolag · 9 = handelsbolag/kommanditbolag · 7 or 8 = ekonomisk förening, BRF,
näringsdrivande ideell förening · 2 or 8 = trossamfund · 20 = statlig myndighet · 3 = utländskt
företag) and publishes this caveat with it, which we repeat rather than hide: *"Det går inte att
helt säkert säga vilken företagsform som finns bakom ett organisationsnummer."* The field is named
`…Indicated` for that reason.

#### Contact

| Field | What it is |
|---|---|
| `email` / `emails` | Deduplicated, on-site addresses first; Cloudflare-obfuscated ones decoded |
| `emailConflict` | Set when the best email is on a *different* non-freemail domain |
| `phone` / `phones` | E.164 (`+46…`) |
| `addressLine` / `postcode` / `city` | The trading/visiting address, when it differs from the registered one |
| `socialLinks` | LinkedIn, Facebook, Instagram, X, YouTube, TikTok — profiles only, never share buttons |
| `officerName` / `officerRole` | A VD, styrelseordförande, firmatecknare or ansvarig utgivare named next to a label |
| `termsUrl` / `privacyPolicyUrl` | The site's own policy pages |

#### Provenance & accounting

| Field | What it is |
|---|---|
| `disclosureUrl` | The exact page the disclosure was read from |
| `disclosureSource` | `homepage` · `homepage-json-ld` · `legal-page` · `merged` · `browser-render` |
| `discoveryChannel` | Which channel found it |
| `renderedWithBrowser` | Whether the Chromium rung was used for this domain |
| `pagesParsed` / `pagesChecked` | How much work the row cost |
| `fieldCount` | How many of the 20 value fields are populated |
| `stableId` | The organisationsnummer when valid, else the registrable domain |
| `status` | `ok` · `no_disclosure` · `sole_trader` · `blocked` · `unreachable` · `filtered_out` · `duplicate` · `suppressed` · `skipped_robots` · `error` |
| `missReason` | Why, in a sentence, for anything that is not `ok` |
| `fieldSources` | Per-field provenance (optional) |
| `elapsedMs` / `fetchedAt` | Timing |

***

### 🔐 Sole traders and personnummer — the one thing we deliberately do not hand you

A personnummer and an organisationsnummer are **both ten digits and both carry the same Luhn check
digit**, so a checksum cannot tell them apart. What can: positions 3–4 of a personnummer are a birth
month (01–12), and an organisationsnummer is issued with that pair **≥ 20** precisely so the two
ranges never collide.

When a site publishes a Luhn-valid number whose date pair is below 20, this Actor:

- **does not** treat it as an organisationsnummer,
- **does not** deliver the row as a company lead,
- **does not** charge you for it,
- returns `status: "sole_trader"` with a plain-English `missReason` if you asked for miss rows,
- and **withholds the digits** unless you explicitly switch on *Include a sole trader's
  personnummer*, because a personnummer is personal data and that business is not bound by
  28 kap. 5 § in the first place.

No other Actor on the store draws that line. It is also the honest answer to "why did my Swedish
list convert at 35% and not 80%" — a big slice of any Swedish SME list is enskild firma.

***

### 📊 Measured results — real numbers, not claims

Every number below comes from a run on the Apify platform through the Apify proxy. None of it is
an estimate.

#### The test corpus

Real Swedish businesses from **OpenStreetMap** — `office` + `shop` + `craft` elements carrying a
`website` tag, intersected with the Swedish national area — **4,231 unique hosts**, picked without
any reference to whether they publish a disclosure.

| TLD | hosts | share |
|---|---|---|
| `.se` | 3,426 | **81.0%** |
| `.com` | 562 | **13.3%** |
| `.nu` | 133 | 3.1% |
| `.org` / `.net` / `.eu` / other | 110 | 2.6% |

**19% of a real Swedish business corpus is not `.se`.** That is why no TLD filter is applied by
default.

The corpus was then split by a stable hash of the domain into **3,404 development** hosts and
**827 held-out** hosts. The held-out slice was never looked at while the parser was being written.
The numbers below are from it.

#### What 299 held-out Swedish domains returned

| outcome | domains | note |
|---|---|---|
| **billed leads** | **72** | 27.9% of the domains that were reachable |
| publishes no disclosure | 174 | never billed |
| host did not resolve | 41 | never billed — stale OSM `website` tags |
| blocked / challenged | 9 | never billed |
| sole trader (personnummer) | 2 | never billed, digits withheld |
| duplicate company | 1 | collapsed before billing |

**Delivered rows: 72. `chargedEventCounts`: 72. Exactly equal**, with the counters polled to
stability — no row was billed that was not delivered, and none delivered that was not billed.

The same 299 domains were run three times across three builds and returned **69, 70 and 72** billed
rows. That spread is transport luck — dead and soft-blocked hosts move between runs — and it is why
the yield below is quoted as a range rather than as one run's number.

#### Where the disclosure actually was

| found on | leads |
|---|---|
| merged across two or more pages | 39 |
| a dedicated legal page | 31 |
| the homepage alone | 2 |

| found via | leads |
|---|---|
| the site's own ranked **Swedish** links | 40 |
| homepage footer (free) | 27 |
| homepage JSON-LD | 3 |
| XML sitemap | 2 |
| guessed paths | **0** |

**More than half of all leads needed two or more pages merged.** A Swedish site carries the
organisationsnummer in the footer, the email on `/kontakt` and the address on `/kopvillkor`.

#### Per-field fill, on the billed rows

| field | fill | | field | fill |
|---|---|---|---|---|
| organisationsnummer | **95.8%** | | address line | 86.1% |
| derived momsnummer | **95.8%** | | postcode / city | 86.1% |
| entity kind (indicated) | 95.8% | | email | **91.7%** |
| legal form | 81.9% | | phone | 81.9% |
| företagsnamn | 80.6% | | published momsnummer | 18.1% |
| registered office | 63.9% | | F-skatt stated | 4.2% |
| säte | **4.2%** | | named officer | 4.2% |
| bankgiro | 1.4% | | plusgiro | 0% |

**Every one of the 69 organisationsnummer delivered passed Luhn.** Zero checksum failures.
All 13 published momsnummer passed; 10 of 13 decompose to the same company's organisationsnummer.

**About `säte` at 4.2%.** 28 kap. 5 § requires it, and Swedish sites overwhelmingly do not publish
it. That is a fact about Swedish compliance, not a parser gap — see the next section, which is how
you can tell the difference.

#### How you can tell that is not a parser gap

Two checks, both run against the raw bytes rather than the output:

**1. Label present → value emitted.** Of the 156 captured pages, 52 carry an `org.nr` label in
their bytes. **51 of 52 emitted a value. The 52nd carried a personnummer and was withheld on
purpose.** The corrected rate is **51/51 — 100%.** Where the field is on the page, it comes out.

**2. The negative direction.** Every row with no organisationsnummer was re-swept with an
**independently written** Luhn — different code from the one that ships — looking for a labelled,
organisation-shaped number the parser had missed. **Zero.**

Then the same was done live: 45 domains were drawn at random from the `no_disclosure` bucket and
**re-fetched from scratch**. 42 answered. **Exactly one of the 42 carried an org.nr label at all,
and zero carried an organisationsnummer the Actor had missed.** The bucket is real.

#### We never repair a number to make it check out

A Swedish organisationsnummer is unusually easy to fabricate: a random ten-digit run passes the
Luhn checksum **by luck about one time in ten**, so "it passes the checksum" is not evidence that
it is the number the site published. The rule here is **output ⊆ input** — every digit delivered
appears, as those digits, in the bytes the page served. The suite brute-forces it:

```
4,680 single-digit mutations of the 52 real organisationsnummer in the test corpus
    0 were repaired into a different number
```

Alongside it: a bare Luhn-valid run with no label is refused · a labelled number that fails Luhn is
refused outright rather than nudged · a momsnummer whose core fails Luhn is published **as found**
with `vatNumberValid: false` and never promoted to `companyNumber` · a labelled Luhn-valid
personnummer is refused as a company id · and two numbers in adjacent HTML elements can never be
welded into one, because the separator inside a digit run is horizontal whitespace only.

Normalising `556035-8672` to `5560358672`, dropping the `16` century prefix, or unwrapping
`SE……01` to recover the core is **canonicalisation** — no digit is invented, and the test asserts
every digit is still in the input. Inferring a missing digit is **completion**, and there is no
code path that does it.

#### Printed on the page, or declared to machines?

A page can carry its organisationsnummer inside a `<script>` tag, so the two surfaces are counted
separately rather than lumped:

```
50 printed on the page (a human reading the site can see it)
 2 schema.org JSON-LD (published for machines; reported as companyNumberSource "json-ld")
```

An analytics token that merely matches the pattern is read by neither: the text extractor strips
`<script>` before any label is looked for, and that is asserted, not assumed.

#### Transport

1,937 requests for 299 domains (6.5 per domain). 1,340 OK · 43 hosts dead · 19 blocked ·
80 stub pages · 14 escalations to residential of which 5 recovered. The whole run took **10m 33s**
at concurrency 10 in 1,024 MB, peaking at **590 MB**.

#### Measured on two independent samples, not one

Whether a field is on a page is a question about the parser, and the answer does not move. **How
many Swedish companies publish at all, and how many hide it behind JavaScript, are questions about
the population** — and those answers move a lot depending on where you draw the sample. So the
whole measurement was run twice, on two frames chosen because their biases run in *opposite*
directions:

- **OpenStreetMap business POIs** — over-represents small high-street traders, many of them
  enskild firma who owe no duty at all.
- **Tranco `.se` tail (rank 504k–1M)** — ranks by real resolver traffic, so it over-represents
  larger companies: more aktiebolag, more modern JS front-ends.

400 domains each, dead hosts removed *first* (a dead host is not a JavaScript problem):

| | OSM POIs | Tranco `.se` tail |
|---|---|---|
| live, after dead / blocked / parked removed | 341 of 400 | 357 of 400 |
| **organisationsnummer over plain HTTP** | **78 — 22.9%** | **121 — 33.9%** |
| publishes nothing at all | 25 — 7.3% | 34 — 9.5% |
| sole traders (no orgnr exists) | 3 | 0 |

**The hit rate swings 11 points between the two frames**, and that is the honest answer to "what
will I get?" — **roughly a quarter to a third of live domains**, depending on whether your list
looks more like a high street or more like a company register. A list of `AB` companies sits at
the top of that range; a scraped high street sits at the bottom.

#### Does a real browser help?

Yes, a little, and consistently — and it still ships **off**, because of what it costs to get it.

| | OSM POIs | Tranco `.se` tail |
|---|---|---|
| domains the browser trigger fires on | 238 — 69.8% of live | 202 — 56.6% of live |
| rendered | 229 | 191 |
| **extra leads recovered** | **17 — 5.0% of live** | **13 — 3.6% of live** |
| recovered per domain rendered | 7.1% | 6.4% |

Those are the numbers a standalone harness reached. **Measured inside this Actor** on 240 domains
with the browser switched on, it attempted 118 renders and recovered **3** — 2.5% of what it
attempted, not 7%. The harness renders and merges every candidate page; the Actor stops at the
first rendered page that carries a disclosure marker. **The lower number is the one that applies to
you**, and it is the one quoted here.

Either way the cost is the problem, not the finding: the trigger has to render **well over half of
all live domains**, at **3.9× the compute per domain** and a peak of 1,118 MB against 460 MB. A
measured browser-on run cost **$0.00315 per delivered row** against **$0.00082** for the plain-HTTP
path — roughly **4× the cost per row** for **a few per cent more rows**.

**Turn it on** if your list is large consumer brands, where client-side footers are common and each
lead is worth more — and raise the run memory to 2048 MB, or the Actor will refuse to launch
Chromium rather than risk an out-of-memory failure. **Leave it off** for an ordinary Swedish B2B
list. The numbers above are so you can make that call on evidence rather than on our default.

#### And guessed paths?

Measured over 240 development domains and again over the 299 held-out ones: guessed paths found
the winning page **zero times**, while producing 28% of a run's requests as 404s. They are off by
default and available as an opt-in.

#### A real row, from that run

```json
{
  "domain": "jula.se",
  "disclosureUrl": "https://www.jula.se/",
  "disclosureSource": "merged",
  "discoveryChannel": "homepage",
  "pagesParsed": 3,
  "registeredName": "Jula Sverige AB",
  "registeredNameSource": "suffix",
  "legalForm": "Aktiebolag (AB)",
  "entityKindIndicated": "Aktiebolag, filial, bank, försäkringsbolag eller europabolag",
  "companyNumber": "5569447856",
  "companyNumberFormatted": "556944-7856",
  "companyNumberSource": "label",
  "companyNumberChecksumValid": true,
  "companyNumberScheme": "SE-ORGNR",
  "jurisdiction": "Sweden",
  "registerUrl": "https://www.allabolag.se/5569447856",
  "registeredOffice": "Jula Sverige AB, Box 363, 53224 Skara",
  "registeredOfficePostcode": "532 24",
  "registeredOfficeCity": "Skara",
  "vatNumber": "SE556944785601",
  "vatNumberValid": true,
  "vatChecksumVariant": "luhn-10 on the organisationsnummer core",
  "vatCore": "5569447856",
  "derivedVatNumber": "SE556944785601",
  "vatMatchesOrgNr": true,
  "inLiquidation": false,
  "email": "info@jula.se",
  "phone": "+46423820143",
  "socialLinks": ["https://www.facebook.com/jula", "https://instagram.com/julasverige"],
  "status": "ok"
}
```

#### What to expect from your own list

Measured on two independent samples, the answer is **roughly a quarter to a third of live
domains** — 22.9% on a high-street sample, 33.9% on a traffic-ranked one. 28 kap. 5 § binds
aktiebolag, so the more of your list is `AB`, the closer you sit to the top of that range. Filter
on `AB` in the company name, or run it against a supplier list rather than a scraped high street.
The rest of the list is not lost money: a domain that publishes nothing, is dead, or turns out to
be a sole trader is never delivered and never billed.

***

### 🧰 Every input

**The list** — `domains` (bare domains, homepage URLs or direct `/kopvillkor` URLs, any TLD),
`startUrls` (requestListSources, so Make / Zapier / Clay / Sheets can hand it over natively),
`sourceDatasetId` + `domainFieldName` (enrich a previous Actor's output; Swedish column headings
`webbplats` / `hemsida` / `webbadress` are auto-detected), `domainsFileUrl` (a CSV/TXT/JSON/JSONL
at a URL), `skipDomains`, `previousDatasetId` (a monthly re-run never re-buys a company you
already paid for).

**Limits & cost guards** — `maxItems` (billed rows), `maxDomains` (domains attempted),
`maxDiscoveryRequestsPerDomain`, `maxPagesParsed`, `perDomainTimeoutSecs`, `requestConcurrency`,
`requestTimeoutSecs`, `maxRequestRetries`.

**Discovery** — `discoveryChannels` (homepage+JSON-LD / ranked Swedish links / sitemap / WP REST /
guessed paths), `followWwwAndRootVariants`, `deepJsDiscovery` (browser-free bundle scan),
`renderJsPages` + `browserConcurrency` + `browserWaitMs` (the Chromium rung), `respectRobotsTxt`.

**Network** — `proxyConfiguration`, `proxyCountry`, `escalateToResidentialOnBlock`,
`escalateToUnblockerOnBlock`, `customUserAgent`, `extraHttpHeaders`.

**Row filters — nothing filtered out is ever billed** — `tldFilterMode` + `tldFilter`,
`requireOrgNr`, `requireSate`, `requireFSkatt`, `requireContact`, `excludeLiquidation`,
`entityKindFilter`, `minFieldsRequired`, `emailPolicy`, `dedupeBy`.

**Field extraction** — `validateTaxId`, `includeSoleTraderId`, `extractOfficer`, `extractSocials`,
`extractPolicyUrls`.

**Output shape** — `includeMissRows`, `includeFieldSources`, `flattenOutput`.

***

### 💰 Pricing

**Pay per event, one event: `disclosure-scraped` — $0.006 per Swedish company disclosure
returned, i.e. $6 per 1,000. No start fee. Both Apify auto-events removed.**

That is the same rate as the UK and Irish actors in this family, which do the same job on the same
shape of corpus, and it is set deliberately: no actor on the store reads a **validated** Swedish
organisationsnummer off a company's own website, so the price is anchored to our own family rather
than to a cheaper product that only looks similar.

You are charged **only** for a row you actually receive that carries a registry identifier
(organisationsnummer or momsregistreringsnummer) or företagsnamn + säte. Delivery and billing are
atomic — `Actor.pushData(row, event)` — so a charge cap truncates delivery too and can never leave
you billed for a row you did not get.

**Never charged:** a host that did not resolve · a host that refused us · a domain that publishes
no disclosure · an enskild näringsidkare · a duplicate · a row your own filters removed · any
unbilled coverage row.

Apify platform compute is separate and is billed to your account as usual. The default path is
plain HTTP; the Chromium rung costs materially more compute per domain, which is why it is
narrowly triggered and why its measured share is published above rather than buried.

***

### ❓ FAQ

**Do I need a Bolagsverket account, an API key or a login?**
No. Everything here is read from the company's own public website.

**Why is `registerUrl` not a Bolagsverket link?**
Because Bolagsverket does not publish a stable per-company URL. Its own "Sök företagsinformation"
service at `foretagsinfo.bolagsverket.se/fisok/` answered a **CAPTCHA interstitial** in a real
Chromium session when we tested it on 2026-09-20 — measured, not assumed. So `registerUrl` points
at a free public lookup keyed on the same organisationsnummer, and this README says exactly that
rather than calling it "the register".

**Is `companyNumberChecksumValid` a real checksum?**
Yes — unlike a UK company number, which carries none. It is the Luhn (mod-10) check digit,
verified against two independently published real numbers: Bolagsverket's own **202100-5489** and
Clas Ohlson Aktiebolag's **556035-8672**. The offline validator re-implements Luhn a *different
way* (reverse the string, double the odd indices) from the shipping code (walk forward, double the
even indices) so a bug in one cannot validate itself.

**What is the difference between `vatNumber` and `derivedVatNumber`?**
`vatNumber` is what the site printed. `derivedVatNumber` is what the momsregistreringsnummer would
be, computed from the organisationsnummer. When both are present, `vatMatchesOrgNr` tells you
whether they agree — a mismatch usually means the page belongs to a group company.

**Can I use it on `.com` domains?**
Yes, and you should. Swedish companies trade on `.com` and `.nu` as well as `.se`; the TLD filter
is an opt-in input, never a built-in assumption. The exact split of the test corpus is above.

**What if the site is a single-page app and the footer only renders in a browser?**
Turn on *Render client-side pages in a real browser*. It fires only for a domain where plain HTTP
found nothing **and** the homepage looked like a shell, so it never runs on the large majority of
Swedish sites that answer plain HTTP perfectly well.

**Does it respect robots.txt?**
There is a switch, off by default. A 28 kap. 5 § disclosure is a page the company is legally
obliged to publish and wants indexed. Turn it on if your own compliance policy requires it.

**Will it find a company that does not comply with the law?**
No. If a company breaks 28 kap. 5 § and publishes nothing, it comes back `no_disclosure` and costs
you nothing. That is the honest ceiling on this product and the numbers above are measured with it
included.

***

### ⚠️ Honest limits

- **28 kap. 5 § binds aktiebolag.** Sole traders, and companies that simply do not comply, are not
  in the output. Measured on the held-out slice: **69-72 billed leads from 299 domains** across three
  runs, with ~175 publishing nothing, ~40 hosts dead and 2-3 sole traders. On two independent
  400-domain samples the rate was 22.9% and 33.9% of live domains.
- **The säte is the weakest field in the row — 4.2% fill.** The statute requires it and Swedish
  sites overwhelmingly do not publish it. Do not build a workflow that depends on it; the
  conditional-emit check above is the evidence that this is the market, not the parser.
- **`entityKindIndicated` is an indication, not a verified company form.** Bolagsverket says so
  itself and this README quotes it.
- **`derivedVatNumber` is a format derivation, not proof of VAT registration.** Only 18.1% of rows
  carry a momsnummer the company actually published.
- **`officerName` is sparse — 4.2%.** 28 kap. 5 § does not require a named person, unlike a German
  Impressum.
- **A domain can resolve somewhere you did not expect.** Measured: `interflora.se` redirected a
  run to an individual florist's `*.interflora.se` storefront, which publishes no disclosure of its
  own, while the main site does. `resolvedUrl` on every row tells you where the Actor actually
  landed.
- **Anti-bot.** A minority of Swedish sites sit behind a managed challenge. The Actor escalates
  datacenter → residential automatically and can escalate to Apify Unblocker on request; the
  residual blocked share is reported per run in `RUN_SUMMARY` and is never billed.
- **A proxy IP being blocked is reported, not hidden.** If you have your own proxies, pass them in
  `proxyConfiguration`.

***

### 🧾 Legal & fair use

This Actor reads pages that Swedish law obliges companies to publish, on their own websites, about
themselves. It does not log in, does not bypass authentication, and does not read anything a
visitor could not read.

You are responsible for complying with each site's terms of service, and for handling any personal
data in the output — a named VD, a personal email, or (if you deliberately switch it on) a sole
trader's personnummer — under the GDPR and the Swedish **dataskyddslag (2018:218)**. Personnummer
are withheld by default for exactly that reason.

The `robots.txt` switch is provided and is off by default; turn it on if your own policy requires
it.

# Actor input Schema

## `domains` (type: `array`):

One entry per company. Accepts all three shapes: a bare domain ("nordicnest.se"), a homepage URL ("https://www.byggmax.se") or a direct legal-page URL ("https://x.se/kopvillkor"). ANY TLD is accepted — Swedish companies trade on .com and .nu as well as .se, and filtering to .se would silently drop them. The Actor finds each site's statutory 28 kap. 5 § disclosure and parses it into one lead per domain. Leave empty to run the built-in Swedish demo batch.

## `startUrls` (type: `array`):

The same domain list handed over as URLs, so Make, Zapier, Clay or a Google Sheet can pass it natively (a link to a text/CSV file of URLs also works). Merged with "Swedish domains".

## `sourceDatasetId` (type: `string`):

Dataset ID of a previous Actor run. Each item's domain/website column is read and enriched — the real agency workflow: run a directory, Maps or Bolagsverket scraper first, then pipe its output here.

## `domainFieldName` (type: `string`):

Which field of the source dataset (or which CSV column of the list file) holds the domain. Leave empty to auto-detect domain / website / webbplats / hemsida / url / site. Dotted paths like "company.website" work.

## `domainsFileUrl` (type: `string`):

URL of a CSV, TXT, JSON or JSONL file holding the domains — for lists too big to paste into the editor. The column is picked with "Domain field / CSV column".

## `skipDomains` (type: `array`):

Domains to skip outright — accounts you already own, competitors, do-not-contact entries. Matched on the registrable domain, so "www.x.se/sida" and "x.se" are the same entry.

## `previousDatasetId` (type: `string`):

Dataset ID of an earlier run of THIS Actor. Its stableId / companyNumber / domain values are loaded as a suppression list, so a monthly re-run never re-delivers — and never re-charges you for — a company you already bought.

## `maxItems` (type: `integer`):

Hard cap on rows DELIVERED AND BILLED this run. Distinct from "Max domains": 28 kap. 5 § binds aktiebolag only, not enskild firma, so a mixed Swedish list yields a fraction of its domains as billed rows — the measured hit rate on a held-out slice of real Swedish business domains is published in the README. 0 = unlimited.

## `maxDomains` (type: `integer`):

Cap on domains ATTEMPTED, applied before any request. Use it to sample a big list cheaply. 0 = attempt every domain supplied.

## `maxDiscoveryRequestsPerDomain` (type: `integer`):

Hard cap on discovery + confirmation requests per domain, counted AFTER the homepage. This is the knob that stops one slow host burning a minute of a run.

## `maxPagesParsed` (type: `integer`):

How many fetched pages may be parsed and merged for one domain. The homepage counts as one. Raising it finds more fields on sites that split the disclosure across köpvillkor, kontakt and integritetspolicy pages; lowering it to 1 makes the run homepage-only and very cheap.

## `perDomainTimeoutSecs` (type: `integer`):

Wall-clock deadline for one domain, discovery and any browser render included.

## `requestConcurrency` (type: `integer`):

How many domains are worked in parallel. Higher is faster; keep it modest to stay polite to small business sites. The Actor lowers it automatically, with a warning, when the memory allocated to the run cannot support it — roughly 70 MB per parallel domain.

## `requestTimeoutSecs` (type: `integer`):

Timeout for a single HTTP request.

## `maxRequestRetries` (type: `integer`):

Retries per request, each on a FRESH proxy IP (got-scraping's own retry reuses the flagged IP, which is useless against a soft block).

## `discoveryChannels` (type: `array`):

Which channels may be used to locate the disclosure, tried in this order. MEASURED over 240 real Swedish domains, counting which channel found the winning page: the site's own ranked links 39, the homepage footer 15, the XML sitemap 4, and guessed paths ZERO - while guessed paths produced 498 of that run's 1,791 requests as 404s, so they are OFF by default and are worth switching on only for a list whose footers carry no legal link at all. Ranked links use SWEDISH scoring - kopvillkor, om oss, juridisk information, integritetspolicy, kontakt - which is what makes this work: the English ladder used by the UK sibling scores zero on every one of those.

## `renderJsPages` (type: `boolean`):

Last rung, wired and working but OFF by default because it was MEASURED and did not pay. A real Chromium is launched only for a domain where plain HTTP found no disclosure AND the homepage behaved like a client-rendered shell. On a hand-picked list of large Swedish e-commerce sites it recovered 4 of 139 reachable domains (2.9%); on a random sample of real Swedish businesses it rendered 25 domains over 52 browser page loads and recovered ZERO, at 3.6x the compute per domain. Turn it on for a list of big consumer brands, and raise the run's memory to 2048 MB - below 1536 MB the Actor refuses to launch Chromium rather than risk an out-of-memory failure.

## `browserConcurrency` (type: `integer`):

How many Chromium contexts may render at once. Chromium needs far more memory than a cheerio parse, so this is deliberately much lower than the request concurrency. Raise it only together with the memory allocated to the run.

## `browserWaitMs` (type: `integer`):

How long the browser waits after DOM content load before reading the page, on top of a scroll to the bottom that mounts lazily-rendered footers. Raise it for slow single-page apps.

## `followWwwAndRootVariants` (type: `boolean`):

If the homepage fails, retry the other host form (www.x.se to x.se and back) before declaring the domain unreachable.

## `deepJsDiscovery` (type: `boolean`):

A cheap, browser-free channel for a footer that only exists after client-side render: the Actor reads the page's inline JSON payloads and its external JavaScript bundles, searching them for a legal-page URL. Costs up to 6 extra requests per domain and is never charged separately.

## `respectRobotsTxt` (type: `boolean`):

Fetch and honour each site's robots.txt before requesting anything. Off by default because it costs one extra request per domain and these pages exist to be indexed. Turn it on if your own compliance policy requires it.

## `proxyConfiguration` (type: `object`):

Proxy settings. Apify DATACENTER is the default, with a residential retry only when a host actually challenges us.

## `proxyCountry` (type: `string`):

Pin the exit country for sites that geo-tailor their content. MEASURED across this fleet: Apify datacenter proxies cannot be pinned to a country (they answer HTTP 407), so choosing a country here switches the run to RESIDENTIAL exit nodes in that country. Leave on None for the cheaper, faster datacenter path.

## `escalateToResidentialOnBlock` (type: `boolean`):

On a 403 / challenge (never on a dead host and never on a 404), retry the request once on RESIDENTIAL + country-SE. It only fires on an actual refusal, so it costs nothing on a clean list.

## `escalateToUnblockerOnBlock` (type: `boolean`):

A SECOND escalation, after residential, for the Cloudflare-managed-challenge tail. Off by default because Unblocker requests are billed to your Apify account on top of the row price.

## `customUserAgent` (type: `string`):

Override the browser User-Agent sent on every request. Leave empty for the built-in Chrome 124 fingerprint.

## `extraHttpHeaders` (type: `object`):

Additional request headers, merged over the defaults (e.g. a From: header identifying your crawler).

## `tldFilterMode` (type: `string`):

No TLD filter is applied by default, deliberately: a real Swedish business corpus is only part .se — the rest is .com, .nu, .net and .org — so filtering to .se would discard a large share of it, including many of the larger companies. Use this only when your own list really is single-TLD.

## `tldFilter` (type: `array`):

The TLD list the mode above applies to, without the dot: se, com, nu, net, org.

## `requireOrgNr` (type: `boolean`):

Deliver (and bill) only companies whose disclosure carries a Luhn-valid organisationsnummer. Rows identified only by a momsregistreringsnummer, or only by företagsnamn plus säte, are filtered out here and never billed.

## `requireSate` (type: `boolean`):

Deliver (and bill) only companies that publish "den ort där styrelsen har sitt säte" — the municipality of the registered seat, which 28 kap. 5 § requires and which is often a different place from the postal address.

## `requireFSkatt` (type: `boolean`):

Deliver (and bill) only companies that state they are godkända för F-skatt. Not a 28 kap. 5 § field, but the standard Swedish signal that a supplier invoices with its own tax liability — useful when you are procuring services.

## `requireContact` (type: `boolean`):

Deliver (and bill) only companies with an email or a phone number.

## `excludeLiquidation` (type: `boolean`):

28 kap. 5 § second sentence obliges a company that has gone into likvidation to say so on its website. Turn this on to drop those rows (and konkurs / företagsrekonstruktion statements with them) so they are never delivered or billed.

## `entityKindFilter` (type: `array`):

Keep only companies whose organisationsnummer INDICATES one of these kinds. Bolagsverket's own words: "Det går inte att helt säkert säga vilken företagsform som finns bakom ett organisationsnummer" — so this is a filter on a strong indication, not on a verified fact. Leave empty for all.

## `minFieldsRequired` (type: `integer`):

Quality floor: a row must carry at least this many of the 20 value fields before it is delivered and billed. 0 = no floor.

## `emailPolicy` (type: `string`):

Agencies split hard on whether info@ / kontakt@ counts as a lead. "Role only" keeps just those; "Exclude role" keeps only named mailboxes.

## `dedupeBy` (type: `string`):

Which key collapses duplicates BEFORE anything is pushed or charged. The default uses the organisationsnummer when the page carries a checksum-valid one and the registrable domain otherwise, so two domains owned by the same company collapse to one billed row — and a repeated input line is dropped at queue time, before it costs a request.

## `validateTaxId` (type: `boolean`):

Run the Luhn (mod-10) check digit on every organisationsnummer and on the ten-digit core of every momsregistreringsnummer found, and emit the result. Unlike a UK company number, a Swedish organisationsnummer really does carry a checksum, so a validated number is a fact you can rely on rather than a format guess.

## `includeSoleTraderId` (type: `boolean`):

OFF by default, on purpose. An enskild näringsidkare has no organisationsnummer — its identity beteckning is the owner's PERSONNUMMER, which is personal data, and such a business is not bound by 28 kap. 5 § at all. The Actor detects these, reports them as status "sole\_trader", never charges for them, and withholds the digits unless you switch this on and take responsibility for handling them under the GDPR.

## `extractOfficer` (type: `boolean`):

A VD, styrelseordförande, firmatecknare or ansvarig utgivare named next to a label on the page. Honest expectation: Swedish sites rarely publish one — 28 kap. 5 § does not require it, unlike a German Impressum. The measured fill is in the README.

## `extractSocials` (type: `boolean`):

LinkedIn, Facebook, Instagram, X, YouTube and TikTok links present on the pages read.

## `extractPolicyUrls` (type: `boolean`):

The site's terms and privacy-policy URLs, linked from the pages read.

## `includeMissRows` (type: `boolean`):

Emit a row for every domain that produced no lead — dead host, blocked, publishes no disclosure, a sole trader, filtered out, a duplicate, or suppressed — with its status and missReason, so you can do coverage accounting. These rows are NEVER charged.

## `includeFieldSources` (type: `boolean`):

Add a fieldSources object naming the exact URL each field was read from. Useful when a disclosure is split across a köpvillkor page and a kontakt page and you need to audit which said what.

## `flattenOutput` (type: `boolean`):

On: one flat row, ready for Google Sheets, Clay or a CSV export. Off: fields grouped into company {}, registry {}, contact {}, people {} and policies {} objects.

## Actor input object example

```json
{
  "domains": [
    "clasohlson.com",
    "nordicnest.se",
    "byggmax.se",
    "stadium.se",
    "lyko.com",
    "inet.se",
    "ellos.se",
    "royaldesign.se",
    "komplett.se",
    "teknikmagasinet.se"
  ],
  "maxItems": 1000,
  "maxDomains": 0,
  "maxDiscoveryRequestsPerDomain": 8,
  "maxPagesParsed": 3,
  "perDomainTimeoutSecs": 60,
  "requestConcurrency": 10,
  "requestTimeoutSecs": 25,
  "maxRequestRetries": 2,
  "discoveryChannels": [
    "homepage",
    "anchor",
    "sitemap",
    "wpJson"
  ],
  "renderJsPages": false,
  "browserConcurrency": 2,
  "browserWaitMs": 2500,
  "followWwwAndRootVariants": true,
  "deepJsDiscovery": false,
  "respectRobotsTxt": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "proxyCountry": "none",
  "escalateToResidentialOnBlock": true,
  "escalateToUnblockerOnBlock": false,
  "tldFilterMode": "none",
  "requireOrgNr": false,
  "requireSate": false,
  "requireFSkatt": false,
  "requireContact": false,
  "excludeLiquidation": false,
  "entityKindFilter": [],
  "minFieldsRequired": 0,
  "emailPolicy": "all",
  "dedupeBy": "orgnr-then-domain",
  "validateTaxId": true,
  "includeSoleTraderId": false,
  "extractOfficer": true,
  "extractSocials": true,
  "extractPolicyUrls": true,
  "includeMissRows": false,
  "includeFieldSources": false,
  "flattenOutput": true
}
```

# Actor output Schema

## `records` (type: `string`):

The dataset of Swedish company leads (one item per domain that published a statutory 28 kap. 5 § disclosure).

## `runSummary` (type: `string`):

Coverage accounting for the run: domains attempted, rows billed, and the four kinds of nothing (unreachable host, blocked host, publishes no disclosure, sole trader with no organisationsnummer) kept apart - none of them charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "clasohlson.com",
        "nordicnest.se",
        "byggmax.se",
        "stadium.se",
        "lyko.com",
        "inet.se",
        "ellos.se",
        "royaldesign.se",
        "komplett.se",
        "teknikmagasinet.se"
    ],
    "maxItems": 1000,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/se-orgnr-website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "clasohlson.com",
        "nordicnest.se",
        "byggmax.se",
        "stadium.se",
        "lyko.com",
        "inet.se",
        "ellos.se",
        "royaldesign.se",
        "komplett.se",
        "teknikmagasinet.se",
    ],
    "maxItems": 1000,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/se-orgnr-website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "clasohlson.com",
    "nordicnest.se",
    "byggmax.se",
    "stadium.se",
    "lyko.com",
    "inet.se",
    "ellos.se",
    "royaldesign.se",
    "komplett.se",
    "teknikmagasinet.se"
  ],
  "maxItems": 1000,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scrapersdelight/se-orgnr-website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/se-orgnr-website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0A5KM0frhaYwjSymy/builds/UEq73i3oIvmHaLHhy/openapi.json
