# Swedish Company Financials & Annual Reports (årsredovisningar) (`vhsgreed/swedish-company-financials`) Actor

Multi-year financials for any Swedish aktiebolag — nettoomsättning, årets resultat, rörelseresultat, summa tillgångar, eget kapital, medelantal anställda — keyed on organisationsnummer, parsed from the official Bolagsverket annual reports (iXBRL).

- **URL**: https://apify.com/vhsgreed/swedish-company-financials.md
- **Developed by:** [Karl Sundström](https://apify.com/vhsgreed) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.95 / 1,000 company records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Swedish Company Financials & Annual Reports (årsredovisningar)

Multi-year financials for any Swedish **aktiebolag** — nettoomsättning, årets
resultat, rörelseresultat, summa tillgångar, eget kapital, medelantal anställda —
keyed on **organisationsnummer**, parsed from the company's official
**Bolagsverket** annual report filings (iXBRL / Inline XBRL).

Most Swedish company actors stop at the registry key-value layer. This one adds
the financial statement layer — the part buyers actually pay for. The figures are
delivered **parsed and normalised**, not as a document: one flat JSON record per
company, with a multi-year series and derived ratios.

> **No API key needed.** The annual reports are read from Bolagsverket's public
> "Värdefulla datamängder" archive. The optional identity enrichment is the only
> part that uses credentials, and it is off by default.

### What it does

Give it one or more organisation numbers. For each company the actor:

1. locates the company's digitally filed **årsredovisningar** in Bolagsverket's
   public bulk archive (`arsredovisningar/<year>/<week>_<part>.zip`), reading
   **only the ZIP central directory and the members it needs** over HTTP range
   requests — tens of KB per archive instead of the archive's ~93 MB;
2. extracts the filing's iXBRL document (a filing may also contain a separate,
   text-only **revisionsberättelse**; the actor selects the document that
   carries the numbers);
3. parses the financial statement line items against the Swedish **BFN K2/K3**
   taxonomy, handling both taxonomy generations with one mapping;
4. assembles one record per company: latest-year figures, the multi-year series
   the filing itself carries, and named derived ratios.

One dataset item = **one company**. Companies with no filing on record still
produce an honest record with `data_status: "no_filing"` and null figures — a
"no data" answer is a real answer, not a failure — and those records are **not
charged**.

### Input

| Field | Type | Default | Meaning |
|---|---|---|---|
| `organisationNumbers` | string\[] | — | Swedish org numbers, `NNNNNN-NNNN` (hyphen optional). **Required in practice.** |
| `archiveKeys` | string\[] | — | Pin exact weekly archives, e.g. `arsredovisningar/2025/47_4.zip`. Fast, deterministic, few KB. |
| `scanMode` | `all` | `newest` | `all` | How to choose archives when `archiveKeys` is empty. |
| `maxArchivesToScan` | int | 50 | Only for `scanMode: "newest"`. |
| `fiscalYears` | string\[] | — | Only return these fiscal years. |
| `maxYearsPerCompany` | int | 5 | Cap the multi-year series. 0 = no cap. |
| `includeSeries` | bool | true | Emit `financials_by_year` (the multi-year series). |
| `downloadDocuments` | bool | true | Parse the filing. Off = filing metadata only, no figures. |
| `includePersons` | bool | **false** | CEO/board signatories. **Personal data** — see below. |
| `reportDocumentRef` | bool | true | Include archive key / object name / document URL per filing. |
| `enrichWithApi` | bool | false | Fill legal form, registration date, address, SNI via the Bolagsverket API (needs credentials). |
| `concurrency` | int | 6 | Archives read in parallel. |
| `maxRunSeconds` | int | 210 | Soft budget for the archive scan; keeps a run inside a 5-minute slot. |
| `maxDownloadMb` | number | 4096 | Hard cap on bytes downloaded per run. |
| `allowFullArchiveDownload` | bool | true | Allow a full-archive fallback if a server ignores range requests. |

The prefilled input pins **one small archive** and one real company, so a
default run finishes in seconds without credentials or a large download.

#### Finding a company's filing

The public archive has **no company index** — it is organised by filing week, and
the organisation number appears only inside the ZIP member names. So the actor
searches archives newest-first and stops as soon as each company has enough
fiscal years:

- `scanMode: "all"` (default) searches every archive's central directory by range
  request. Complete coverage; measured **~26 s** for a company whose filings sit
  in older archives (it stops early once satisfied).
- `archiveKeys` skips the search entirely — use it when you know the filing
  period, or on the Apify scheduler where predictable run time matters.
- `scanMode: "newest"` trades coverage for the cheapest possible run.

### Output

One record per company. Money is **SEK**; every field copied verbatim from a
filing is suffixed `_reported`, everything we computed is suffixed `_derived`.

```json
{
  "organisationsnummer": "559514-2257",
  "company_name": "…",
  "company_name_source": "filing",
  "legal_form": null,
  "registration_date": null,
  "address": null,
  "sni_codes": null,

  "fiscal_year": "2025",
  "period_start": "2025-01-08",
  "period_end": "2025-08-31",

  "revenue_sek_reported": 0,
  "profit_loss_sek_reported": 198196,
  "operating_profit_sek_reported": -1813,
  "profit_before_tax_sek_reported": 198196,
  "result_after_financial_sek_reported": 198196,
  "total_assets_sek_reported": 1183196,
  "equity_sek_reported": 223196,
  "employees_reported": 0,
  "soliditet_reported": 0.19,
  "kassalikviditet_reported": null,

  "profit_margin_derived": null,
  "operating_margin_derived": null,
  "equity_ratio_derived": 0.1886,
  "revenue_growth_derived": null,
  "equity_growth_derived": null,

  "financials_by_year": [
    {
      "fiscal_year": "2025",
      "period_start": "2025-01-08",
      "period_end": "2025-08-31",
      "revenue_sek_reported": 0,
      "profit_loss_sek_reported": 198196,
      "total_assets_sek_reported": 1183196,
      "equity_sek_reported": 223196,
      "employees_reported": 0
    }
  ],
  "years_available": 1,

  "data_status": "ok",
  "has_filing": true,
  "missing_fields": ["kassalikviditet_reported"],
  "taxonomy_generation": "gen-base/2021-10-31+ar-base",
  "taxonomy_framework": "K2",
  "facts_used": 41,
  "dedupe_conflicts": 2,
  "notes": [],

  "report_archive_key": "arsredovisningar/2025/47_4.zip",
  "report_object_name": "5595142257_2025-08-31.zip",
  "report_document_name": "d9a6fbbc-….xhtml",
  "report_format": "ixbrl-xhtml",
  "report_period_ends": "2025-08-31",
  "report_document_url": "https://vardefulla-datamangder.bolagsverket.se/arsredovisningar-bulkfiler/arsredovisningar/2025/47_4.zip",
  "reports": [ { "period_end": "…", "archive_key": "…", "object_name": "…", "document_names": ["…"] } ],

  "persons": null,
  "source": "bolagsverket",
  "source_url": "https://vardefulla-datamangder.bolagsverket.se/arsredovisningar-bulkfiler/arsredovisningar/2025/47_4.zip",
  "fetched_at": "2026-09-27T19:20:00Z",
  "actor_version": "0.1.0"
}
```

#### Field conventions

- **`_reported`** = copied verbatim from the filing. **`_derived`** = computed by
  the actor, with the formula below. No invented values, ever.
- **`data_status`** is `ok` (figures parsed), `metadata_only`
  (`downloadDocuments: false`), `parse_failed` (filing found, no numbers read), or
  `no_filing` (nothing on record).
- **`employees_reported`** is the filing's *medelantal anställda* (average number
  of employees), not a headcount on a given day.
- **Nulls are honest.** A field the filing does not carry is `null`, and it is
  listed in `missing_fields`. Measured coverage over 1,021 real filings: equity
  99.8 %, total assets 99.5 %, net profit 99.3 %, revenue 94.5 % (holding and
  service companies often report no net revenue), employees 63.9 % (no-staff
  companies omit it).
- **`legal_form` / `registration_date` / `address` / `sni_codes`** are registry
  fields the iXBRL filing does not carry. They are `null` unless
  `enrichWithApi: true` with credentials supplied.

#### Derived ratios (`_derived`)

| Field | Formula | Null when |
|---|---|---|
| `profit_margin_derived` | `profit_loss_sek_reported / revenue_sek_reported` | revenue missing or 0 |
| `operating_margin_derived` | `operating_profit_sek_reported / revenue_sek_reported` | revenue missing or 0 |
| `equity_ratio_derived` | `equity_sek_reported / total_assets_sek_reported` (soliditet) | assets missing or 0 |
| `revenue_growth_derived` | `(rev_t − rev_{t−1}) / |rev_{t−1}|` | no prior year, or prior revenue 0 |
| `equity_growth_derived` | `(eq_t − eq_{t−1}) / |eq_{t−1}|` | no prior year, or prior equity 0 |

Growth compares the two newest fiscal years in `financials_by_year`. Ratios are
rounded to 4 decimals. `soliditet_reported` / `kassalikviditet_reported` are the
values the **filing itself** publishes, so they can be compared against our
derived ratio.

#### Multi-year series

A Swedish annual report contains a **flerårsöversikt** — typically ~5 fiscal
years of net revenue/result and 2–5 years of balance items. The actor emits that
whole series in `financials_by_year` at no extra cost, merged across every filing
it found for the company (newest filing wins on conflicts). One download can
therefore produce several company-years.

#### Correctness notes

- **Duplicate facts are deduped by `(concept, period)`**, keeping the *most
  precise* copy: the flerårsöversikt repeats values rounded to thousands
  (`decimals="-3" scale="3"`) while the statement carries the exact figure
  (`decimals="INF"`). The exact one wins, and the number of disagreements is
  reported as `dedupe_conflicts`.
- **Both taxonomy generations** are handled by one mapping. K2 and K3 use
  identical concept names and differ only in extension namespaces, so matching is
  on the concept's local name; the generation is detected from the namespace URI
  (`se-gen-base/2017-09-30` vs `…/2021-10-31`, with or without the `se-ar-base`
  wrapper) and reported in `taxonomy_generation` / `taxonomy_framework`.
- **No dimensions** are used in K2/K3 filings, so no dimension handling is needed.
- A filing holding **two documents** (årsredovisning + revisionsberättelse) is
  resolved by selecting the document with the numbers and merging numeric facts;
  the text-only auditor copy is skipped.

### Pricing

**Actor start `$0.05` + `$0.0015` per record** (`$1.50 / 1,000` companies).
Records with no filing are free.

This deliberately **undercuts every financials-bearing competitor** found on
Apify on 2026-09-27 while sitting well above cost (the source archive is public
and free):

| Competitor (verified 2026-09-27) | Rate | vs ours |
|---|---|---|
| `regdata/poland-krs-financial-scraper` (closest category analogue) | **$0.06 / statement** + $0.01 start | ~40× higher |
| `scrapeworks/danish-annual-reports-cvr` | from **$5.00 / 1,000** results | 3.3× higher |
| `stealth_mode/proff-bussiness-search-scraper` | **$3.00 / 1,000** | 2× higher |
| `logiover/allabolag-scraper` | **$1.99 / 1,000** | 1.3× higher |
| our own `vhsgreed/allabolag-fresh-details` (registry only, no statements) | $0.20 / 1,000 | 7.5× lower — the financials layer and the series are what justify the step up |

See `PLAN.md` → *Pricing* for the full justification and unit economics.

### Credentials (optional)

Only needed for `enrichWithApi: true`. Runtime environment variables, never
committed:

```
BOLAGSVERKET_CLIENT_ID
BOLAGSVERKET_CLIENT_SECRET
```

OAuth2 **client-credentials**; the actor fetches a bearer token and re-uses it
until expiry. Without them, `enrichWithApi` is skipped and the identity fields
stay null — the financials path does not need any credentials.

### Responsible use

- This actor reads **Bolagsverket's official public archive and API**, not a
  scraped website. Respect their terms and rate limits in your own runs.
- We publish a tool; you are the operator and are responsible for how you run it
  and for complying with the source's terms and with GDPR in your own use.
- We read only the ZIP members needed and cap downloads per run; archive
  requests are made at modest concurrency. Keep it that way if you modify it.
- **Personal data:** the default output contains **no personal data** — only
  company data. `includePersons: true` adds the CEO/board signatories that appear
  in the public filing; that is personal data and you must have a lawful basis
  before enabling it (the actor records its stated basis as GDPR Art. 6(1)(f),
  legitimate interest in already-published public-register data, minimised to
  name + role). Beneficial owners (verkliga huvudmän) are **never** returned.
- Do not republish personal data unlawfully.

### Notes / limitations

- Only **digitally filed** annual reports exist in the corpus. Digital filing was
  voluntary until recently, so coverage for older years and small companies is
  incomplete; it becomes mandatory for fiscal years beginning after
  **2025-12-31** (~587,000 aktiebolag), which steadily closes the gap.
- **No company-name search.** The archive has no name index, so
  `companyNames` is not wired (it is logged and ignored). Supply organisation
  numbers. See `PLAN.md` open question 7.
- The bulk archive is the *only* source: if Bolagsverket is down or changes the
  key layout, the actor logs and returns `no_filing` rather than failing.
- **ESEF/IFRS** reports (listed groups) were not found in the sampled corpus and
  are out of scope for v0.1.0 — the mapping targets BFN K2/K3.
- Raw reports are also free in bulk (`…/arsredovisningar-bulkfiler`: 1,435 ZIPs,
  \~133 GB). The value here is the parsing, multi-year assembly, dedupe,
  cross-generation mapping and query layer — not access to the documents.

### Verification

The iXBRL parser was validated **2,115/2,115 numeric facts** against the
independent `ixbrlparse` library (0 value mismatches) on 14 real filings, and is
re-verified on every release against real filings pulled from the live archive
(`tests/test_smoke_live.py`). Run the offline suite with:

```bash
pip install -r requirements.txt
python -m pytest              # offline, no credentials
RUN_LIVE_SMOKE=1 python -m pytest tests/test_smoke_live.py   # real archive, no credentials
```

# Actor input Schema

## `organisationNumbers` (type: `array`):

One or more Swedish organisation numbers, format NNNNNN-NNNN (the hyphen is optional on input; the actor always normalises it). Each company produces one dataset record.

## `companyNames` (type: `array`):

Optional. Intended for name -> organisation number resolution. NOT IMPLEMENTED in v0.1.0: the public archive carries no name index, so names are logged and ignored. Supply organisationNumbers instead (see PLAN.md open question 7).

## `archiveKeys` (type: `array`):

Advanced. Pin the exact weekly archive objects to search, e.g. arsredovisningar/2025/47\_4.zip. This is the fast, deterministic path: it skips the bucket listing and reads only the archives given (a single archive is typically a few hundred KB to a few MB via range requests). Use it when you know the filing period, or to keep run time and download volume predictable.

## `scanMode` (type: `string`):

How to choose archives when archiveKeys is empty. 'all' searches every archive (complete coverage; archives are read lazily by range request, newest first, stopping as soon as every company has enough fiscal years). 'newest' searches only the newest maxArchivesToScan archives (cheapest, partial coverage).

## `maxArchivesToScan` (type: `integer`):

Only used when scanMode is 'newest': cap how many of the newest archives are searched. Ignored when scanMode is 'all' or archiveKeys is set.

## `fiscalYears` (type: `array`):

Only return reports whose fiscal period ends in these years (e.g. \["2023", "2024"]). Also restricts which archives are listed. Leave empty for all available years.

## `maxYearsPerCompany` (type: `integer`):

Cap the multi-year series per company (newest first). A single annual report already contains several fiscal years, so this is usually satisfied by the first filing found. 0 = no cap.

## `includeSeries` (type: `boolean`):

Emit financials\_by\_year — the multi-year series taken from the filing's own flerarsoversikt (typically ~5 years of revenue/result and 2-5 years of balance items). Off = headline figures for the latest year only.

## `downloadDocuments` (type: `boolean`):

Fetch the iXBRL (Inline XBRL) filing and parse the financial statement line items. When off, the actor only locates the filing and returns its metadata (archive key, object name, period) with no figures — much cheaper, no numbers.

## `includePersons` (type: `boolean`):

OFF by default. When on, returns the CEO and board signatories published in the filing. These are personal data; only enable when you have a lawful basis (see the legal/GDPR notes in README.md). Beneficial owners (verkliga huvudman) are never returned.

## `reportDocumentRef` (type: `boolean`):

Include the archive key, filing object name, document name and download URL for each filing found, so every figure can be traced back to its source document.

## `enrichWithApi` (type: `boolean`):

OFF by default. When on, fills legal\_form, registration\_date, address and sni\_codes from the free Bolagsverket 'Vardefulla datamangder' API. Requires BOLAGSVERKET\_CLIENT\_ID / BOLAGSVERKET\_CLIENT\_SECRET in the environment. Fails soft: without credentials the fields stay null.

## `concurrency` (type: `integer`):

How many archives to read in parallel. Higher is faster but uses more connections.

## `maxRunSeconds` (type: `integer`):

Soft wall-clock budget for the archive scan. When reached, the actor stops scanning and emits what it has found (companies with nothing found become honest no\_filing records). Keeps a run comfortably inside a 5-minute slot.

## `maxDownloadMb` (type: `number`):

Hard cap on bytes downloaded in one run. Reached = the actor stops scanning and reports what it has. Guards against runaway cost/time.

## `allowFullArchiveDownload` (type: `boolean`):

The actor normally reads only the ZIP central directory and the members it needs via HTTP range requests. Set this off to forbid falling back to a full download if a server ignores range requests.

## Actor input object example

```json
{
  "organisationNumbers": [
    "559514-2257"
  ],
  "archiveKeys": [
    "arsredovisningar/2025/47_4.zip"
  ],
  "scanMode": "all",
  "maxArchivesToScan": 50,
  "maxYearsPerCompany": 5,
  "includeSeries": true,
  "downloadDocuments": true,
  "includePersons": false,
  "reportDocumentRef": true,
  "enrichWithApi": false,
  "concurrency": 6,
  "maxRunSeconds": 210,
  "maxDownloadMb": 4096,
  "allowFullArchiveDownload": true
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset of Swedish company financial records, one item per company (with a multi-year series and derived ratios). Keyed on organisationsnummer.

## `resultsJson` (type: `string`):

Full dataset items as raw JSON, including financials\_by\_year and provenance.

## `runView` (type: `string`):

Inspect this run, its logs and storages in Apify Console.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "organisationNumbers": [
        "559514-2257"
    ],
    "archiveKeys": [
        "arsredovisningar/2025/47_4.zip"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("vhsgreed/swedish-company-financials").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "organisationNumbers": ["559514-2257"],
    "archiveKeys": ["arsredovisningar/2025/47_4.zip"],
}

# Run the Actor and wait for it to finish
run = client.actor("vhsgreed/swedish-company-financials").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "organisationNumbers": [
    "559514-2257"
  ],
  "archiveKeys": [
    "arsredovisningar/2025/47_4.zip"
  ]
}' |
apify call vhsgreed/swedish-company-financials --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,vhsgreed/swedish-company-financials"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LQ4MLflhuF9yOJD8X/builds/bD1XcSel9rA8GtUeu/openapi.json
