# Companies House iXBRL Period Extractor (`calm_geometry_p3l/companies-house-ixbrl-period-extractor`) Actor

Extract tagged facts from public Companies House iXBRL filing URLs. Preserve reporting periods, employee comparatives, units and source provenance. Export JSON and CSV. Electronic iXBRL filings only; no PDF OCR or estimated figures.

- **URL**: https://apify.com/calm\_geometry\_p3l/companies-house-ixbrl-period-extractor.md
- **Developed by:** [Kevin Lozada Santos](https://apify.com/calm_geometry_p3l) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$50.00 / 1,000 parsed filings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Companies House iXBRL Period Extractor

Extract numeric facts from public Companies House XHTML filings while preserving their reporting periods, units, dimensions and source document. The employee summary keeps current and comparative periods separate. Missing or conflicting values remain explicit instead of becoming an assumed number.

### How to use

1. Paste **1–20 distinct Companies House XHTML document URLs** into `documentUrls`. Supply specific filing documents with `format=xhtml`; company profile URLs, company numbers and PDFs are not accepted.
2. Keep the default limits or reduce them, then run the Actor.
3. Open the run's **Storage**. Download the **Dataset** as JSON for complete nested evidence, or download **facts.csv** and **periods.csv** from the default key-value store for flat tables. The `OUTPUT` record contains counts and `chargeLimitReached`.

No Companies House API key is required for these public document URLs.

```json
{
  "documentUrls": [
    "https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQ2ODExODg4OGFkaXF6a2N4/document?format=xhtml&download=1"
  ],
  "maxBytesPerDocument": 2097152,
  "requestTimeoutSeconds": 20,
  "selectedMetrics": ["average_employees"]
}
```

### Example and tested scope

The linked FY2024 filing produces these employee rows:

| Period start | Period end | Average employees |
|---|---|---:|
| 2023-10-01 | 2024-09-30 | 6 |
| 2022-10-01 | 2023-09-30 | 9 |

On **2026-09-28**, an Apify platform run tested this [FY2024 filing](https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQ2ODExODg4OGFkaXF6a2N4/document?format=xhtml\&download=1) and the [FY2023 filing](https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQyNTg1MzMyNmFkaXF6a2N4/document?format=xhtml\&download=1) for company 10956764. The run returned **44 numeric facts and 4 source-linked employee-period rows**, matching the source values: 6/9 in FY2024 and 9/18 in FY2023, with no parser errors for those documents. The 2023 period appears in both filings. This two-document test does not establish coverage or accuracy for every filer or taxonomy.

### What you receive

Each processed document produces one dataset item with `rowType`, `status`, stable `sourceUrl`, sanitized `finalUrl`, and:

- **`document`**: exact received-byte SHA-256, byte length, parser version and source provenance. Temporary delivery query strings and fragments are removed.
- **`facts`**: concept and expanded namespace, raw text, normalized `numericValue`, exact decimal string `numericValueExact`, context ID, start/end/instant, dimensions, entity identifier, declared company number, unit, sign, scale and decimals. Invalid or nil facts have null numeric values. Use the exact string when downstream numeric precision matters.
- **`periods`**: average employees per entity and undimensioned duration period, with supporting `factIndexes` and `contextIds`. Missing, invalid or ambiguous results have null `averageEmployees`. Only `AverageNumberEmployeesDuringPeriod` in dated FRC core namespaces with a simple `xbrli:pure` unit is selected. Dimensional facts remain in `facts`.
- **`errors`**: structured diagnostics. A fetch/parse failure has `status: "error"` and empty fact/period collections. A parsed document without usable numeric facts has `status: "no_usable_facts"`. A `parsed` status does not mean every fact is supported; inspect individual statuses and diagnostics.

CSV exports include successfully persisted documents with usable numeric facts. Nested CSV fields are JSON strings and nulls are blank. Duplicate identical employee facts are accepted; competing values or incompatible candidates are marked ambiguous. Filings remain separate: there is no cross-filing deduplication or restatement selection.

### Limits and coverage

| Limit | Value |
|---|---|
| Documents | 1–20, fetched sequentially |
| Bytes per document | 2 MiB default; 5 MiB maximum |
| Total fetch time per document | 20 seconds default; 30 seconds maximum |
| Redirects | At most 3, destination checked |
| XML | UTF-8; depth 128; 100,000 nodes; no DTD/entity declarations |
| Numeric facts | At most 5,000 per document |
| Numeric raw text / normalized fact output | 5 MiB / 8 MiB per document |
| Combined CSV exports | 10 MiB |

Initial URLs must use HTTPS on `find-and-update.company-information.service.gov.uk`, with `/company/{8-character-number}/filing-history/{document-id}/document?format=xhtml` and optional `download=1`. Redirects allow that document path or `https://s3.eu-west-2.amazonaws.com/document-api-images-live.ch.gov.uk/docs/`. Other destinations and compressed responses are rejected. Exceeding a limit fails explicitly; earlier dataset items may remain, but truncated CSV files are not exported.

Supports `ix:nonFraction` in the 2008/2013 namespaces and conservative numeric forms for `numcommadot`, `numdotcomma`, `numspacedot`, `numspacecomma`, `numdash`, `num-dot-decimal`, `num-comma-decimal`, and `zero-dash`. Scale is limited to −12 through 12. Unsupported formats, continued numeric facts and inline fractions are marked invalid.

Company search, filing discovery, bulk ZIPs, PDFs and officer/PSC extraction are outside scope. Turnover is not inferred, and registered addresses are not substituted for trading addresses. This is an extraction tool, not a complete XBRL validator or standardized accounting model.

### Pricing

See **Actor pricing** for the configured price. When pay-per-event pricing is configured, one `document-parsed` event applies to a useful parsed document containing at least one supported numeric fact. Fetch/parse failures and zero-usable-fact documents request no event. There is no per-fact or start event. Reaching the event limit stops further fetching; a document not persisted at the cap is excluded from CSV exports.

# Actor input Schema

## `documentUrls` (type: `array`):

One to 20 explicit HTTPS document URLs on find-and-update.company-information.service.gov.uk with a company filing-history document path and format=xhtml query. Only optional download=1 is accepted. Each URL is checked before requests.

## `maxBytesPerDocument` (type: `integer`):

Maximum downloaded document size in bytes. Default 2 MiB; hard maximum 5 MiB.

## `requestTimeoutSeconds` (type: `integer`):

Download timeout in seconds. Default 20; hard maximum 30.

## `selectedMetrics` (type: `array`):

This version supports only average\_employees. Numeric facts retain their context and reporting period.

## Actor input object example

```json
{
  "documentUrls": [
    "https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQ2ODExODg4OGFkaXF6a2N4/document?format=xhtml&download=1"
  ],
  "maxBytesPerDocument": 2097152,
  "requestTimeoutSeconds": 20,
  "selectedMetrics": [
    "average_employees"
  ]
}
```

# Actor output Schema

## `documents` (type: `string`):

Document records with facts, period metrics, errors and source provenance.

## `factsCsv` (type: `string`):

All extracted numeric facts, with nested evidence encoded as JSON cells. Available after a successful run.

## `periodsCsv` (type: `string`):

Average employee metrics with separate reporting periods and source provenance. Available after a successful run.

## `summary` (type: `string`):

Document, fact, period and error counts plus CSV record keys.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documentUrls": [
        "https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQ2ODExODg4OGFkaXF6a2N4/document?format=xhtml&download=1"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("calm_geometry_p3l/companies-house-ixbrl-period-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "documentUrls": ["https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQ2ODExODg4OGFkaXF6a2N4/document?format=xhtml&download=1"] }

# Run the Actor and wait for it to finish
run = client.actor("calm_geometry_p3l/companies-house-ixbrl-period-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documentUrls": [
    "https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQ2ODExODg4OGFkaXF6a2N4/document?format=xhtml&download=1"
  ]
}' |
apify call calm_geometry_p3l/companies-house-ixbrl-period-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,calm_geometry_p3l/companies-house-ixbrl-period-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FO8BUOJCNwL2pPO6f/builds/vkwQVqhWIgwm94eXb/openapi.json
