Companies House iXBRL Period Extractor avatar

Companies House iXBRL Period Extractor

Pricing

$50.00 / 1,000 parsed filings

Go to Apify Store
Companies House iXBRL Period Extractor

Companies House iXBRL Period Extractor

Extract tagged facts from public Companies House iXBRL filing URLs. Preserve reporting periods, employee comparatives, units and source provenance. Export JSON and CSV. Electronic iXBRL filings only; no PDF OCR or estimated figures.

Pricing

$50.00 / 1,000 parsed filings

Rating

0.0

(0)

Developer

Kevin Lozada Santos

Kevin Lozada Santos

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Extract numeric facts from public Companies House XHTML filings while preserving their reporting periods, units, dimensions and source document. The employee summary keeps current and comparative periods separate. Missing or conflicting values remain explicit instead of becoming an assumed number.

How to use

  1. Paste 1–20 distinct Companies House XHTML document URLs into documentUrls. Supply specific filing documents with format=xhtml; company profile URLs, company numbers and PDFs are not accepted.
  2. Keep the default limits or reduce them, then run the Actor.
  3. Open the run's Storage. Download the Dataset as JSON for complete nested evidence, or download facts.csv and periods.csv from the default key-value store for flat tables. The OUTPUT record contains counts and chargeLimitReached.

No Companies House API key is required for these public document URLs.

{
"documentUrls": [
"https://find-and-update.company-information.service.gov.uk/company/10956764/filing-history/MzQ2ODExODg4OGFkaXF6a2N4/document?format=xhtml&download=1"
],
"maxBytesPerDocument": 2097152,
"requestTimeoutSeconds": 20,
"selectedMetrics": ["average_employees"]
}

Example and tested scope

The linked FY2024 filing produces these employee rows:

Period startPeriod endAverage employees
2023-10-012024-09-306
2022-10-012023-09-309

On 2026-09-28, an Apify platform run tested this FY2024 filing and the FY2023 filing for company 10956764. The run returned 44 numeric facts and 4 source-linked employee-period rows, matching the source values: 6/9 in FY2024 and 9/18 in FY2023, with no parser errors for those documents. The 2023 period appears in both filings. This two-document test does not establish coverage or accuracy for every filer or taxonomy.

What you receive

Each processed document produces one dataset item with rowType, status, stable sourceUrl, sanitized finalUrl, and:

  • document: exact received-byte SHA-256, byte length, parser version and source provenance. Temporary delivery query strings and fragments are removed.
  • facts: concept and expanded namespace, raw text, normalized numericValue, exact decimal string numericValueExact, context ID, start/end/instant, dimensions, entity identifier, declared company number, unit, sign, scale and decimals. Invalid or nil facts have null numeric values. Use the exact string when downstream numeric precision matters.
  • periods: average employees per entity and undimensioned duration period, with supporting factIndexes and contextIds. Missing, invalid or ambiguous results have null averageEmployees. Only AverageNumberEmployeesDuringPeriod in dated FRC core namespaces with a simple xbrli:pure unit is selected. Dimensional facts remain in facts.
  • errors: structured diagnostics. A fetch/parse failure has status: "error" and empty fact/period collections. A parsed document without usable numeric facts has status: "no_usable_facts". A parsed status does not mean every fact is supported; inspect individual statuses and diagnostics.

CSV exports include successfully persisted documents with usable numeric facts. Nested CSV fields are JSON strings and nulls are blank. Duplicate identical employee facts are accepted; competing values or incompatible candidates are marked ambiguous. Filings remain separate: there is no cross-filing deduplication or restatement selection.

Limits and coverage

LimitValue
Documents1–20, fetched sequentially
Bytes per document2 MiB default; 5 MiB maximum
Total fetch time per document20 seconds default; 30 seconds maximum
RedirectsAt most 3, destination checked
XMLUTF-8; depth 128; 100,000 nodes; no DTD/entity declarations
Numeric factsAt most 5,000 per document
Numeric raw text / normalized fact output5 MiB / 8 MiB per document
Combined CSV exports10 MiB

Initial URLs must use HTTPS on find-and-update.company-information.service.gov.uk, with /company/{8-character-number}/filing-history/{document-id}/document?format=xhtml and optional download=1. Redirects allow that document path or https://s3.eu-west-2.amazonaws.com/document-api-images-live.ch.gov.uk/docs/. Other destinations and compressed responses are rejected. Exceeding a limit fails explicitly; earlier dataset items may remain, but truncated CSV files are not exported.

Supports ix:nonFraction in the 2008/2013 namespaces and conservative numeric forms for numcommadot, numdotcomma, numspacedot, numspacecomma, numdash, num-dot-decimal, num-comma-decimal, and zero-dash. Scale is limited to −12 through 12. Unsupported formats, continued numeric facts and inline fractions are marked invalid.

Company search, filing discovery, bulk ZIPs, PDFs and officer/PSC extraction are outside scope. Turnover is not inferred, and registered addresses are not substituted for trading addresses. This is an extraction tool, not a complete XBRL validator or standardized accounting model.

Pricing

See Actor pricing for the configured price. When pay-per-event pricing is configured, one document-parsed event applies to a useful parsed document containing at least one supported numeric fact. Fetch/parse failures and zero-usable-fact documents request no event. There is no per-fact or start event. Reaching the event limit stops further fetching; a document not persisted at the cap is excluded from CSV exports.