Australia Building Approvals by SA2 - ABS, Per Observation
Pricing
from $33.50 / 1,000 approval observations
Australia Building Approvals by SA2 - ABS, Per Observation
Australia building approvals at SA2 grain (ABS Data API, BA_SA2 SDMX-CSV) as per-observation records - measure, SA2 region, building type, period, value. 10.8M observations; required key + start-period filter and record cap keep a run buyable. CC BY 4.0. $0.05/record.
Pricing
from $33.50 / 1,000 approval observations
Rating
0.0
(0)
Developer
NexGen Signal
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Australia's building approvals at SA2 (small-area) grain from the Australian Bureau of Statistics (ABS), as clean per-observation records - one SA2 x building/work/value measure x period, with the observation value.
What one record represents
The source is the ABS Data API at data.api.abs.gov.au, read as keyless SDMX-CSV from the ABS,BA_SA2,2.0.0 dataflow. Each record is one observation: measure, sector, work type, building type, region type, SA2 region, frequency, period and the observation value.
Coverage, the partition inputs and the default run
The complete dataflow is 10,800,070 observations (2.66 GB) - far too large to pull in one run. So the run is scoped by two inputs plus the cap: an SDMX key filter (default measure 1; the dot-separated key is MEASURE.SECTOR.WORK_TYPE.BUILDING_TYPE.REGION_TYPE.REGION.FREQ, blank positions meaning all) and a Start period (default 2025). The Actor streams the filtered SDMX-CSV and stops at Maximum records, so a default run streams the first 1,000 rows for about US$50. Widen the key or the period and raise the cap to pull more. Only statistical aggregates are emitted - no microdata.
Licence
Australian Bureau of Statistics (ABS). Creative Commons Attribution 4.0 International (CC BY 4.0): commercial reuse and adaptation permitted with attribution to the ABS. Only statistical aggregates are emitted; logos, unit-record microdata and identified third-party material are excluded.
Sibling Actors
The fleet's US building-permits construction-leads cell is US permit records at address grain (a different country and grain); this is Australian small-area aggregates. The companion ABS goods-trade cell shares this door but is a different dataflow. No row overlap.
Fields, scheduling and integration
Every field in the record is either a source-native identifier, a source-native attribute, or one of the six
provenance fields (source, source_dataset, licence, attribution, caveat, observed_at) the fleet attaches to
every record. Nothing is derived or inferred beyond the small, documented transforms noted above, and nothing is
dropped silently - the handling section spells out exactly what is excluded and why. The grain is one record per the
natural unit of the source, which keeps each row independently meaningful, keeps the key stable across runs so
re-running is a cheap upsert rather than a re-import, and lets you aggregate up to whatever unit you need without
unpicking a pre-joined table.
Because the source republishes on its own cadence, a scheduled run keeps a downstream table current: new and changed
records upsert over the old ones on the stable key, and the observed_at stamp tells you when each was last seen live.
Set Maximum records low to sample the shape of the data cheaply, then raise it once the cell fits your use; the
Actor streams or partitions its source, so memory stays flat regardless of how many records you request, and you are
billed only for what is delivered. The output is a flat table of typed records, so it drops straight into whatever you
already use: load the run's dataset over the API or an export, key on the record id, and upsert. Because identifiers
are preserved exactly as the source publishes them, joins across the fleet's cells - and onto your own systems - work
without a mapping layer. There is no subscription and no minimum: the per-record price and the record cap together
mean the spend on any run is known in advance and matched exactly to the data you receive.
Reconciling counts honestly
Where the live count differs from any previously published figure, the live measure is the honest one and is what this listing quotes; sources re-issue and re-version their data over time. The run receipt always states what was actually delivered and charged and confirms the two agree, so every run is auditable against itself regardless of what any external index expected.
Scaling, scheduling and support
A common pattern is a light scheduled run that pulls the newest slice into a staging table, then a merge on the stable key into the table your product reads, so you never re-pay for rows you already hold and your history grows cleanly over time. Because the record shape assumes no particular warehouse, language or tool, the integration work is a load and a merge, not a cleaning project: the same code path handles a 40-row sample and a full pull, and the only thing that changes between them is the record cap. If you only need a slice, the cap and any partition or filter inputs bound the run precisely, so a targeted pull costs cents rather than the price of the whole set, and a broad pull is simply a higher cap left to run. Nothing about the delivery is subscription-gated: each run stands alone, priced at exactly the records it returns, so you can dial spend up or down run by run as your needs change, and a scheduled cadence keeps a downstream table current without any standing commitment. When the source publishes a correction or a new period, the next run picks it up and upserts it over the stale row on the same key, so the table you maintain stays both complete and current with no manual reconciliation.
Provenance and compliance
Every run reads the door host's robots.txt at runtime and records the result (URL, status, byte length and, where
a policy is served, its SHA-256) in the run's RUN_RECEIPT. Where the host serves no applicable policy - a 404, a
403, or a homepage redirect - the gate records that as a flag and proceeds on the licence, which grants re-use; a flag
is never treated as permission in itself. The endpoint is keyless and the Actor reads only the public data door -
never a mirror, and it never bypasses a block.
Data quality and freshness
Numbers arrive as real numbers, booleans as real booleans, and every other value as a string or null, so the dataset
loads without a cleaning pass. Each record is keyed on a stable composite of the source's own identifiers, so it is
safe to diff, deduplicate or upsert. Every run re-reads the live door, so the data is as fresh as the source
publishes, and each record's observed_at stamp dates the snapshot. The receipt records how many rows were delivered
and charged and confirms charge_equals_delivered.
Billing, delivery and joins
Pricing is per record: you are billed only for records the Actor actually delivers, and the charge is raised after each record is pushed (push-then-charge), so a failed or empty run costs nothing. The Maximum records cap bounds every run, so spend is known before you start - sample cheaply, then raise it. Every record is a flat, typed object keyed on a stable id, so it loads without a cleaning pass, diffs cleanly between runs, and upserts into a table you keep over time; re-running keeps that table current without re-paying for rows you already hold, and each receipt reconciles delivered against charged. Because the source's own identifiers are preserved verbatim, the dataset joins onto other sources keyed on the same identifier.
Who buys this, and how they use it
This cell is bought by teams that need the source's published set as a typed, keyed table they can hold and refresh rather than a page they scrape: market- and macro-intelligence teams sizing and tracking a market, data engineers wiring a clean upstream feed into a warehouse, and compliance and research teams building on a stable identifier. The grain and the key are chosen so the output is a building block, not a one-off export - you run it on a schedule, keep the delta, and join it to your other sources on the identifiers it preserves verbatim. The spend on any run is the per-record price times the records delivered, matched exactly to what you receive.