GovXRay Scraper — City Government Fiscal Data
Pricing
from $7.00 / 1,000 result scrapeds
GovXRay Scraper — City Government Fiscal Data
Scrape consolidated government finance profiles for 240+ world cities from GovXRay.com: spending & revenue per capita, deficit, public debt & assets by government tier, spending and revenue by category, and GovXRay's module index (economy, housing, healthcare, education). No login, no cookies.
Pricing
from $7.00 / 1,000 result scrapeds
Rating
0.0
(0)
Developer
Studio Amba
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 days ago
Last modified
Categories
Share
GovXRay Scraper
Scrape city government fiscal-transparency data from GovXRay.com — a "Government Financial X-Ray" covering 118 cities across 19 countries. For each city you get per-capita public finances broken down by government tier (federal/national, state/region, municipal, and sometimes social security), spending and revenue by category (where GovXRay has that data connected), and the site's own "what stands out" narrative callouts. No login, no cookies.
Why use this actor?
GovXRay does the hard work of normalizing public finance data that otherwise lives scattered across national statistics offices, Eurostat/OECD tables and municipal budget documents into one consistent per-capita view per city. This actor turns that into structured records you can drop into a spreadsheet, a BI tool, or a research database.
Typical users: policy researchers and think tanks, journalists covering municipal finance or debt, urban economists, relocation and site-selection analysts comparing tax burden across cities, and public-finance consultants who need a quick fiscal-health snapshot for many cities at once.
What you get
- Municipal-tier figures — revenue and spending per capita for the city government alone, plus the resulting surplus/deficit, fiscal year and accounting status (audited vs. running budget).
- Multi-tier consolidation — where GovXRay has full government-tier data connected for a city, also the consolidated per-resident view (what a resident pays in and receives across every displayed tier), spending broken down by category (Social, Santé, Éducation, ...) and revenue broken down by category (Personal income taxes, Sales taxes & excise, ...).
- Balance Sheet by tier — revenue, spending and balance per capita for every government tier that applies to the city (national, sub-national "including ..." breakdowns, social security, municipal), each with its own fiscal year, accounting status and source.
- Assets and debt by tier — assets, debt and net position per resident for each tier GovXRay publishes (national, regional, local sector, social security), plus the debt the city itself publishes for its own unit.
- Profile score — GovXRay removed its composite profile score from every
page in its 24 September 2026 relaunch, after dropping the letter grade and
the "what stands out" highlights earlier that month.
fiscalHealthGrade,fiscalHealthCompositeScoreandhighlightsstay empty unless the site brings them back. Spending, funding, the module index and assets/debt moved to the city's own section pages; the actor reads those pages too, at no extra cost per result. - Module directory — the 30-module deep-dive index GovXRay tracks for every city, grouped by category (Public finances, Economy, Health, ...), which modules were removed, and how many are written for this city.
- City vitals — population, country GDP per capita (World Bank, USD), and the government-tier chain that applies to the city.
- No login, no cookies — nothing to configure beyond a list of cities.
How to scrape GovXRay data
- Add the actor to your Apify account.
- Enter a list of Cities (names or GovXRay slugs, e.g.
Berlin,new_york_city,Copenhague) — or leave it empty for a default set of major world cities, or pass["all"]to scrape every city GovXRay covers (476 as of September 2026). - Set Max Cities to how many you want in this run.
- Run it. Download the results as JSON, CSV, Excel, or feed them to an API.
The pages are plain server-rendered HTML, so the actor reads them with a plain HTTP client, four cities at a time. It tries a direct connection first, then your Apify proxy, then Apify's datacenter proxy, each on a fresh session. A Bright Data key is optional: when one is set, Bright Data is the last route and is only used for a page every other route refused. No key is needed to run the actor.
Input
| Field | Type | Required | Description |
|---|---|---|---|
cities | Array of strings | No | City names or GovXRay slugs to scrape (default: 10 major world cities). Pass ["all"] to scrape every city GovXRay covers. |
maxResults | Integer | No | Maximum number of cities to scrape in this run (default: 20). |
brightDataApiKey | String | No | Optional Bright Data Web Unlocker key, used only for a page every other route refused. Falls back to the BRIGHT_DATA_API_KEY environment variable. |
proxyConfiguration | Object | No | Apify proxy settings. Used when the direct connection is refused. |
Leave everything empty and the actor scrapes a default set of major cities
(New York City, London, Paris, Berlin, Copenhague, Zürich, Tokyo, Sydney,
Toronto, Madrid), so an empty input {} still returns data.
Output
Each result is one city's fiscal profile. Key fields:
| Field | Type | Example |
|---|---|---|
city | String | "Copenhague" |
country | String | "Danemark" |
currency | String | "DKK" |
population | Number | 659350 |
cityOnlyRevenuePerCapita / cityOnlySpendingPerCapita | Number | 93112 / 88869 |
govHierarchy | Array | ["Denmark (all tiers)", "Copenhague"] |
consolidatedRevenuePerCapita / consolidatedSpendingPerCapita | Number | 279308 / 280583 (only where GovXRay has multi-tier data connected) |
spendingByCategory | Object | {"Social": 10455, "HEALTH": 3925, ...} |
revenueByCategory | Object | {"Personal income taxes": 5964, "Sales taxes & excise": 4191, ...} |
balanceSheetTiers | Array | Per-tier revenue/spending/balance per capita, fiscal year, status, source |
assetsAndDebtTiers | Array | Per-tier assets, debt and net position per resident |
municipalDebtPerCapita | Number | Debt the city publishes for its own unit, per resident |
moduleCategories | Object | The 30-module directory grouped by category |
modulesRetired | Array | Modules GovXRay has removed |
modulesWrittenCount / modulesTotalCount | Number | 4 / 30 |
namedGapsCount | Number | Data gaps GovXRay names for this city |
url | String | Source page URL |
scrapedAt | String | ISO 8601 timestamp |
Example output (abridged — real run, Berlin, 2026-09-19)
{"city": "Berlin","citySlug": "berlin_de","country": "Allemagne","url": "https://govxray.com/city/berlin_de/","currency": "EUR","population": 3685265,"populationScope": "Municipal","cityOnlyRevenuePerCapita": 10044,"cityOnlySpendingPerCapita": 10874,"cityOnlyFiscalYear": 2024,"cityOnlyStatus": "AUDITED","countryGdpPerCapitaUsd": 60496,"govHierarchy": ["Germany (all tiers)","Berlin"],"consolidatedRevenuePerCapita": 33967,"consolidatedSpendingPerCapita": 36140,"consolidatedGapPerCapita": 2173,"balanceSheetTiers": [{"tier": "Germany (all tiers)","isSubTier": false,"revenuePerCapita": 24215,"spendingPerCapita": 25594,"balancePerCapita": -1379,"fiscalYear": 2024,"status": "AUDITED","source": "API"},{"tier": "including central government","isSubTier": true,"revenuePerCapita": 6941,"spendingPerCapita": 7670,"balancePerCapita": -729,"fiscalYear": 2024,"status": "audited","source": "COUNTRY · Eurostat"},{"tier": "Berlin","isSubTier": false,"revenuePerCapita": 10044,"spendingPerCapita": 10874,"balancePerCapita": -830,"fiscalYear": 2024,"status": "AUDITED","source": "document"}],"assetsAndDebtTiers": [{"tier": "FEDERAL","assetsPerCapita": 8097,"debtPerCapita": 20649,"netPerCapita": -12552},{"tier": "Berlin","assetsPerCapita": 20871,"debtPerCapita": 18168,"netPerCapita": 2703}],"municipalDebtPerCapita": 17326,"modulesWrittenCount": 4,"modulesTotalCount": 30,"namedGapsCount": 65}
How it works
GovXRay's city pages are fully server-rendered — every figure is present in the raw HTML of a plain page load, French-only, no hidden JSON API. The actor:
- Reads the live
cities.jsonindex (replacedsitemap.xmlin the site's August 2026 "v17" relaunch) to build the current city list, filtering out the handful of non-city navigation pages the same file also lists. - Fetches each requested city's page, four at a time, on the first route that serves it (direct, your proxy, Apify datacenter, then Bright Data if a key is set). A page is accepted on its content, not on its HTTP status. If the run gets close to its timeout it stops starting new pages, keeps what it has and says how many cities it did not reach.
- Converts the HTML into an ordered list of text blocks that mirrors the
page's visual structure, then walks it with a small parser anchored on
known French section labels (Population, Fiche de santé, DÉPENSES,
RECETTES, Balance Sheet, Ce qui ressort, Les 29 modules) to pull out
structured fields. It also reads the fiscal-health composite score out
of a
titleattribute the visible text doesn't carry.
Two page layouts exist depending on how much government-tier data GovXRay
has connected for a given city: cities with full consolidation show a "tous
paliers" (all tiers) resident view and a spending/revenue category
breakdown; cities with only municipal-level data show a "ville seule"
(city-only) resident view and no category breakdown. The parser detects and
handles both — consolidatedRevenuePerCapita, consolidatedSpendingPerCapita
and spendingByCategory/revenueByCategory are simply absent for
city-only cities, since GovXRay itself doesn't publish that data for them.
Verified against New York City, Copenhague and Zürich (full consolidation)
and Berlin (city-only) — 96-100% field coverage on all four.
Cost estimate
One plain HTTP request per city page, plus one for cities.json per run.
A run of 50 cities is roughly 51 requests and takes seconds, not minutes.
Bright Data is only charged to your own key, and only for pages every other
route refused.
Usage cost only settles after the run reports SUCCEEDED — reading the
dataset mid-run will undercount what the run actually cost.
Limitations
- Covers the city root/scorecard page only — GovXRay also has deeper per-module pages per city (a full economy or housing breakdown with many more metrics); this actor surfaces the module directory rather than crawling all of them.
- GovXRay's own data coverage varies by city — some cities only have municipal-tier figures connected, with no consolidated multi-tier view or category breakdown; the actor reflects whatever GovXRay currently publishes and doesn't fill gaps.
- Figures are GovXRay's own estimates/aggregations from public sources (Eurostat, national accounts, curated municipal audits) — treat this as a research and comparison tool, not an official audited financial statement.
- GovXRay dropped a number of cities (including Antwerp and Brussels) in
its August 2026 relaunch; only the 118 cities in the live
cities.jsonindex are scrapeable.
Related scrapers
Studio AMBA also publishes scrapers for other European/global regulatory
and public-data sources: belgian-procurement-scraper,
ted-eu-procurement-scraper, eurlex-scraper, handelsregister-scraper,
and kbo-enrichment for company/registry data alongside government data.