Italy Open Data Scraper - dati.gov.it CKAN Catalog avatar

Italy Open Data Scraper - dati.gov.it CKAN Catalog

Pricing

Pay per event

Go to Apify Store
Italy Open Data Scraper - dati.gov.it CKAN Catalog

Italy Open Data Scraper - dati.gov.it CKAN Catalog

Scrape all 65,960 datasets from Italy's national open data portal dati.gov.it: full DCAT-AP_IT metadata, distributions, 399 publishers, licences and live link checks.

Pricing

Pay per event

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

ParseForge

Italy Open Data Scraper - dati.gov.it CKAN Catalog

Scrape all 65,960 datasets on dati.gov.it, Italy's national open data portal, in one run. Every row carries the complete DCAT-AP_IT record: publishing organization, the data holder with its IPA code, EU themes, licence, update frequency, temporal and geographic coverage, and every downloadable file. No login, no API key. Export to CSV, JSON, Excel, or XML.

dati.gov.it is a CKAN catalog that harvests 328 regional, municipal and ministerial data portals, and its own search page renders results in the browser a page at a time. This reads the catalog API directly, filters on the fields the portal actually indexes, and returns each match as one flat row with 69 columns.

Who uses itWhat they scrape dati.gov.it for
Data engineersA machine-readable index of every open dataset Italy publishes, refreshed on a schedule
GIS and geospatial teamsThe 6,208 datasets with a spatial extent, their bounding boxes and their WMS, SHP and GeoJSON files
Civic tech and journalistsWhat each comune, regione and ministry publishes, and which of their links have gone dead
ResearchersCorpora filtered by theme, licence, publisher or date, with EuroVoc subthemes already resolved
Procurement and policy analystsWhich public bodies publish what, under which licence, and how often they promise to update it

What it does

This Actor reads the dati.gov.it CKAN catalog and returns each dataset as a flat row. Every dataset carries:

  • 🧾 Core fields: title, description, slug, portal URL, publishing organization, national identifier and both metadata dates.
  • 🇮🇹 Italian DCAT-AP_IT fields: the titolare del dato with its IPA code, the DCAT publisher and creator, conformsTo, access rights and the source catalog the record was harvested from.
  • 🏷️ Classification: EU data themes with English labels, EuroVoc subthemes, CKAN groups, tags, normalised licence code and update frequency.
  • 🗺️ Coverage: temporal start and end, temporal resolution, geographic name, GeoNames link, spatial type and bounding box.
  • 📦 Files: how many distributions, which formats, total size in bytes, and the primary download URL. Tick one box to get a full row per file.
  • 🔗 Optional live checks: whether each file still downloads, and what the bytes actually are versus what the portal claims.

Results export to CSV, JSON, Excel, or XML, or stream from the API.

What you can do with Italy open data

🗂️ Build a searchable index of Italian public data.

Sweep the whole catalog once, then re-run filtered by modifiedFrom on a schedule to pick up only what changed. Every row has a stable datasetId and a portal URL.

🗺️ Find the geospatial layers.

Turn on Geospatial datasets only and filter formats to GeoJSON, Shapefile, GML or WMS. You get the bounding box and the service endpoint on the same row.

🔍 Audit a public body's open data.

Filter by organization slug or by data holder name, add the publisher profile, and you have every dataset that body publishes with its certified email, website and region.

🧪 Check the catalog is telling the truth.

Turn on the two live checks and the run reports which download links are dead and which files are not the format the metadata declares.

Why choose this scraper

What you get
69 columns per datasetThe full DCAT-AP_IT record flattened, not the raw CKAN package dumped into a cell.
Filters measured against totalsEvery dropdown value was checked against the live catalog count, so nothing silently returns zero.
Spelling variants folded inThe portal writes CC BY 4.0 four ways and uses 92 format spellings. One choice matches all of them.
Nine extra row typesDistributions, organizations, themes, tags, licences, formats, data holders, publishers and source catalogs.
Live link and byte checksOptional. The portal harvests 328 catalogs and never revalidates a link.
Four export formatsCSV, JSON, Excel, and XML, from the dashboard or the API.

How it compares

Two other Actors read dati.gov.it and both return the CKAN package roughly as it arrives, billed per item. The generic CKAN exporters do the same for any portal. The difference here is that the Italian-specific fields are parsed out rather than left nested, the filter vocabularies were measured against live counts instead of copied from the CKAN docs, and the two live checks exist at all.

FeatureParseForgebenthepythondev/italy-dati-gov-scraperdatapilot/open-data-portal-harvesterstraightforward_hydra/ckan-open-data-scraper
DCAT-AP_IT fields parsed out (holder, IPA code, themes, EuroVoc)Yes, 69 columnsRaw packageRaw packageRaw package
Filters verified against catalog totalsYes, all of themNoNoNo
Licence and format spelling variants foldedYesNoNoNo
Dead-link check on distributionsOptionalNoNoNo
Byte probe: does the file match its declared formatOptionalNoNoNo
Directory exports (holders, publishers, source catalogs)9 row typesNoNoNo
Price per 1,000 rows$7.00Not published$2.00$2.00

The generic exporters are cheaper per row. If all you need is the raw CKAN package for one portal, use them. This one costs more because the row is parsed, the vocabularies are resolved, and the optional checks go out to the publishers' own servers.

What a dataset looks like

{
"datasetId": "d74a05d3-8056-4291-91f9-34a8168c5b24",
"name": "rilevazione-dei-prezzi-al-consumo-del-comune-di-firenze-del-2018-gennaio1",
"url": "https://www.dati.gov.it/dataset/rilevazione-dei-prezzi-al-consumo-del-comune-di-firenze-del-2018-gennaio1",
"title": "Rilevazione dei prezzi al consumo del Comune di Firenze del 2018 - Gennaio",
"description": "Il dataset contiene i dati relativi alla rilevazione dei prezzi al consumo del Comune di Firenze del 2018 nel mese di gennaio. Il separatore delle risorse CSV è \";\" (lett. punto e virgola).",
"organization": "regione-toscana",
"organizationTitle": "Regione Toscana",
"organizationId": "9bcd4050-eecd-400b-8196-cd4db7b38c58",
"holderName": "Regione Toscana",
"holderIdentifier": "r_toscan",
"publisherName": "Comune di Firenze",
"publisherIdentifier": "c_d612",
"publisherUri": "Not Disclosed",
"publisherEmail": "Not Disclosed",
"creatorNames": ["Comune di Firenze - Direzione Generale - Servizio Pianificazione, Controllo e Statistica"],
"rightsHolder": "Not Disclosed",
"themeCodes": ["SOCI"],
"themeLabels": ["Population and society"],
"subThemes": [],
"groups": ["societa"],
"groupTitles": ["Società"],
"tags": ["08prezzi", "comune-firenze", "consumo", "finanze", "indici", "opendata", "prezzi"],
"tagCount": 7,
"licenceCode": "CC-BY-4.0",
"licenceLabel": "Creative Commons Attribution 4.0",
"licenceRaw": "Creative Commons Attribuzione 4.0 Internazionale (CC BY 4.0)",
"licenceUrl": "Not Disclosed",
"isOpenLicence": "No",
"accessRights": "Not Disclosed",
"updateFrequency": "NOT_PLANNED",
"updateFrequencyLabel": "Not planned",
"updateFrequencyRaw": "NOT_PLANNED",
"language": "ITA",
"nationalIdentifier": "c_d612:D.6052",
"alternateIdentifiers": [],
"conformsTo": ["http://dati.gov.it/onto/dcatapit#"],
"issued": "2018-01-31",
"modified": "2018-01-31",
"metadataCreated": "2026-06-28T03:40:05.254742",
"metadataModified": "2026-08-22T16:56:52.194367",
"temporalStart": "2018-01-31",
"temporalEnd": "Not Disclosed",
"temporalResolution": "Not Disclosed",
"geographicalName": "ITA_FLR",
"geoNamesUrl": "https://www.geonames.org/6542285",
"spatialType": "N/A",
"boundingBox": "N/A",
"contactName": "Comune di Firenze",
"contactEmail": "opendata@comune.fi.it",
"contactUri": "https://dati.toscana.it/organization/881795ff-b47f-4c85-923c-b67b2d86fd8d",
"landingPage": "https://dati.toscana.it/dataset/rilevazione-dei-prezzi-al-consumo-del-comune-di-firenze-del-2018-gennaio#",
"sourceUri": "https://dati.toscana.it/dataset/rilevazione-dei-prezzi-al-consumo-del-comune-di-firenze-del-2018-gennaio",
"sourceCatalogTitle": "Not Disclosed",
"sourceCatalogHomepage": "Not Disclosed",
"sourceCatalogPublisher": "Not Disclosed",
"sourceCatalogModified": "Not Disclosed",
"sourceCatalogType": "Not Disclosed",
"harvestSourceTitle": "Regione Toscana",
"isNativeDataset": "No",
"resourceCount": 1,
"resourceFormats": ["CSV"],
"hasMachineReadable": "Yes",
"totalSizeBytes": 688,
"primaryDownloadUrl": "https://data.comune.fi.it/datastore/download.php?id=6052&type=1&format=csv",
"viewCount": "N/A",
"downloadCount": "N/A",
"searchCount": "N/A",
"rowType": "dataset",
"scrapedAt": "2026-08-27T16:15:56.950Z"
}

Fields the portal does not fill for a given dataset come back as Not Disclosed; fields that do not apply come back as N/A. Nothing is ever null.

Configure the run

Leave everything empty and the Actor sweeps the whole catalog, most recently changed first. Add filters to narrow it, or paste dataset URLs to fetch specific records. Filters run in the portal's own index, so only matching datasets are read and billed. The Input tab lists every parameter.

Sweep the catalog for a theme and keep only datasets that ship a real data file:

{ "themes": ["ambiente"], "formats": ["CSV", "JSON", "GeoJSON"], "maxItems": 2000 }

Everything one city publishes, with its files and the publisher profile:

{ "organizations": ["comune-di-milano"], "includeResources": true, "includeOrganizationProfile": true, "maxItems": 5000 }

Audit the geospatial layers and check the downloads still work:

{ "onlyGeospatial": true, "formats": ["SHP", "WMS", "GeoJSON"], "includeLinkCheck": true, "includeFileProbe": true, "maxItems": 300 }

Pricing

Pay-per-event: $0.007 per dataset, plus $0.004 per catalog page read (one page covers 200 datasets) and a $0.054 run-start fee. Extra row types and the optional checks have their own prices and are charged only when they return something. You pay only for rows written to your dataset.

Datasets collectedApproximate cost
100$0.76
1,000$7.07
10,000$70.25

At higher monthly volume the per-row rate drops on Apify's usual tiers. New Apify accounts start with $5 in free credit.

Free users

Free-plan runs return up to 10 rows as a preview. Upgrade your Apify plan to collect up to 1,000,000 rows per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Open the Italy Open Data Scraper.
  3. Pick your themes, organizations or formats, set maxItems, tick any extra row types, and click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to Italy's open data catalog through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/dati-gov-it-italy-open-data-scraper"

Then prompt it in plain language:

  • "Find Italian open datasets about air quality published as CSV and list who publishes them."
  • "List every dataset the Comune di Milano publishes with a CC BY licence, newest first."
  • "Pull the geospatial datasets from Regione Toscana and tell me which download links are dead."

Copy this into ChatGPT, Claude, or Cursor to start:

Use the Apify Actor "parseforge/dati-gov-it-italy-open-data-scraper" to search Italy's national open data catalog dati.gov.it. Input: { "searchTerms": ["<keyword>"], "themes": ["<governo|ambiente|salute|trasporti|...>"], "organizations": ["<slug>"], "formats": ["CSV"], "maxItems": <n> }. It returns title, url, organization, holderName, themeLabels, licenceLabel, updateFrequencyLabel, resourceCount, resourceFormats and primaryDownloadUrl per dataset. Call it with the ApifyClient and my APIFY_TOKEN.

Troubleshooting

Why am I getting no results?

A filter value the portal does not know returns zero rows rather than an error. Organization slugs, holder names, publisher names and source catalog titles must match exactly; run the matching directory box once to see the real values with their counts.

Why fewer rows than I asked for?

maxItems is a cap on total rows, not on datasets. If you also ticked distribution rows or a directory, they draw from the same budget. Directories are collected smallest first, so tags, which run to thousands of values, take whatever is left.

Why is a field empty?

Not Disclosed means the publisher did not fill that DCAT field, and N/A means it does not apply to that record. Both are the record's real state. isOpenLicence says No on most rows because the portal's own open flag is set on only 4,026 of 65,960 datasets even though 61,448 carry a CC BY licence; use the licenceCode column instead.

Why is a download link reported dead?

The portal harvests metadata from 328 separate catalogs and never revalidates the URLs. Some publishers' servers are also firewalled against non-Italian traffic, which the check reports as a timeout rather than an HTTP error.

Why is the run slow?

The catalog API answers a 200-dataset page in about eight seconds, so a full sweep of 65,960 datasets is bounded by paging. The two optional live checks go out to the publishers' own servers and are much slower than the catalog itself; raise maxItems gradually or lower maxResourcesPerDataset.

FAQ

QuestionAnswer
Do I need an account or API key for dati.gov.it?No. The catalog API is open and anonymous, and this Actor uses no proxy.
How many datasets are there?65,960 as of 27 August 2026, from 447 registered organizations and 328 source catalogs.
What is the difference between organization and data holder?The organization is the account that published the record on the portal, usually a region. The titolare del dato is the body legally responsible for the data, often a single comune. There are 1,763 holders against 447 organizations.
Can I get the actual data files, not just the metadata?The rows give you every download URL. Tick the byte probe to also get each file's real type, delimiter and column headers without downloading it whole.
Can I filter by region or province?Not directly: the portal has no region filter. Filter by organization slug, or turn on the organizations directory, which carries the Italian region for 364 of 447 bodies.
Why do some formats appear twice in the formats directory?The index is case sensitive and the catalog stores 92 spellings for about 34 real formats. GML matches 1,521 datasets and gml another 92. The format filter folds them together for you.
Can I run a raw CKAN query?Yes. customFilterQuery takes a Solr fq clause and is ANDed with the other filters.
How many rows per run?Free plan: 10. Paid: up to 1,000,000, bounded by what the catalog and your filters return.
Is this an official AgID product?No. It is unofficial and reads only the public dati.gov.it catalog API.

Browse the full ParseForge collection for more scrapers.

🆘 Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by AgID or the Italian government. It collects only publicly available dati.gov.it catalog metadata. You are responsible for using the data in compliance with the portal's terms, the licence on each dataset, and applicable laws including GDPR, CCPA, and PIPL. Do not use it to identify, profile, or target individuals.