EU Commission Open Data Index - CC BY 4.0
Pricing
from $33.50 / 1,000 eu dataset index records
EU Commission Open Data Index - CC BY 4.0
Catalogue of European Commission open datasets (CC BY 4.0) on data.europa.eu: search by text, EU theme (energy, health, transport...), publisher (Eurostat, JRC...) or download format, most recently modified first - title, dates, publisher, formats, download URLs. Pay per record.
Pricing
from $33.50 / 1,000 eu dataset index records
Rating
0.0
(0)
Developer
NexGen Signal
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
15 hours ago
Last modified
Categories
Share
An index of the European Commission's open datasets on data.europa.eu that carry a CC BY 4.0 distribution - one record per dataset. This is a catalogue of metadata: each record names a dataset, its theme, dates, publishing Directorate-General and its distributions (formats, licences, download URLs). The records point at the datasets; the product is the index itself.
What one record represents
The source is data.europa.eu (the EU Open Data Portal) hub search API. This cell is scoped to datasets
whose publisher is a European Commission Directorate-General (corporate-body-classification=DIR_GEN) and
which carry a CC BY 4.0 distribution. Each record is one dataset's metadata: its id, title, DCAT
data-theme categories, modified and issued dates, the publishing Directorate-General, the source catalogue,
the landing page and resource URI, and a summary of its distributions - how many, their distinct formats,
their distinct licence ids, and the download URLs.
Sample output

Real rows from a live run of this actor (first 5 rows, selected columns).
One full record from the same run, exactly as delivered:
{"dataset_id": "eu-food-additives","title": "EU Food Additives","categories": "TECH;AGRI","modified": "2024-08-13","issued": "2015-08-18","publisher_name": "Directorate-General for Health and Food Safety","catalog": "sante","landing_page": "https://food.ec.europa.eu/safety/food-improvement-agents/additives_en","resource_uri": "http://data.europa.eu/88u/dataset/eu-food-additives","distribution_count": 3,"distribution_formats": "JSON","distribution_licences": "CC_BY_4_0","distribution_download_urls": "https://developer.datalake.sante.service.ec.europa.eu/api-details#api=228d6fda-9092-4c25-af9a-d537666ed0e5&operation=66382a6d-6a0d-4c17-99a5-6156f6436ded https://api.datalake.sante.service.ec.europa.eu/food-additives/download?format=json&api-version=v2.0","licence": "CC BY 4.0","record_id": "eu-food-additives","source": "data.europa.eu (EU Open Data Portal)","source_dataset": "hub/search filter=dataset DIR_GEN CC_BY_4_0","attribution": "European Commission - data.europa.eu (EU Open Data Portal)","caveat": "One record per European Commission (Directorate-General) dataset on data.europa.eu that carries a CC BY 4.0 distribution. This is a catalogue of METADATA: dataset id, title, DCAT category, modified/issued dates, publishing Directorate-General, and distribution formats/licences/download URLs - the records point at the datasets, they are not the underlying data. Contact-point and maintainer person fields are structurally excluded; the publisher name is an organisation.","observed_at": "2026-09-25T17:32:09Z","licence_notice": "Dataset metadata from data.europa.eu (the EU Open Data Portal). This cell delivers ONLY datasets carrying a Creative Commons Attribution 4.0 (CC BY 4.0) distribution; the licence filter is baked into the query. Attribution required. The records are METADATA that point at the datasets; the product is the index itself."}
Search the catalogue
Question: Which European Commission energy datasets about electricity can be downloaded as CSV, most recently updated first?
{"searchText": "electricity", "categories": ["ENER"], "formats": ["CSV"], "newestFirst": true, "maxRecords": 10}
Run on 30 Sep 2026 the portal matched 10 datasets and the run returned all 10 (10 × $0.05 = $0.50 on Apify's Free plan):
| Title | Modified | Publisher | Formats |
|---|---|---|---|
| Final energy consumption in households per capita | 2026-07-21 | Eurostat | HTML;XML;CSV;TSV |
| Energy taxes | 2026-07-21 | Eurostat | CSV;HTML;XML;TSV |
| Primary energy consumption | 2026-07-21 | Eurostat | XML;HTML;TSV;CSV |
| Market share of the largest generator in the electricity market | 2026-06-02 | Eurostat | XML;HTML;TSV;CSV |
| Final energy consumption in households by type of fuel | 2026-06-02 | Eurostat | XML;TSV;CSV;HTML |
| Waste electrical and electronic equipment (WEEE) by waste management operations | 2026-04-08 | Eurostat | XML;CSV;TSV;HTML |
Search options (all optional; a record must match every option you set, and several values in one option match any of them):
- Search text (
searchText) — the portal's own full-text search over titles and descriptions. - Themes (
categories) — EU data-theme codes:AGRI,ECON,EDUC,ENER,ENVI,GOVE,HEAL,INTR,JUST,REGI,SOCI,TECH,TRAN. Several = datasets tagged with all of them (the portal combines values this way). - Publisher (
publisher) — one Commission catalogue:estatEurostat,jrcJoint Research Centre,commuDG Communication (Eurobarometer),cnect,empl,grow,comp,sante,digit,just,ener,rtd. - Download formats (
formats) — e.g.CSV,JSON,XLSX; several = datasets offering all of them. - Most recently modified first (
newestFirst) — the portal's sort by modification date.
Every option goes to the data.europa.eu search API, so filtering happens before Maximum records; the status message gives the portal's match count. "No datasets match" (nothing charged) is a successful run; if the portal cannot be reached the run fails as a source failure. With no search option the request is exactly the one the actor always sent. The licence (CC BY 4.0) and Commission-publisher filters always apply.
Coverage and volume
The live filtered set is 10,536 datasets (measured at build time: corporate-body-classification=DIR_GEN
AND license=CC_BY_4_0). The European Commission is organised into Directorates-General, so this classification
is precisely "datasets published by a Commission DG."
Sol's Wave-3 index put this door at 10,533 CC BY 4.0 Commission records; measured live at build time the filter returns 10,536 - the live figure is what this listing quotes.
The Actor pages the hub search API 100 datasets at a time and stops as soon as your Maximum records cap is met.
Licence and attribution
This cell delivers only datasets that carry a CC BY 4.0 distribution. The licence filter
(license=CC_BY_4_0) is baked into the query, not a buyer option, and a per-record assertion rejects
any dataset whose distributions do not include CC BY 4.0 (verified with a planted-row test). The full notice
travels on every record:
Dataset metadata from data.europa.eu (the EU Open Data Portal). This cell delivers ONLY datasets carrying a Creative Commons Attribution 4.0 (CC BY 4.0) distribution; the licence filter is baked into the query. Attribution required. The records are METADATA that point at the datasets; the product is the index itself.
The licence field is set to CC BY 4.0 on every record, and distribution_licences lists the distinct
licence ids the dataset's distributions actually carry (a dataset may offer additional distributions under
other reuse terms; the CC BY 4.0 one is what qualifies it here). Attribution is to the European Commission and
the publishing Directorate-General, named on each record.
Person-data policy
data.europa.eu dataset records carry a contact_point (and sometimes a maintainer) that can hold a person's
name, email or phone. This Actor structurally excludes the contact point and maintainer: they are never
read. The publisher_name that is kept is the Directorate-General - an organisation, never a person - and a
per-record assertion rejects any contact/maintainer/person field (verified with a planted-field test). No
natural-person data is processed.
Interpretation caveat
One record per European Commission (Directorate-General) dataset on data.europa.eu that carries a CC BY 4.0 distribution. A catalogue of METADATA: dataset id, title, DCAT category, modified/issued dates, publishing Directorate-General, and distribution formats/licences/download URLs - the records point at the datasets, they are not the underlying data. Contact-point and maintainer person fields are structurally excluded.
Values are reproduced verbatim from the API; the Actor never rewrites a field. Titles are multilingual in the
source; the record carries the English title where available, otherwise the first available language. The
download URLs are capped at 25 per record to keep rows bounded; distribution_count reports the true total.
Data quality and freshness
distribution_count is a real number; dates are the portal's own modified/issued stamps. The record_id
is the dataset id, so the index is safe to diff, deduplicate or upsert. Every run re-reads the live API, so the
index is as fresh as data.europa.eu publishes, and each record's observed_at stamp dates the snapshot. The
run's RUN_RECEIPT records the live match count alongside how many records were delivered and charged.
Provenance and compliance
Every run reads data.europa.eu/robots.txt at runtime; the gate result (URL, status, byte length, SHA-256 of
the policy) is written to the run's RUN_RECEIPT, and the hub search path is confirmed crawlable before any
data request. The API is keyless. The Actor never bypasses a block or fetches through a mirror.
Inputs
Default changed on 30 Sep 2026: if you leave maxRecords out, a run now returns up to 10 records (it was 500). Set maxRecords yourself to get more — the maximum is unchanged.
- Maximum records (
maxRecords) - hard cap on dataset-index records delivered and billed. - Search options - see Search the catalogue above; all optional.
Output
Records land in the Actor's default dataset and export as JSON, CSV, Excel or via the Apify API. A tabular overview view surfaces dataset id, title, categories, modified date, publishing Directorate-General, distribution formats and licence.
Fields in detail
The record leads with dataset_id and title, then categories (DCAT data-theme ids), modified, issued,
publisher_name (the Directorate-General), catalog, landing_page, resource_uri, and the distribution
summary - distribution_count, distribution_formats, distribution_licences, distribution_download_urls -
plus the licence label and full licence_notice. The provenance block closes every record.
Typical uses
Data brokers, open-data teams and procurement analysts use this cell to source reusable Commission datasets:
one queryable index of what the EU's Directorates-General publish under CC BY 4.0, with the theme to filter by
topic, the modified date to spot fresh releases, and the download URLs to fetch the data itself. Because it is
metadata, the index is small and cheap to refresh, and a scheduled run surfaces new Commission datasets as they
are published. Join dataset_id back to data.europa.eu to pull the full DCAT record when you need more.
Scaling and limits
Set Maximum records low to sample or high to pull the whole ~10,500-dataset index. The Actor pages the API
100 datasets at a time and delivers incrementally, so memory stays flat and you are billed only for what is
delivered. Because the index is small and republished continuously, a scheduled run keeps a downstream
catalogue current; each record's observed_at stamp dates the snapshot.
The distribution summary explained
Each dataset on data.europa.eu ships one or more distributions - the actual downloadable files or API
endpoints. This cell folds them into a compact summary rather than exploding the row: distribution_count is
the true number of distributions, distribution_formats lists the distinct format ids (CSV, XML, JSON, HTML,
PDF, ...), distribution_licences lists the distinct licence ids the distributions carry, and
distribution_download_urls gives up to 25 download URLs so you can fetch the data itself. A dataset qualifies
for this index because at least one of its distributions is CC BY 4.0; other distributions may carry additional
reuse terms, which is why distribution_licences can list more than one id while the record's headline
licence is CC BY 4.0. If you need the full distribution detail, the dataset_id and resource_uri join
straight back to the portal's DCAT record.
Why an index, not the data
The product here is deliberately the catalogue, not the payload. A broker or an analyst does not want ten thousand heterogeneous files dumped into one dataset; they want a clean, queryable map of what the Commission publishes under a reusable licence, so they can decide what to pull. This cell is that map - small, cheap to refresh, and keyed on the dataset id - and it hands you the download URLs when you are ready to fetch a specific dataset yourself.
Sibling Actors
It sits beside the fleet's EU regulatory-change (CELLAR) cell, which indexes EU legal acts rather than datasets. It shares its engineering - the runtime robots gate, offset paging with the licence filter baked in, push-then-charge billing and verbatim-value discipline - with the fleet's other public-data records Actors. Where the CELLAR cell indexes EU legal acts, this cell indexes EU datasets - a different corpus at data.europa.eu, so the two do not overlap.