# EU Commission Open Data Index - CC BY 4.0 (`nexgensignal/eu-commission-open-dataset-index-records`) Actor

Catalogue of European Commission open datasets (CC BY 4.0) on data.europa.eu: search by text, EU theme (energy, health, transport...), publisher (Eurostat, JRC...) or download format, most recently modified first - title, dates, publisher, formats, download URLs. Pay per record.

- **URL**: https://apify.com/nexgensignal/eu-commission-open-dataset-index-records.md
- **Developed by:** [NexGen Signal](https://apify.com/nexgensignal) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $33.50 / 1,000 eu dataset index records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## EU Commission Open Data Index - CC BY 4.0

An index of the **European Commission's open datasets** on data.europa.eu that carry a **CC BY 4.0**
distribution - **one record per dataset**. This is a catalogue of metadata: each record names a dataset, its
theme, dates, publishing Directorate-General and its distributions (formats, licences, download URLs). The
records point at the datasets; **the product is the index itself.**

### What one record represents

The source is **data.europa.eu** (the EU Open Data Portal) hub search API. This cell is scoped to datasets
whose publisher is a European Commission **Directorate-General** (`corporate-body-classification=DIR_GEN`) and
which carry a **CC BY 4.0** distribution. Each record is **one dataset's metadata**: its id, title, DCAT
data-theme categories, modified and issued dates, the publishing Directorate-General, the source catalogue,
the landing page and resource URI, and a summary of its distributions - how many, their distinct formats,
their distinct licence ids, and the download URLs.

### Sample output

![Sample output — EU Commission Open Data Index - CC BY 4.0](https://api.apify.com/v2/key-value-stores/IXCaMKjxSmUTLHhmq/records/eu-commission-open-dataset-index-records.png)

*Real rows from a live run of this actor (first 5 rows, selected columns).*

One full record from the same run, exactly as delivered:

```json
{
  "dataset_id": "eu-food-additives",
  "title": "EU Food Additives",
  "categories": "TECH;AGRI",
  "modified": "2024-08-13",
  "issued": "2015-08-18",
  "publisher_name": "Directorate-General for Health and Food Safety",
  "catalog": "sante",
  "landing_page": "https://food.ec.europa.eu/safety/food-improvement-agents/additives_en",
  "resource_uri": "http://data.europa.eu/88u/dataset/eu-food-additives",
  "distribution_count": 3,
  "distribution_formats": "JSON",
  "distribution_licences": "CC_BY_4_0",
  "distribution_download_urls": "https://developer.datalake.sante.service.ec.europa.eu/api-details#api=228d6fda-9092-4c25-af9a-d537666ed0e5&operation=66382a6d-6a0d-4c17-99a5-6156f6436ded https://api.datalake.sante.service.ec.europa.eu/food-additives/download?format=json&api-version=v2.0",
  "licence": "CC BY 4.0",
  "record_id": "eu-food-additives",
  "source": "data.europa.eu (EU Open Data Portal)",
  "source_dataset": "hub/search filter=dataset DIR_GEN CC_BY_4_0",
  "attribution": "European Commission - data.europa.eu (EU Open Data Portal)",
  "caveat": "One record per European Commission (Directorate-General) dataset on data.europa.eu that carries a CC BY 4.0 distribution. This is a catalogue of METADATA: dataset id, title, DCAT category, modified/issued dates, publishing Directorate-General, and distribution formats/licences/download URLs - the records point at the datasets, they are not the underlying data. Contact-point and maintainer person fields are structurally excluded; the publisher name is an organisation.",
  "observed_at": "2026-09-25T17:32:09Z",
  "licence_notice": "Dataset metadata from data.europa.eu (the EU Open Data Portal). This cell delivers ONLY datasets carrying a Creative Commons Attribution 4.0 (CC BY 4.0) distribution; the licence filter is baked into the query. Attribution required. The records are METADATA that point at the datasets; the product is the index itself."
}
```

### Search the catalogue

**Question:** Which European Commission energy datasets about electricity can be downloaded as CSV, most recently updated first?

```json
{"searchText": "electricity", "categories": ["ENER"], "formats": ["CSV"], "newestFirst": true, "maxRecords": 10}
```

Run on 30 Sep 2026 the portal matched **10** datasets and the run returned all 10 (10 × $0.05 = $0.50 on Apify's Free plan):

| Title | Modified | Publisher | Formats |
|---|---|---|---|
| Final energy consumption in households per capita | 2026-07-21 | Eurostat | HTML;XML;CSV;TSV |
| Energy taxes | 2026-07-21 | Eurostat | CSV;HTML;XML;TSV |
| Primary energy consumption | 2026-07-21 | Eurostat | XML;HTML;TSV;CSV |
| Market share of the largest generator in the electricity market | 2026-06-02 | Eurostat | XML;HTML;TSV;CSV |
| Final energy consumption in households by type of fuel | 2026-06-02 | Eurostat | XML;TSV;CSV;HTML |
| Waste electrical and electronic equipment (WEEE) by waste management operations  | 2026-04-08 | Eurostat | XML;CSV;TSV;HTML |

Search options (all optional; a record must match every option you set, and several values in one option match any of them):

- **Search text** (`searchText`) — the portal's own full-text search over titles and descriptions.
- **Themes** (`categories`) — EU data-theme codes: `AGRI`, `ECON`, `EDUC`, `ENER`, `ENVI`, `GOVE`, `HEAL`, `INTR`, `JUST`, `REGI`, `SOCI`, `TECH`, `TRAN`. Several = datasets tagged with **all** of them (the portal combines values this way).
- **Publisher** (`publisher`) — one Commission catalogue: `estat` Eurostat, `jrc` Joint Research Centre, `commu` DG Communication (Eurobarometer), `cnect`, `empl`, `grow`, `comp`, `sante`, `digit`, `just`, `ener`, `rtd`.
- **Download formats** (`formats`) — e.g. `CSV`, `JSON`, `XLSX`; several = datasets offering all of them.
- **Most recently modified first** (`newestFirst`) — the portal's sort by modification date.

Every option goes to the data.europa.eu search API, so filtering happens before **Maximum records**; the status message gives the portal's match count. "No datasets match" (nothing charged) is a successful run; if the portal cannot be reached the run fails as a source failure. With **no** search option the request is exactly the one the actor always sent. The licence (CC BY 4.0) and Commission-publisher filters always apply.

### Coverage and volume

The live filtered set is **10,536 datasets** (measured at build time: `corporate-body-classification=DIR_GEN`
AND `license=CC_BY_4_0`). The European Commission is organised into Directorates-General, so this classification
is precisely "datasets published by a Commission DG."

**Sol's Wave-3 index put this door at 10,533 CC BY 4.0 Commission records; measured live at build time the
filter returns 10,536 - the live figure is what this listing quotes.**

The Actor pages the hub search API 100 datasets at a time and stops as soon as your **Maximum records** cap is
met.

### Licence and attribution

This cell delivers **only** datasets that carry a **CC BY 4.0** distribution. The licence filter
(`license=CC_BY_4_0`) is **baked into the query**, not a buyer option, and a per-record assertion **rejects
any dataset whose distributions do not include CC BY 4.0** (verified with a planted-row test). The full notice
travels on every record:

> Dataset metadata from data.europa.eu (the EU Open Data Portal). This cell delivers ONLY datasets carrying a Creative Commons Attribution 4.0 (CC BY 4.0) distribution; the licence filter is baked into the query. Attribution required. The records are METADATA that point at the datasets; the product is the index itself.

The `licence` field is set to `CC BY 4.0` on every record, and `distribution_licences` lists the distinct
licence ids the dataset's distributions actually carry (a dataset may offer additional distributions under
other reuse terms; the CC BY 4.0 one is what qualifies it here). Attribution is to the European Commission and
the publishing Directorate-General, named on each record.

### Person-data policy

data.europa.eu dataset records carry a `contact_point` (and sometimes a maintainer) that can hold a person's
name, email or phone. This Actor **structurally excludes** the contact point and maintainer: they are never
read. The `publisher_name` that is kept is the Directorate-General - an organisation, never a person - and a
per-record assertion rejects any contact/maintainer/person field (verified with a planted-field test). No
natural-person data is processed.

### Interpretation caveat

One record per European Commission (Directorate-General) dataset on data.europa.eu that carries a CC BY 4.0 distribution. A catalogue of METADATA: dataset id, title, DCAT category, modified/issued dates, publishing Directorate-General, and distribution formats/licences/download URLs - the records point at the datasets, they are not the underlying data. Contact-point and maintainer person fields are structurally excluded.

Values are reproduced verbatim from the API; the Actor never rewrites a field. Titles are multilingual in the
source; the record carries the English title where available, otherwise the first available language. The
download URLs are capped at 25 per record to keep rows bounded; `distribution_count` reports the true total.

### Data quality and freshness

`distribution_count` is a real number; dates are the portal's own `modified`/`issued` stamps. The `record_id`
is the dataset id, so the index is safe to diff, deduplicate or upsert. Every run re-reads the live API, so the
index is as fresh as data.europa.eu publishes, and each record's `observed_at` stamp dates the snapshot. The
run's `RUN_RECEIPT` records the live match count alongside how many records were delivered and charged.

### Provenance and compliance

Every run reads `data.europa.eu/robots.txt` at runtime; the gate result (URL, status, byte length, SHA-256 of
the policy) is written to the run's `RUN_RECEIPT`, and the hub search path is confirmed crawlable before any
data request. The API is keyless. The Actor never bypasses a block or fetches through a mirror.

### Inputs

**Default changed on 30 Sep 2026:** if you leave `maxRecords` out, a run now returns up to **10** records (it was 500). Set `maxRecords` yourself to get more — the maximum is unchanged.

- **Maximum records** (`maxRecords`) - hard cap on dataset-index records delivered and billed.
- **Search options** - see *Search the catalogue* above; all optional.

### Output

Records land in the Actor's default dataset and export as JSON, CSV, Excel or via the Apify API. A tabular
**overview view** surfaces dataset id, title, categories, modified date, publishing Directorate-General,
distribution formats and licence.

### Fields in detail

The record leads with `dataset_id` and `title`, then `categories` (DCAT data-theme ids), `modified`, `issued`,
`publisher_name` (the Directorate-General), `catalog`, `landing_page`, `resource_uri`, and the distribution
summary - `distribution_count`, `distribution_formats`, `distribution_licences`, `distribution_download_urls` -
plus the `licence` label and full `licence_notice`. The provenance block closes every record.

### Typical uses

Data brokers, open-data teams and procurement analysts use this cell to source reusable Commission datasets:
one queryable index of what the EU's Directorates-General publish under CC BY 4.0, with the theme to filter by
topic, the modified date to spot fresh releases, and the download URLs to fetch the data itself. Because it is
metadata, the index is small and cheap to refresh, and a scheduled run surfaces new Commission datasets as they
are published. Join `dataset_id` back to data.europa.eu to pull the full DCAT record when you need more.

### Scaling and limits

Set **Maximum records** low to sample or high to pull the whole ~10,500-dataset index. The Actor pages the API
100 datasets at a time and delivers incrementally, so memory stays flat and you are billed only for what is
delivered. Because the index is small and republished continuously, a scheduled run keeps a downstream
catalogue current; each record's `observed_at` stamp dates the snapshot.

### The distribution summary explained

Each dataset on data.europa.eu ships one or more distributions - the actual downloadable files or API
endpoints. This cell folds them into a compact summary rather than exploding the row: `distribution_count` is
the true number of distributions, `distribution_formats` lists the distinct format ids (CSV, XML, JSON, HTML,
PDF, ...), `distribution_licences` lists the distinct licence ids the distributions carry, and
`distribution_download_urls` gives up to 25 download URLs so you can fetch the data itself. A dataset qualifies
for this index because at least one of its distributions is CC BY 4.0; other distributions may carry additional
reuse terms, which is why `distribution_licences` can list more than one id while the record's headline
`licence` is `CC BY 4.0`. If you need the full distribution detail, the `dataset_id` and `resource_uri` join
straight back to the portal's DCAT record.

### Why an index, not the data

The product here is deliberately the catalogue, not the payload. A broker or an analyst does not want ten
thousand heterogeneous files dumped into one dataset; they want a clean, queryable map of what the Commission
publishes under a reusable licence, so they can decide what to pull. This cell is that map - small, cheap to
refresh, and keyed on the dataset id - and it hands you the download URLs when you are ready to fetch a specific
dataset yourself.

### Sibling Actors

It sits beside the fleet's EU regulatory-change (CELLAR) cell, which indexes EU legal acts rather than datasets. It shares its engineering - the runtime robots gate, offset paging with the licence
filter baked in, push-then-charge billing and verbatim-value discipline - with the fleet's other public-data
records Actors. Where the CELLAR cell indexes EU **legal acts**, this cell indexes EU **datasets** - a
different corpus at data.europa.eu, so the two do not overlap.

# Actor input Schema

## `searchText` (type: `string`):

Words searched in titles and descriptions by the portal, e.g. <code>electricity</code>, <code>unemployment</code>, <code>Eurobarometer</code>.

## `categories` (type: `array`):

EU data themes: <code>AGRI</code> Agriculture, fisheries, forestry and food, <code>ECON</code> Economy and finance, <code>EDUC</code> Education, culture and sport, <code>ENER</code> Energy, <code>ENVI</code> Environment, <code>GOVE</code> Government and public sector, <code>HEAL</code> Health, <code>INTR</code> International issues, <code>JUST</code> Justice, legal system and public safety, <code>REGI</code> Regions and cities, <code>SOCI</code> Population and society, <code>TECH</code> Science and technology, <code>TRAN</code> Transport. Several = datasets tagged with all of them.

## `publisher` (type: `string`):

One Commission publisher (the portal catalogue).

## `formats` (type: `array`):

e.g. <code>CSV</code>, <code>JSON</code>, <code>XLSX</code>, <code>TSV</code>, <code>XML</code>. Several = datasets offering all of them.

## `newestFirst` (type: `boolean`):

Order by the dataset's modification date, newest first.

## `maxRecords` (type: `integer`):

Maximum dataset-index records delivered and billed. You are billed only for records delivered. About 10,536 Commission (Directorate-General) datasets carry a CC BY 4.0 distribution.

## Actor input object example

```json
{
  "searchText": "electricity",
  "categories": [
    "ENER"
  ],
  "formats": [
    "CSV"
  ],
  "newestFirst": true,
  "maxRecords": 10
}
```

# Actor output Schema

## `results` (type: `string`):

The delivered EU Commission open dataset index record.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchText": "electricity",
    "categories": [
        "ENER"
    ],
    "formats": [
        "CSV"
    ],
    "newestFirst": true,
    "maxRecords": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("nexgensignal/eu-commission-open-dataset-index-records").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchText": "electricity",
    "categories": ["ENER"],
    "formats": ["CSV"],
    "newestFirst": True,
    "maxRecords": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("nexgensignal/eu-commission-open-dataset-index-records").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchText": "electricity",
  "categories": [
    "ENER"
  ],
  "formats": [
    "CSV"
  ],
  "newestFirst": true,
  "maxRecords": 10
}' |
apify call nexgensignal/eu-commission-open-dataset-index-records --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nexgensignal/eu-commission-open-dataset-index-records"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4hbkHKHcz53FSOaoL/builds/WbgJqO7wfaVHdMRRh/openapi.json
