# data.gouv.fr Dataset Catalogue Scraper (`compass_lab/data-gouv-datasets-scraper`) Actor

data.gouv.fr scraper and API: export French open-data dataset records (title, publisher, licence, update frequency, tags, file formats, description) to JSON, CSV or Excel. Official public API, no login, search by keyword, organisation or tag.

- **URL**: https://apify.com/compass\_lab/data-gouv-datasets-scraper.md
- **Developed by:** [COMPASSLAB](https://apify.com/compass_lab) (community)
- **Categories:** Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Get data.gouv.fr dataset records as clean JSON, CSV or Excel, from the Apify API or on a schedule. No login, about $1.00 for 1,000 datasets.**

### What does data.gouv.fr Dataset Catalogue Scraper do?

**data.gouv.fr Dataset Catalogue Scraper** extracts structured data from **[data.gouv.fr](https://data.gouv.fr)**. Collect metadata of open datasets published on data.gouv.fr (the French national open data portal) to build catalogues, monitor new publications and analyse open data by organization, tag or format. It works
as an **API for data.gouv.fr data**: run it from Apify Console, on a schedule, or from your own code, and get clean,
typed JSON with numbers as numbers and dates in ISO 8601.

### What you get

| | |
|---|---|
| **Data** | 13 fields per item: `id`, `title`, `slug`, `url`, `organization`, `license`, ... |
| **Formats** | JSON, CSV, Excel, HTML, or the Apify API |
| **Price** | $1.00 per 1,000 datasets, pay per result |
| **Access** | Public data.gouv.fr data only: no login, no cookies, robots.txt respected |

### Why use data.gouv.fr Dataset Catalogue Scraper?

- **Data journalists and researchers**: find every dataset on a topic, with its licence, publisher and files.
- **Govtech and civic tech**: watch a portal for new or updated datasets and their download links.
- **Data catalogue aggregation**: merge several open-data portals into one searchable index.
- **AI agents and RAG**: give an LLM the portal's metadata and file links as clean JSON.

Main features:

- Follows pagination up to `maxPages` pages per start URL and stops at `maxItems` results.
- Filters: `q` (search text), `organization` (organization), `tag` (tag), `pageSize` (page size), so you only get (and pay for) the results you need.
- Polite by default: respects robots.txt, at most `maxConcurrency` parallel requests and a delay between requests.
- Checks every result against field validators, so layout changes show up as clear data-quality warnings.
- Runs on the Apify platform: scheduling, API access, integrations, monitoring and datasets you can export.

### What data can data.gouv.fr Dataset Catalogue Scraper extract?

| Field | Type | Description |
|---|---|---|
| `id` | string | Unique dataset identifier |
| `title` | string | Dataset title |
| `slug` | string | Dataset slug |
| `url` | string | Public dataset page URL on data.gouv.fr |
| `organization` | string | Name of the publishing organization (null if the dataset is owned by a person) |
| `license` | string | License identifier (e.g. fr-lo for Licence Ouverte) |
| `createdAt` | ISO 8601 date | Creation date (ISO 8601) |
| `lastUpdate` | ISO 8601 date | Last update date (ISO 8601) |
| `frequency` | string | Declared update frequency (e.g. daily, monthly, unknown) |
| `tags` | array | Dataset tags. Empty when the dataset has no tags. |
| `resourcesCount` | integer | Number of resources (files) in the dataset |
| `formats` | array | Distinct formats of the dataset's resources (csv, json, xlsx...) |
| `description` | string | Plain-text description (markdown stripped, max 2000 chars) |

### How to scrape data.gouv.fr

1. Open data.gouv.fr Dataset Catalogue Scraper in Apify Console and go to the Input tab.
2. Enter what to scrape (see the Input section below), for example the start URLs.
3. Set **Max items** to the number of results you need.
4. Click **Start** and wait for the run to finish.
5. Download the results from the Output tab, or fetch them with the API.

### How much will it cost to scrape data.gouv.fr?

This Actor is priced **per result**: $1.00 per 1,000 results, with no extra charge for platform usage. That is about $1.00 for 1,000 datasets: 100 results cost $0.10 and 10,000 results cost $10.00. Set a maximum cost per run and the Actor stops when it is reached.

### Input

See the Input tab for full configuration options.

| Field | Type | Required | Description |
|---|---|---|---|
| `q` | string | no | Free-text search query applied to dataset titles and descriptions. |
| `organization` | string | no | Organization id or slug to restrict results to (optional). |
| `tag` | string | no | Only return datasets carrying this tag (optional). |
| `maxItems` | integer | no | Maximum number of items to return (0 = unlimited). |
| `startUrls` | array | no | Optional data.gouv.fr API dataset list URLs, e.g. https://www.data.gouv.fr/api/1/datasets/?page\_size=50. If empty, the actor builds the request from the search fields below. |
| `pageSize` | integer | no | Number of datasets requested per API page (1-100). |
| `maxPages` | integer | no | Maximum listing pages to follow per start URL (pagination). |
| `maxConcurrency` | integer | no | Maximum parallel requests (politeness; 1-10). |
| `requestDelayMs` | integer | no | Minimum delay between requests, in milliseconds (at least 250). |
| `proxyType` | string | no | none (direct connection), datacenter (Apify Proxy, cheapest) or residential (opt-in, billed per GB, fewer blocks). The actor never switches by itself. |
| `proxyCountry` | string | no | Two-letter country code for the proxy IP (optional). |

Example input:

```json
{
  "startUrls": [
    {
      "url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"
    }
  ],
  "maxItems": 75,
  "pageSize": 50,
  "maxPages": 3,
  "maxConcurrency": 2,
  "requestDelayMs": 1000,
  "proxyType": "none"
}
```

### Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Example results
from a real run:

```json
[
  {
    "id": "6abe2e358daf5a52fd678da0",
    "title": "Numéros de garde et de permanence des soins en France : relevé vérifié des numéros, coûts réels et sources officielles (2026)",
    "slug": "numeros-de-garde-et-de-permanence-des-soins-en-france-releve-verifie-des-numeros-couts-reels-et-sources-officielles-2026",
    "url": "https://www.data.gouv.fr/datasets/numeros-de-garde-et-de-permanence-des-soins-en-france-releve-verifie-des-numeros-couts-reels-et-sources-officielles-2026",
    "organization": "OPTICO",
    "license": "lov2",
    "createdAt": "2026-10-01T09:56:05.688000+00:00",
    "lastUpdate": "2026-10-01T09:56:06.238000+00:00",
    "frequency": "annual",
    "tags": [
      "116-117",
      "annuaire",
      "annuaire-sante",
      "docteur",
      "docteur-de-garde",
      "france",
      "medecin",
      "medecin-de-garde",
      "medecins",
      "numeros-de-garde",
      "numeros-utiles",
      "permanence-des-soins",
      "pharmacie",
      "pharmacie-de-garde",
      "pharmacien",
      "pharmacies",
      "sante",
      "sos-medecin",
      "sos-medecins",
      "urgences"
    ],
    "resourcesCount": 3,
    "formats": [
      "html",
      "csv"
    ],
    "description": "Relevé vérifié des numéros de garde et de permanence des soins en France : médecin de garde, pharmacie de garde, urgences et lignes nationales associées. Chaque ligne porte le numéro, le service couvert, le coût réel relevé (gratuit, prix d'un appel, ou tarification par minute pour les services audi…"
  },
  {
    "id": "6abe2be38304142b24f8389b",
    "title": "Numéros d'écoute en France : audit de 49 pages publiques et relevé des numéros obsolètes (2026)",
    "slug": "numeros-decoute-en-france-audit-de-49-pages-publiques-et-releve-des-numeros-obsoletes-2026",
    "url": "https://www.data.gouv.fr/datasets/numeros-decoute-en-france-audit-de-49-pages-publiques-et-releve-des-numeros-obsoletes-2026",
    "organization": "OPTICO",
    "license": "lov2",
    "createdAt": "2026-10-01T09:46:10.993000+00:00",
    "lastUpdate": "2026-10-01T09:46:11.601000+00:00",
    "frequency": "quarterly",
    "tags": [
      "coaching",
      "lignes-decoute",
      "numeros-decoute",
      "numeros-gratuits",
      "numeros-obsoletes",
      "parler",
      "parler-a-quelquun",
      "prevention-du-suicide",
      "sante-mentale",
      "soutien",
      "soutien-psychologique"
    ],
    "resourcesCount": 2,
    "formats": [
      "csv",
      "html"
    ],
    "description": "Relevé brut d'un audit de 49 pages publiques françaises listant des numéros d'écoute et de soutien (associations, institutions, médias), contrôlées une par une : 73 % contiennent au moins une erreur, et 36 pages sur 49 affichent au moins un numéro obsolète. Chaque ligne du relevé porte la page audit…"
  },
  {
    "id": "6abe2b1f510b3ba9b0454b96",
    "title": "Données essentielles des marchés publics - Région Auvergne Rhône-Alpes",
    "slug": "donnees-essentielles-des-marches-publics-region-auvergne-rhone-alpes-1985",
    "url": "https://www.data.gouv.fr/datasets/donnees-essentielles-des-marches-publics-region-auvergne-rhone-alpes-1985",
    "organization": "Région Auvergne-Rhône-Alpes",
    "license": "fr-lo",
    "createdAt": "2026-10-01T09:42:54.913000+00:00",
    "lastUpdate": "2026-10-01T09:42:59.926000+00:00",
    "frequency": null,
    "tags": [
      "commande-publique",
      "donnees-essentielles"
    ],
    "resourcesCount": 1,
    "formats": [
      "json"
    ],
    "description": "L'arrêté du 14 avril 2017 (https://www.legifrance.gouv.fr/eli/arrete/2017/4/14/ECFM1637256A/jo/texte), modifié par l'arrêté du 27 juillet 2018 (https://www.legifrance.gouv.fr/affichTexte.do?cidTexte=JORFTEXT000037282994&dateTexte=&categorieLien=id), impose à tous les acheteurs publics la publication…"
  }
]
```

### Integrations and API

- **Apify API**: start a run and get the results in one HTTP request:

```bash
curl -X POST "https://api.apify.com/v2/acts/compass_lab~data-gouv-datasets-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" -d '{"startUrls": [{"url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"}], "maxItems": 75, "pageSize": 50}'
```

- **Python** (`pip install apify-client`):

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("compass_lab/data-gouv-datasets-scraper").call(run_input={"startUrls": [{"url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"}], "maxItems": 75, "pageSize": 50})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

- **JavaScript** (`npm install apify-client`):

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('compass_lab/data-gouv-datasets-scraper').call({"startUrls": [{"url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"}], "maxItems": 75, "pageSize": 50});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

- **Make, Zapier, n8n, Google Sheets, webhooks**: use the Apify integrations (Integrations tab) to send each
  run's results where you need them, or to start a run from your workflow.
- **Schedules**: run it hourly, daily or weekly from Apify Console (Schedules) and always have fresh datasets.

### Tips and advanced options

- Keep **Max items** and **Max pages** as low as you need: fewer pages means a faster, cheaper run.
- Raise **Delay between requests** if the site responds slowly; keep **Max concurrency** low to stay polite.
- Missing values are `null`. Fields that often come back empty are listed in the run log as data-quality warnings.

### FAQ, disclaimers and support

#### Is it legal to scrape data.gouv.fr?

Vetted by the autonomy policy (green tier) on 2026-10-01: Green tier (autonomy policy, 2026-10-01): robots.txt exists and permits /api/1/datasets/ (no matching Disallow); terms https://www.data.gouv.fr/pages/legal/cgu read by the rule parser and a sandboxed model, no restriction found; public endpoint

Our Actors are ethical and do not extract any private user data, such as email addresses, gender, or location.
They only extract what the user has chosen to share publicly. We therefore believe that our Actors, when used
for ethical purposes by Apify users, are safe. However, you should be aware that your results could contain
personal data. Personal data is protected by the GDPR in the European Union and by other regulations around the
world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether
your reason is legitimate, consult your lawyers.

#### How many results can I get?

Up to `maxItems` per run (0 means no limit), as many as the source lists. Each result is one dataset item, and
you are only charged for items that are saved.

#### Can I run it on a schedule or from my own code?

Yes. Schedule it in Apify Console (Schedules), or call it with the `run-sync-get-dataset-items` endpoint or the
Python/JavaScript clients shown in **Integrations and API** above.

#### What are the limitations?

- Very large result sets (the whole catalogue holds tens of thousands of datasets) make long runs; use maxItems or filters to limit them.
- Some datasets have no organization, license or description, so those fields may be empty.
- The number of resources and formats reflects what the publisher declared and may be incomplete or inconsistently written (e.g. CSV vs csv).
- The API may rate-limit very fast crawling; the actor slows down and retries when this happens.
- Individual datasets can carry their own license terms; check the license field before reusing the underlying data.

#### Where can I get help?

Report problems or ideas on the Issues tab. To call this Actor from your own code, see the API tab.

### Related actors

the same clean, typed output across sources, so you can combine them in one dataset.

| Actor | What it scrapes | Price |
|---|---|---|
| [Django Weblog Scraper](https://apify.com/compass_lab/django-weblog-scraper) | Posts from Django Weblog | $1.00 / 1,000 |
| [Greenhouse Jobs Scraper](https://apify.com/compass_lab/greenhouse-jobs-scraper) | Job listings from Greenhouse | $1.60 / 1,000 |
| [Lever Jobs Scraper](https://apify.com/compass_lab/lever-jobs-scraper) | Job listings from Lever | $1.60 / 1,000 |
| [Python Jobs Scraper (python.org)](https://apify.com/compass_lab/python-job-board-scraper) | Job listings from Python Jobs Scraper (python.org) | $2.00 / 1,000 |
| [We Work Remotely Jobs Scraper](https://apify.com/compass_lab/weworkremotely-jobs-scraper) | Job listings from We Work Remotely | $2.50 / 1,000 |

# Actor input Schema

## `q` (type: `string`):

Free-text search query applied to dataset titles and descriptions.

## `organization` (type: `string`):

Organization id or slug to restrict results to (optional).

## `tag` (type: `string`):

Only return datasets carrying this tag (optional).

## `maxItems` (type: `integer`):

Maximum number of items to return (0 = unlimited). You pay per result, so this also caps the cost.

## `startUrls` (type: `array`):

Optional data.gouv.fr API dataset list URLs, e.g. https://www.data.gouv.fr/api/1/datasets/?page\_size=50. If empty, the actor builds the request from the search fields below.

## `pageSize` (type: `integer`):

Number of datasets requested per API page (1-100).

## `maxPages` (type: `integer`):

Maximum listing pages to follow per start URL (pagination).

## `maxConcurrency` (type: `integer`):

Maximum parallel requests (politeness; 1-10).

## `requestDelayMs` (type: `integer`):

Minimum delay between requests, in milliseconds (at least 250).

## `proxyType` (type: `string`):

none (direct connection), datacenter (Apify Proxy, cheapest) or residential (opt-in, billed per GB, fewer blocks). The actor never switches by itself.

## `proxyCountry` (type: `string`):

Two-letter country code for the proxy IP (optional).

## Actor input object example

```json
{
  "maxItems": 20,
  "startUrls": [
    {
      "url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"
    }
  ],
  "pageSize": 50,
  "maxPages": 3,
  "maxConcurrency": 2,
  "requestDelayMs": 1000,
  "proxyType": "none"
}
```

# Actor output Schema

## `dataset` (type: `string`):

All scraped items (overview view)

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 20,
    "startUrls": [
        {
            "url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"
        }
    ],
    "pageSize": 50,
    "maxPages": 3,
    "maxConcurrency": 2,
    "requestDelayMs": 1000,
    "proxyType": "none"
};

// Run the Actor and wait for it to finish
const run = await client.actor("compass_lab/data-gouv-datasets-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxItems": 20,
    "startUrls": [{ "url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50" }],
    "pageSize": 50,
    "maxPages": 3,
    "maxConcurrency": 2,
    "requestDelayMs": 1000,
    "proxyType": "none",
}

# Run the Actor and wait for it to finish
run = client.actor("compass_lab/data-gouv-datasets-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 20,
  "startUrls": [
    {
      "url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"
    }
  ],
  "pageSize": 50,
  "maxPages": 3,
  "maxConcurrency": 2,
  "requestDelayMs": 1000,
  "proxyType": "none"
}' |
apify call compass_lab/data-gouv-datasets-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,compass_lab/data-gouv-datasets-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qRMif6C65AGdczwGO/builds/fseLBybnzgu9MIo9H/openapi.json
