data.gouv.fr Dataset Catalogue Scraper avatar

data.gouv.fr Dataset Catalogue Scraper

Pricing

$1.00 / 1,000 results

Go to Apify Store
data.gouv.fr Dataset Catalogue Scraper

data.gouv.fr Dataset Catalogue Scraper

data.gouv.fr scraper and API: export French open-data dataset records (title, publisher, licence, update frequency, tags, file formats, description) to JSON, CSV or Excel. Official public API, no login, search by keyword, organisation or tag.

Pricing

$1.00 / 1,000 results

Rating

0.0

(0)

Developer

COMPASSLAB

COMPASSLAB

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 hours ago

Last modified

Categories

Share

Get data.gouv.fr dataset records as clean JSON, CSV or Excel, from the Apify API or on a schedule. No login, about $1.00 for 1,000 datasets.

What does data.gouv.fr Dataset Catalogue Scraper do?

data.gouv.fr Dataset Catalogue Scraper extracts structured data from data.gouv.fr. Collect metadata of open datasets published on data.gouv.fr (the French national open data portal) to build catalogues, monitor new publications and analyse open data by organization, tag or format. It works as an API for data.gouv.fr data: run it from Apify Console, on a schedule, or from your own code, and get clean, typed JSON with numbers as numbers and dates in ISO 8601.

What you get

Data13 fields per item: id, title, slug, url, organization, license, ...
FormatsJSON, CSV, Excel, HTML, or the Apify API
Price$1.00 per 1,000 datasets, pay per result
AccessPublic data.gouv.fr data only: no login, no cookies, robots.txt respected

Why use data.gouv.fr Dataset Catalogue Scraper?

  • Data journalists and researchers: find every dataset on a topic, with its licence, publisher and files.
  • Govtech and civic tech: watch a portal for new or updated datasets and their download links.
  • Data catalogue aggregation: merge several open-data portals into one searchable index.
  • AI agents and RAG: give an LLM the portal's metadata and file links as clean JSON.

Main features:

  • Follows pagination up to maxPages pages per start URL and stops at maxItems results.
  • Filters: q (search text), organization (organization), tag (tag), pageSize (page size), so you only get (and pay for) the results you need.
  • Polite by default: respects robots.txt, at most maxConcurrency parallel requests and a delay between requests.
  • Checks every result against field validators, so layout changes show up as clear data-quality warnings.
  • Runs on the Apify platform: scheduling, API access, integrations, monitoring and datasets you can export.

What data can data.gouv.fr Dataset Catalogue Scraper extract?

FieldTypeDescription
idstringUnique dataset identifier
titlestringDataset title
slugstringDataset slug
urlstringPublic dataset page URL on data.gouv.fr
organizationstringName of the publishing organization (null if the dataset is owned by a person)
licensestringLicense identifier (e.g. fr-lo for Licence Ouverte)
createdAtISO 8601 dateCreation date (ISO 8601)
lastUpdateISO 8601 dateLast update date (ISO 8601)
frequencystringDeclared update frequency (e.g. daily, monthly, unknown)
tagsarrayDataset tags. Empty when the dataset has no tags.
resourcesCountintegerNumber of resources (files) in the dataset
formatsarrayDistinct formats of the dataset's resources (csv, json, xlsx...)
descriptionstringPlain-text description (markdown stripped, max 2000 chars)

How to scrape data.gouv.fr

  1. Open data.gouv.fr Dataset Catalogue Scraper in Apify Console and go to the Input tab.
  2. Enter what to scrape (see the Input section below), for example the start URLs.
  3. Set Max items to the number of results you need.
  4. Click Start and wait for the run to finish.
  5. Download the results from the Output tab, or fetch them with the API.

How much will it cost to scrape data.gouv.fr?

This Actor is priced per result: $1.00 per 1,000 results, with no extra charge for platform usage. That is about $1.00 for 1,000 datasets: 100 results cost $0.10 and 10,000 results cost $10.00. Set a maximum cost per run and the Actor stops when it is reached.

Input

See the Input tab for full configuration options.

FieldTypeRequiredDescription
qstringnoFree-text search query applied to dataset titles and descriptions.
organizationstringnoOrganization id or slug to restrict results to (optional).
tagstringnoOnly return datasets carrying this tag (optional).
maxItemsintegernoMaximum number of items to return (0 = unlimited).
startUrlsarraynoOptional data.gouv.fr API dataset list URLs, e.g. https://www.data.gouv.fr/api/1/datasets/?page_size=50. If empty, the actor builds the request from the search fields below.
pageSizeintegernoNumber of datasets requested per API page (1-100).
maxPagesintegernoMaximum listing pages to follow per start URL (pagination).
maxConcurrencyintegernoMaximum parallel requests (politeness; 1-10).
requestDelayMsintegernoMinimum delay between requests, in milliseconds (at least 250).
proxyTypestringnonone (direct connection), datacenter (Apify Proxy, cheapest) or residential (opt-in, billed per GB, fewer blocks). The actor never switches by itself.
proxyCountrystringnoTwo-letter country code for the proxy IP (optional).

Example input:

{
"startUrls": [
{
"url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"
}
],
"maxItems": 75,
"pageSize": 50,
"maxPages": 3,
"maxConcurrency": 2,
"requestDelayMs": 1000,
"proxyType": "none"
}

Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Example results from a real run:

[
{
"id": "6abe2e358daf5a52fd678da0",
"title": "Numéros de garde et de permanence des soins en France : relevé vérifié des numéros, coûts réels et sources officielles (2026)",
"slug": "numeros-de-garde-et-de-permanence-des-soins-en-france-releve-verifie-des-numeros-couts-reels-et-sources-officielles-2026",
"url": "https://www.data.gouv.fr/datasets/numeros-de-garde-et-de-permanence-des-soins-en-france-releve-verifie-des-numeros-couts-reels-et-sources-officielles-2026",
"organization": "OPTICO",
"license": "lov2",
"createdAt": "2026-10-01T09:56:05.688000+00:00",
"lastUpdate": "2026-10-01T09:56:06.238000+00:00",
"frequency": "annual",
"tags": [
"116-117",
"annuaire",
"annuaire-sante",
"docteur",
"docteur-de-garde",
"france",
"medecin",
"medecin-de-garde",
"medecins",
"numeros-de-garde",
"numeros-utiles",
"permanence-des-soins",
"pharmacie",
"pharmacie-de-garde",
"pharmacien",
"pharmacies",
"sante",
"sos-medecin",
"sos-medecins",
"urgences"
],
"resourcesCount": 3,
"formats": [
"html",
"csv"
],
"description": "Relevé vérifié des numéros de garde et de permanence des soins en France : médecin de garde, pharmacie de garde, urgences et lignes nationales associées. Chaque ligne porte le numéro, le service couvert, le coût réel relevé (gratuit, prix d'un appel, ou tarification par minute pour les services audi…"
},
{
"id": "6abe2be38304142b24f8389b",
"title": "Numéros d'écoute en France : audit de 49 pages publiques et relevé des numéros obsolètes (2026)",
"slug": "numeros-decoute-en-france-audit-de-49-pages-publiques-et-releve-des-numeros-obsoletes-2026",
"url": "https://www.data.gouv.fr/datasets/numeros-decoute-en-france-audit-de-49-pages-publiques-et-releve-des-numeros-obsoletes-2026",
"organization": "OPTICO",
"license": "lov2",
"createdAt": "2026-10-01T09:46:10.993000+00:00",
"lastUpdate": "2026-10-01T09:46:11.601000+00:00",
"frequency": "quarterly",
"tags": [
"coaching",
"lignes-decoute",
"numeros-decoute",
"numeros-gratuits",
"numeros-obsoletes",
"parler",
"parler-a-quelquun",
"prevention-du-suicide",
"sante-mentale",
"soutien",
"soutien-psychologique"
],
"resourcesCount": 2,
"formats": [
"csv",
"html"
],
"description": "Relevé brut d'un audit de 49 pages publiques françaises listant des numéros d'écoute et de soutien (associations, institutions, médias), contrôlées une par une : 73 % contiennent au moins une erreur, et 36 pages sur 49 affichent au moins un numéro obsolète. Chaque ligne du relevé porte la page audit…"
},
{
"id": "6abe2b1f510b3ba9b0454b96",
"title": "Données essentielles des marchés publics - Région Auvergne Rhône-Alpes",
"slug": "donnees-essentielles-des-marches-publics-region-auvergne-rhone-alpes-1985",
"url": "https://www.data.gouv.fr/datasets/donnees-essentielles-des-marches-publics-region-auvergne-rhone-alpes-1985",
"organization": "Région Auvergne-Rhône-Alpes",
"license": "fr-lo",
"createdAt": "2026-10-01T09:42:54.913000+00:00",
"lastUpdate": "2026-10-01T09:42:59.926000+00:00",
"frequency": null,
"tags": [
"commande-publique",
"donnees-essentielles"
],
"resourcesCount": 1,
"formats": [
"json"
],
"description": "L'arrêté du 14 avril 2017 (https://www.legifrance.gouv.fr/eli/arrete/2017/4/14/ECFM1637256A/jo/texte), modifié par l'arrêté du 27 juillet 2018 (https://www.legifrance.gouv.fr/affichTexte.do?cidTexte=JORFTEXT000037282994&dateTexte=&categorieLien=id), impose à tous les acheteurs publics la publication…"
}
]

Integrations and API

  • Apify API: start a run and get the results in one HTTP request:
curl -X POST "https://api.apify.com/v2/acts/compass_lab~data-gouv-datasets-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
-H "Content-Type: application/json" -d '{"startUrls": [{"url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"}], "maxItems": 75, "pageSize": 50}'
  • Python (pip install apify-client):
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("compass_lab/data-gouv-datasets-scraper").call(run_input={"startUrls": [{"url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"}], "maxItems": 75, "pageSize": 50})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)
  • JavaScript (npm install apify-client):
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('compass_lab/data-gouv-datasets-scraper').call({"startUrls": [{"url": "https://www.data.gouv.fr/api/1/datasets/?page_size=50"}], "maxItems": 75, "pageSize": 50});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
  • Make, Zapier, n8n, Google Sheets, webhooks: use the Apify integrations (Integrations tab) to send each run's results where you need them, or to start a run from your workflow.
  • Schedules: run it hourly, daily or weekly from Apify Console (Schedules) and always have fresh datasets.

Tips and advanced options

  • Keep Max items and Max pages as low as you need: fewer pages means a faster, cheaper run.
  • Raise Delay between requests if the site responds slowly; keep Max concurrency low to stay polite.
  • Missing values are null. Fields that often come back empty are listed in the run log as data-quality warnings.

FAQ, disclaimers and support

Vetted by the autonomy policy (green tier) on 2026-10-01: Green tier (autonomy policy, 2026-10-01): robots.txt exists and permits /api/1/datasets/ (no matching Disallow); terms https://www.data.gouv.fr/pages/legal/cgu read by the rule parser and a sandboxed model, no restriction found; public endpoint

Our Actors are ethical and do not extract any private user data, such as email addresses, gender, or location. They only extract what the user has chosen to share publicly. We therefore believe that our Actors, when used for ethical purposes by Apify users, are safe. However, you should be aware that your results could contain personal data. Personal data is protected by the GDPR in the European Union and by other regulations around the world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether your reason is legitimate, consult your lawyers.

How many results can I get?

Up to maxItems per run (0 means no limit), as many as the source lists. Each result is one dataset item, and you are only charged for items that are saved.

Can I run it on a schedule or from my own code?

Yes. Schedule it in Apify Console (Schedules), or call it with the run-sync-get-dataset-items endpoint or the Python/JavaScript clients shown in Integrations and API above.

What are the limitations?

  • Very large result sets (the whole catalogue holds tens of thousands of datasets) make long runs; use maxItems or filters to limit them.
  • Some datasets have no organization, license or description, so those fields may be empty.
  • The number of resources and formats reflects what the publisher declared and may be incomplete or inconsistently written (e.g. CSV vs csv).
  • The API may rate-limit very fast crawling; the actor slows down and retries when this happens.
  • Individual datasets can carry their own license terms; check the license field before reusing the underlying data.

Where can I get help?

Report problems or ideas on the Issues tab. To call this Actor from your own code, see the API tab.

the same clean, typed output across sources, so you can combine them in one dataset.

ActorWhat it scrapesPrice
Django Weblog ScraperPosts from Django Weblog$1.00 / 1,000
Greenhouse Jobs ScraperJob listings from Greenhouse$1.60 / 1,000
Lever Jobs ScraperJob listings from Lever$1.60 / 1,000
Python Jobs Scraper (python.org)Job listings from Python Jobs Scraper (python.org)$2.00 / 1,000
We Work Remotely Jobs ScraperJob listings from We Work Remotely$2.50 / 1,000