Netherlands Open Data Scraper - data.overheid.nl API avatar

Netherlands Open Data Scraper - data.overheid.nl API

Pricing

from $6.23 / 1,000 datasets

Go to Apify Store
Netherlands Open Data Scraper - data.overheid.nl API

Netherlands Open Data Scraper - data.overheid.nl API

Scrape all 20,380 datasets from data.overheid.nl, the national open data portal of the Netherlands: full DCAT-AP-DONL metadata, distributions, 4,667 data requests, registers and live link checks.

Pricing

from $6.23 / 1,000 datasets

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 days ago

Last modified

Categories

Share

ParseForge

Netherlands Open Data Scraper - data.overheid.nl CKAN API

Scrape every one of the 20,380 datasets on data.overheid.nl, the national open data portal of the Netherlands, with all 86 metadata fields the CKAN API carries. Each row is a flat DCAT-AP-DONL record: the data owner as an OWMS authority, the EU themes in Dutch and English, licence, access rights, update frequency, contact point, legal foundation and every file with its format. No login, no API key, no rate limit. Export to CSV, JSON, Excel, or XML.

The portal's own search shows you ten results a page and hides most of the record behind them, and its government open data API answers only in Dutch URIs: a licence comes back as http://creativecommons.org/licenses/by/4.0/deed.nl and a theme as .../owms/terms/Natuur_en_milieu. This Actor pages the whole open data catalogue in one run, resolves every URI to a readable name, and adds four registers the CKAN API does not expose at all, including the 4,667 public data requests Dutch citizens have filed asking the government to open a dataset.

Who uses itWhat they scrape data.overheid.nl for
Data engineersA complete inventory of Netherlands government datasets to pipe into a warehouse
Open data researchersLicence, access-rights and update-frequency coverage across 187 public bodies
Civic technologistsLive download links by format, checked, before building on top of them
JournalistsThe 4,667 data requests and the government's own answer to each one
Compliance and policy teamsThe 229 EU High Value Datasets and the 34 statutory basisregistraties

What it does

This Actor reads the CKAN open data API behind data.overheid.nl and the five community registers on the site itself, and returns each entry as a flat row tagged with rowType:

  • 🇳🇱 Datasets: all 20,380, with 86 fields each: title, description, national identifier, data owner and publisher, harvest organization, source catalogue, OWMS themes in Dutch and English, keywords, licence with an open-licence flag, access rights, availability status, update frequency, contact point with email, phone and website, issue and modification dates, temporal coverage, spatial extent and the legal foundation with its wetten.overheid.nl link.
  • 📦 Files: around 74,600 distributions, 3.66 per dataset, with access and download URLs, the EU file-type URI resolved to a plain name, a machine-readable flag, media type, preview URL and licence.
  • 📣 Data requests: all 4,667 public requests to open a dataset, with the phase they reached in Dutch and English, the body that was asked, and optionally the full request and the government's written answer.
  • 🏛️ Registers: 1,671 public bodies with their governance layer (Rijk, Gemeente, Provincie, Waterschap), 91 curated groups, 36 source catalogues and the registered applications built on the data.
  • 📚 Directories: 187 data owners, 99 themes, 40 formats, 8 licences, 23 source catalogues and 10,403 keywords, each with the dataset count and a working link to its slice of the portal.
  • Flags the portal buries: 115 datasets flagged high value and 229 carrying an EU High Value Dataset category, 34 basisregistraties, 6,909 nationally-scoped datasets, 51 reference-data sets, and the 2,020 datasets catalogued but not openly downloadable.
  • 🔎 Optional checks: fetch each file to see whether it still downloads (9.0 percent are not), and read its first bytes to see whether it really is what it claims (10.0 percent are not: 24 files declared JSON in a 469 file sample served an HTML page instead).

Results export to CSV, JSON, Excel, or XML, or stream from the API.

What you can do with data.overheid.nl data

🗂️ Build a complete inventory of Netherlands government data. One run with no filters returns all 20,380 datasets with their owners, licences, formats and update frequencies. Tick the file rows and you also get every download URL in the country's open data catalogue, ready to schedule.

🔗 Find the dead links before your users do. data.overheid.nl harvests 23 separate catalogues and records a link status on only 8 percent of files, so for the other 92 percent nobody has ever checked. Turn on the link check and each file comes back with its real HTTP status, redirect target, content type and size. In a 647 file sample, 9.0 percent were dead: mostly WFS and WMS services answering 500, plus spreadsheets behind a 403.

⚖️ Audit open data policy compliance. Filter to the 229 datasets in an EU High Value Dataset category, or to the 2,020 datasets whose access rights are not public, or to the 1,376 published under a closed licence, and you have the evidence for a coverage report in one export.

📣 Read what the public is asking for. The data request register is 4,667 requests from citizens, journalists and companies asking a named public body to open a dataset, each with the phase it reached and the answer it got. 336 were refused as not public and 1,970 ended with the government pointing at data that already existed.

Why choose this scraper

What you get
Full coverageAll 20,380 datasets, 86 fields each, plus five registers the CKAN API does not expose
Readable, not URIsEvery OWMS and EU URI resolved to a name, themes in Dutch and English
Filters measured against the live totals28 filters, each one checked against the catalogue count, with the real number in the dropdown
Clean rowsISO dates, real numbers, Yes/No booleans, "Not Disclosed" where the portal withheld a value, zero always-empty columns
A register nobody else carries4,667 public data requests with the government's own answer
Priced per rowYou pay for rows written, and the optional checks only when they returned something

How it compares

One other Actor covers this exact source, and two more cover CKAN portals generically. The generic ones are cheaper per row and give you what a bare CKAN call gives you: an id, a name, a few tags. The named data.overheid.nl Actor lists six fields in its own store description. This Actor costs more per row than either, and the difference buys the other 80 fields, the URI-to-name resolution the portal makes you do yourself, the five community registers, and filters that were each measured against the live catalogue instead of copied from CKAN documentation.

This Actorbenthepythondev/netherlands-data-overheid-scraperGeneric CKAN ActorsThe portal's own API
Fields per dataset866, per its own descriptionWhatever CKAN returns rawRaw URIs only
Filters28, measured against totalsNone documentedFree textYou write the Solr query
Data requests4,667NoNoNot exposed
Registers5NoNoOnly organizations
Link checkingYes, optionalNoNoNo
Users in the last 30 daysNew11 eachn/a

What a Netherlands open dataset looks like

One real row from the verified run, unedited apart from truncating the description:

{
"rowType": "dataset",
"url": "https://data.overheid.nl/dataset/wpozittenblijvers-v1",
"apiUrl": "https://data.overheid.nl/data/api/3/action/package_show?id=wpozittenblijvers-v1",
"id": "7cc95d10-bb52-4211-bbe8-8a18bb6e0f0d",
"slug": "wpozittenblijvers-v1",
"title": "Zittenblijvers in het bo en sbo per vestiging",
"description": "Dit document beschrijft de bestanden over leerlingen in het Primair Onderwijs die via een API (A ...",
"identifier": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1",
"alternateIdentifiers": [],
"organization": "dienst-uitvoering-onderwijs",
"organizationTitle": "DUO",
"organizationId": "3b966f6a-4bd6-4f63-95b4-76adcdf31657",
"dataOwner": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs",
"dataOwnerName": "Dienst Uitvoering Onderwijs",
"publisher": "http://standaarden.overheid.nl/owms/terms/Dienst_Uitvoering_Onderwijs",
"publisherName": "Dienst Uitvoering Onderwijs",
"sourceCatalog": "https://onderwijsdata.duo.nl",
"sourceCatalogName": "Dienst Uitvoering Onderwijs (DUO)",
"themes": [
"http://standaarden.overheid.nl/owms/terms/Onderwijs_en_wetenschap"
],
"themeLabels": [
"Onderwijs en wetenschap"
],
"themeLabelsEn": [
"Education and science"
],
"keywords": [
"Leerlingen"
],
"keywordCount": 1,
"licence": "http://creativecommons.org/licenses/by/4.0/deed.nl",
"licenceLabel": "CC-BY (4.0)",
"licenceUrl": "http://creativecommons.org/licenses/by/4.0/deed.nl",
"isOpenLicence": "Yes",
"accessRights": "http://publications.europa.eu/resource/authority/access-right/PUBLIC",
"accessRightsLabel": "Public",
"datasetStatus": "http://data.overheid.nl/status/beschikbaar",
"datasetStatusLabel": "Available",
"restrictionsStatement": "Not Disclosed",
"updateFrequency": "http://publications.europa.eu/resource/authority/frequency/ANNUAL",
"updateFrequencyLabel": "Annual",
"languages": [
"http://publications.europa.eu/resource/authority/language/NLD"
],
"languageLabels": [
"Dutch"
],
"metadataLanguage": "http://publications.europa.eu/resource/authority/language/NLD",
"highValueDataset": "No",
"hvdCategories": [],
"hvdCategoryLabels": [],
"basisRegister": "No",
"nationalCoverage": "No",
"referenceData": "No",
"datasetQuality": "N/A",
"contactName": "Informatieproducten",
"contactTitle": "Not Disclosed",
"contactEmail": "informatieproducten@duo.nl",
"contactPhone": "Not Disclosed",
"contactWebsite": "Not Disclosed",
"contactAddress": "Not Disclosed",
"author": "Informatieproducten",
"authorEmail": "informatieproducten@duo.nl",
"maintainer": "Not Disclosed",
"maintainerEmail": "Not Disclosed",
"version": "1.0.0",
"versionNotes": "Not Disclosed",
"metadataCreated": "2020-04-02T22:20:24.932693",
"metadataModified": "2026-08-27T08:10:20.178083",
"modified": "2026-01-09T07:43:36",
"issued": "Not Disclosed",
"datePlanned": "Not Disclosed",
"temporalStart": "Not Disclosed",
"temporalEnd": "Not Disclosed",
"temporalLabel": "Not Disclosed",
"spatialValues": [],
"spatialSchemes": [],
"legalFoundationLabel": "Not Disclosed",
"legalFoundationRef": "Not Disclosed",
"legalFoundationUrl": "Not Disclosed",
"provenance": [],
"documentation": [],
"samples": [],
"sources": [],
"conformsTo": [],
"relatedResources": [],
"landingPage": "https://onderwijsdata.duo.nl/datasets/wpozittenblijvers-v1",
"syncChecksum": "Not Disclosed",
"extraFields": [],
"changeType": "updated",
"state": "active",
"isPrivate": "No",
"fileCount": 1,
"fileFormats": [
"http://publications.europa.eu/resource/authority/file-type/CSV"
],
"fileFormatLabels": [
"CSV"
],
"files": [
{
"name": "Aantal zittenblijvers bo en sbo per schoolvestiging",
"url": "https://onderwijsdata.duo.nl/dataset/f0c79c94-6ffb-44bf-a3d5-9a05aef4ade4/resource/df7297a9-70a3-4a3a-b4a6-d289c7d16014/download/brin6_zittenblijvers.csv",
"format": "CSV",
"position": 0
}
],
"scrapedAt": "2026-08-27T19:34:28.746Z"
}

Configure the run

Leave everything empty and the Actor sweeps the whole open data catalogue, newest change first. Every dropdown carries the live dataset count next to each value, so you can see what a filter is worth before you run it. Filters combine with AND; several values inside one filter combine with OR.

Every dataset, newest change first, with its files:

{
"maxItems": 20000,
"sortBy": "metadata_modified desc",
"includeDistributions": true
}

Environmental datasets published as CSV under an open licence, checked to see whether they still download:

{
"themes": ["http://standaarden.overheid.nl/owms/terms/Natuur_en_milieu"],
"formats": ["http://publications.europa.eu/resource/authority/file-type/CSV"],
"licences": ["http://creativecommons.org/licenses/by/4.0/deed.nl"],
"includeDistributions": true,
"includeLinkCheck": true,
"maxResourcesPerDataset": 5,
"maxItems": 2000
}

The data request register with the government's full answer to each request:

{
"datasetIds": [],
"includeDataRequests": true,
"includeDataRequestDetails": true,
"dataRequestPhases": ["Data is niet openbaar"],
"maxItems": 400
}

Pricing

Pay per event. You are charged for rows the Actor actually wrote, plus a small page fee for paging the catalogue, and the optional checks only when they returned something.

EventPrice
Actor start$0.054, once per run
Catalogue page scanned$0.004 per page of up to 500 datasets
Dataset$0.007
File$0.003
Register page scanned$0.004 per page of 10 entries
Data request$0.006
Data request detail read$0.006
Application, catalogue, group, organization$0.005 to $0.006
Data owner, theme, format, licence, source catalogue$0.003
Keyword$0.002
File link checked$0.004
File bytes probed$0.012
Data owner profile attached$0.006
DatasetsWhat it costs
100$0.76
1,000$7.06
10,000$70.13

Paid Apify plans get a discount on every event: 3.8 percent on Bronze, 7.4 percent on Silver, 11 percent on Gold and above.

Free users

Free Apify accounts get 10 rows per run as a preview, which is enough to see every column and decide. Any paid plan lifts the cap to whatever you set in Max Items, up to 1,000,000. Upgrade here.

Run it

  1. Create a free Apify account. New accounts come with $5 in free credit, which is about 700 datasets.
  2. Open the Actor, leave the input empty for a full sweep or pick a theme, a data owner or a format from the dropdowns.
  3. Click Start. A 10-row preview finishes in about 2 seconds; a full 20,380-dataset sweep with files runs at roughly 267 rows per second.
  4. Download the results as CSV, JSON, Excel or XML, or pull them from the dataset API.

Use with AI agents (MCP)

claude mcp add apify --transport http https://mcp.apify.com --header "Authorization: Bearer YOUR_APIFY_TOKEN"

Then ask your agent in plain language:

  • "List every Netherlands open dataset about air quality that is published as CSV under an open licence."
  • "Which Dutch public bodies publish the most datasets, and how many of theirs are not openly downloadable?"
  • "Show me the data requests where the government refused because the data is not public."

Troubleshooting

No results at all. A filter value the portal does not know returns zero rows rather than an error. Every dropdown here only offers values that exist, so the usual cause is combining two filters that have no overlap, for example a theme and a data owner that never publish together. Clear one filter and run again.

Fewer rows than I asked for. Max Items is a single budget shared by every row type. If you tick the extra registers, the datasets are collected first and can use the whole budget before the registers start. Raise Max Items or run a register on its own with no dataset filters.

A field is empty on every row. Most DCAT fields are optional upstream and the portal leaves them blank. legalFoundationLabel is filled on 3 percent of datasets, datasetQuality on 9 percent, spatialValues on under 1 percent. "Not Disclosed" means the portal withheld it; "N/A" means it does not apply. No column is empty on every dataset in the catalogue.

No community column on my rows. The Communities filter works but does not export: the portal indexes the community list for search and never returns it on the dataset record.

The run is slow. Plain dataset collection runs at about 267 rows per second, so 8,000 rows take 28 seconds. The link check and the byte probe fetch files from 23 different government servers, some of which take seconds to answer: 1,500 rows with both on and three files checked per dataset took 560 seconds. Lower "Files to check per dataset" to speed it up.

FAQ

QuestionAnswer
Do I need an API key for data.overheid.nl?No. The open data API is anonymous, and so is every page this Actor reads.
Is this the official CKAN API?It reads the portal's CKAN API and its own website. It is not run by or affiliated with data.overheid.nl.
How many datasets are there?20,380 as of 27 August 2026, across 26 harvest organizations and 23 source catalogues.
Is the data in Dutch?The records are, since the portal is Dutch-first: 18,956 are in Dutch, 1,389 in English, 35 in German. Themes, statuses, access rights, licences, formats and data request phases are also given in English.
Can I get only the machine-readable files?Yes. Filter by format, and every file row also carries a machineReadable flag.
What is a basisregistratie?One of the ten statutory Dutch national registers. 34 datasets are flagged as belonging to one, and there is a filter for them.
What are EU High Value Datasets?Categories from the EU Open Data Directive that member states must publish for free. 229 datasets carry one, in five categories.
Does it check that the downloads work?Optionally. The portal itself records a link status on only 8 percent of files; the link check does the other 92 percent, and found 9.0 percent of a 647 file sample dead. OGC services are asked for their GetCapabilities, so a live WFS is not reported as broken.
How fresh is it?Live. Every run reads the portal at that moment, and 5,825 catalogue records changed in the last 30 days.
Can I scrape one dataset by URL?Yes. Paste data.overheid.nl dataset or community URLs into Dataset or community URLs, or put CKAN slugs into Dataset IDs.

Browse the full ParseForge collection for more scrapers.

Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

Disclaimer: this is an unofficial tool, not affiliated with or endorsed by data.overheid.nl, KOOP or the government of the Netherlands. It collects only data that is already published openly, without logging in and without circumventing any access control. Records may contain names and contact details of public officials acting in their professional capacity; if you process them, GDPR, CCPA and PIPL obligations are yours as the data controller.